<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[accelerate shipping software]]></title><description><![CDATA[accelerate shipping software]]></description><link>https://dromeas.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>accelerate shipping software</title><link>https://dromeas.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 30 Sep 2026 13:34:31 GMT</lastBuildDate><atom:link href="https://dromeas.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Behavioral Bugs Are Still Slipping Through AI Code Review. Here's What Actually Catches Them.]]></title><description><![CDATA[Originally published on the Dromeas blog.
The pull request looks clean, the linter's green, the tests pass — and the code still does the wrong thing the moment someone hands it an input nobody thought]]></description><link>https://dromeas.hashnode.dev/behavioral-bugs-are-still-slipping-through-ai-code-review-here-s-what-actually-catches-them</link><guid isPermaLink="true">https://dromeas.hashnode.dev/behavioral-bugs-are-still-slipping-through-ai-code-review-here-s-what-actually-catches-them</guid><category><![CDATA[AI]]></category><category><![CDATA[Programming Blogs]]></category><category><![CDATA[Testing]]></category><category><![CDATA[webdev]]></category><dc:creator><![CDATA[Manos Saratsis]]></dc:creator><pubDate>Thu, 17 Sep 2026 15:36:47 GMT</pubDate><content:encoded><![CDATA[<p><em>Originally published on the <a href="https://dromeas.ai/blog/behavioral-bugs-ai-code-static-analysis-bug-tracing">Dromeas blog</a>.</em></p>
<p>The pull request looks clean, the linter's green, the tests pass — and the code still does the wrong thing the moment someone hands it an input nobody thought about. Not a crash, not a lint warning. Just the wrong answer, quietly.</p>
<p>I've been digging into why that keeps happening even as AI writes more and more of our code, and it turns out there's decent data on it now — not just vibes. So let's go through what the numbers actually say, why the tools most of us already run can't really catch this category of bug, and what I think actually works.</p>
<h2>What the 2026 data says</h2>
<p>First: people don't fully trust this code, and it's specifically about correctness, not whether it runs. Sonar's 2026 State of Code Developer Survey found 96% of developers don't fully trust that AI-generated code is functionally correct, and 61% agree that "AI often produces code that looks correct but isn't reliable." Only 48% say they always verify AI-assisted code before committing it. Worth sitting with that for a second: the complaint isn't "it doesn't compile." It's "it runs, it looks fine, and I don't actually know if it's right."</p>
<p>Second, when you break down what kind of bugs show up, it lines up with that. A 2026 empirical study ("Debt Behind the AI Boom") looked across five widely used AI coding tools — Copilot, Claude, Cursor, Gemini, Devin — and found code smells made up 89.3% of flagged issues (the stuff any linter catches fine). Correctness issues were a smaller slice, 6.0% — but the single most common one was "undefined variable or reference," almost 24,000 instances, which the researchers describe as code that "may look locally correct, but still fails to stay consistent with the surrounding context." That's a pretty good one-line description of the whole problem. More than 15% of commits from every single tool they studied introduced at least one issue, and 22.7% of the AI-introduced issues they tracked were still sitting in the repo, unnoticed, at the time of the study — some for nine months or longer.</p>
<p>Third, when this reaches production, it's not staying theoretical. New Relic's 2026 State of AI Coding report found 82% of organizations had at least one major production failure caused by AI code in the past six months, and 78% report a measurable spike in incidents tied to AI code overall — with AI-generated code introducing roughly 1.7x more critical runtime issues than human-reviewed code. CloudBees' 2026 State of Code Abundance report found something similar from a different angle: 81% of enterprise leaders report increased production issues tied to AI code, and their own summary of it stuck with me — "writing code is no longer the primary bottleneck, governing it is." Same conclusion, three separate 2026 surveys.</p>
<p>And the backdrop makes all of this harder to catch by hand. GitClear's 2026 research tracked eight quality signals across 623 million code changes from 2023 to 2026, and the trend lines aren't great: duplicated code blocks up 81% since 2023 (highest on record), copy-paste share of changed lines up from 9.4% to 15.7%, actual refactoring down about 70% over the same stretch. GitClear's own read on it: block duplication is the single risk signal most tied to defects and propagated bugs in the research they looked at. So more code is shipping, less of it is getting consolidated or re-examined, and reviewers have less bandwidth per line than ever.</p>
<p>Put simply: it's not that AI writes "bad" code. It clears the bars we're good at checking automatically — style, structure, the obviously dangerous patterns — and it's specifically weaker on whether the logic holds up for a real input, in a way that's now showing up as real production failures, not just review comments. That's worth being precise about, because it tells you exactly why the tools most teams already run don't solve it.</p>
<h2>What these bugs actually look like</h2>
<p>"Behavioral bug" is the term I keep reaching for, but it's worth knowing it's not the only name people use for this — you'll see the same basic idea called logic bugs or logic errors (the plain-English default), semantic bugs (more of an academic/compiler-literature term — the code is syntactically fine but means the wrong thing), correctness bugs or correctness issues (the term the arxiv paper above uses), business logic bugs or business logic vulnerabilities (the AppSec framing), and functional bugs (QA/testing terminology). Some people just call them silent bugs or silent failures, which honestly might be the most useful name of the bunch — it points at the actual property that makes them dangerous: nothing crashes, nothing throws, nothing shows red in CI.</p>
<p>In practice, most of what falls under that umbrella breaks down into a handful of recurring shapes:</p>
<p><strong>Boolean logic bugs</strong> — an inverted condition, the wrong operator, a guard clause that lets through exactly the case it was written to block. Classic version: a permission check written with <code>||</code> where it needed <code>&amp;&amp;</code>, so access gets granted if any one condition is met instead of requiring all of them.</p>
<p><strong>Reference correctness bugs</strong> — stale closures, the wrong variable captured, the wrong object mutated once two calls overlap. Classic version: a loop that captures its loop variable by reference in a callback, so every callback ends up pointing at the last value instead of the one that was current when it was created.</p>
<p><strong>Boundary and indexing bugs</strong> — the everyday off-by-one: loop bounds, first/last-element handling, pagination math. Classic version: a "load more" offset computed as <code>page * pageSize</code> instead of <code>(page - 1) * pageSize</code>.</p>
<p><strong>Normalization-symmetry bugs</strong> — a value gets normalized on write but not on read (case, whitespace, encoding), so a lookup silently misses. Classic version: emails lowercased at signup but not at login.</p>
<p><strong>Nullability bugs</strong> — a new code path where a value can legitimately be absent, and nothing downstream accounts for it.</p>
<p><strong>Shared-state concurrency bugs</strong> — races between concurrent writes, a rollback to a stale captured value, a cache write that clobbers something fresher.</p>
<p>None of these are exotic. Every engineer has shipped at least one of each at some point. What's changed is the volume and the plausibility — code that reads as confidently correct, generated fast enough that nobody's tracing each new input by hand anymore.</p>
<h2>Why static analysis can't really get at this</h2>
<p>Tools like SonarQube, Semgrep, and CodeQL work by pattern matching — an AST shape, a taint path from untrusted input to a dangerous sink, something that matches a known CWE. Genuinely useful, and it's why these tools are worth running. But it only works if the bug matches a pattern someone already wrote a rule for.</p>
<p>Most behavioral bugs don't work that way. Gecko Security put this well in a piece on why static analysis struggles with business logic: a lot of these bugs are about something being <em>missing</em> — a missing auth check, a missing validation step — rather than something dangerous being <em>present</em> that a rule can flag. Taint analysis can tell you untrusted data reaches a sensitive spot. It can't tell you whether the authorization logic guarding that spot is actually correct. And a lot of the bugs that cause real incidents span multiple files — their example is a real Cal.com auth-bypass bug that needed three separate issues chained together across different files, each looking fine on its own.</p>
<p>So that's the real limitation: pattern matching can tell you something <em>looks like</em> a shape it's seen before. It can't tell you that a specific input, actually walked through the logic, produces a different result than it should. That requires tracing behavior against intent — and it happens to be exactly where the data above says AI-generated code is weakest.</p>
<p>It's also why I think "detect that AI wrote this, then review it harder" — which a few vendors have shipped, SonarQube's AI Code Assurance being one — is sorting on the wrong axis. It's routing on <em>who</em> wrote the code, when the thing that actually matters is <em>what the code does</em>. The question was never "who wrote this" — it's "does this do the right thing for a real input," and that needs tracing, not detection.</p>
<h2>Where this leaves Bug Tracing</h2>
<p>This is basically the problem we built Bug Tracing to solve inside Dromeas — trace real inputs through the actual changed code and check what comes out, instead of scanning for patterns. The six categories above are exactly what it scores every change against, and it only reports a finding if it can name the concrete input that triggers it — no reproducible trace, no report.</p>
<p>And because "review a diff" and "audit a whole system" are genuinely different jobs, it runs at whichever scope your team actually works at: on the pull request itself (changed lines plus their blast radius), on a branch or release after merge for teams doing trunk-based development, or as an on-demand sweep across a whole repository when you just want a second opinion on a system nobody's looked at closely in a while.</p>
<p>If this is a problem you're running into, happy to show you how it works — the feature page is at <a href="https://dromeas.ai/bug-tracing">dromeas.ai/bug-tracing</a>.</p>
]]></content:encoded></item><item><title><![CDATA[Shadow AI in Your Codebase: The Governance Gap Most CISOs Haven't Mapped Yet]]></title><description><![CDATA[Your AI code governance policy probably covers the tools you approved. It says nothing about the ones your developers are actually using. That gap has a name now, shadow AI, and in 2026 it's stopped b]]></description><link>https://dromeas.hashnode.dev/shadow-ai-in-your-codebase-the-governance-gap-most-cisos-haven-t-mapped-yet</link><guid isPermaLink="true">https://dromeas.hashnode.dev/shadow-ai-in-your-codebase-the-governance-gap-most-cisos-haven-t-mapped-yet</guid><category><![CDATA[Security]]></category><category><![CDATA[AI]]></category><category><![CDATA[Devops]]></category><category><![CDATA[cybersecurity]]></category><dc:creator><![CDATA[Manos Saratsis]]></dc:creator><pubDate>Sun, 13 Sep 2026 13:38:32 GMT</pubDate><content:encoded><![CDATA[<p>Your AI code governance policy probably covers the tools you approved. It says nothing about the ones your developers are actually using. That gap has a name now, shadow AI, and in 2026 it's stopped being a vague compliance worry and turned into something you can actually measure: developers self-report significant use of unsanctioned code assistants, often specifically to get around a review process or licensing restriction they find slow.</p>
<p>Most CISO-facing writing about AI code risk focuses on the code that goes through review — how good the model is, whether vulnerabilities get caught before merge (we've covered that ground ourselves in our <a href="https://dromeas.ai/blog/ciso-guide-ai-generated-code">CISO's guide to AI-generated code</a>). This piece is about the code that never reaches that conversation at all: written by a tool nobody approved, reviewed by no one, invisible to the governance program on paper because the governance program doesn't know it exists.</p>
<p>A note on sourcing before we go further: the stats below are drawn from a mix of primary surveys and industry compilations, linked inline so you can weigh them yourself. Shadow AI research is still a young field, and precision varies by source.</p>
<h2>The visibility gap is bigger than most security teams assume</h2>
<p>Across recent industry surveys compiled by <a href="https://airia.com/blog/shadow-ai-statistics-key-data-points-every-ciso-needs-in-2026/">Airia</a>: a majority of AI users bring their own tools without IT approval, and a majority of employees use at least one non-sanctioned AI tool, with shadow AI usage growing sharply year over year. On the security-team side of the ledger, the picture is just as lopsided: most security leaders say they lack visibility into what AI tools employees are actually using, and a large share say they can't even produce an inventory of AI tools in use across the organization.</p>
<p>That's not really a policy failure, it's a detection failure. You can't govern what you can't see, and right now most organizations can't see most of it.</p>
<h2>Why source code is the highest-stakes category of shadow AI, specifically</h2>
<p>General shadow-AI risk usually gets framed around things like customer data or confidential documents pasted into a public chatbot. Source code deserves its own category, for two reasons that don't really apply to those other cases.</p>
<p>First, the exposure compounds. When a developer pastes proprietary code into an unsanctioned tool to get a suggestion, that code may become training data for a model neither your company nor your competitors' companies control — an exposure that data classification programs built for documents and spreadsheets typically aren't built to catch. <a href="https://airia.com/blog/shadow-ai-statistics-key-data-points-every-ciso-needs-in-2026/">Industry compilations</a> put source code leakage at a disproportionate share of sensitive-data events involving AI tools, given that code is a fraction of the total data most orgs handle day to day.</p>
<p>Second, and more specific to engineering: shadow AI in the codebase isn't just a data-handling problem, it's a governance blind spot on the thing you actually ship. A document that leaks is embarrassing. A code change that ships to production with no review record, written by a tool with no logged provenance, is a different category of risk entirely — it's not about what left the building, it's about what's now running inside it.</p>
<h2>The policy vacuum is the real story, not the tool sprawl</h2>
<p>It'd be easy to read those adoption numbers and conclude the fix is "ban unsanctioned tools." The data suggests that's not really where most organizations are stuck. Most CISOs, <a href="https://airia.com/blog/shadow-ai-statistics-key-data-points-every-ciso-needs-in-2026/">per Airia's compilation</a>, already name shadow AI a top concern — so awareness isn't the gap. The gap shows up downstream of awareness: roughly half of organizations have no formal policy governing external AI tool use at all, only a minority have any kind of formal detection program in place to know when policy is being violated, and <a href="https://jumpcloud.com/blog/11-stats-about-shadow-ai-in-2026">JumpCloud's research</a> puts AI-related incidents at taking meaningfully longer to identify and contain, specifically because tracking data flows to and from third-party AI tools adds a layer of complexity most incident response processes weren't built for.</p>
<p>In other words: most CISOs already know the risk exists, most organizations haven't written down a policy for it, and even where a policy exists, most can't yet detect whether it's being followed. That's three separate gaps, and "buy a detection tool" only closes one of them.</p>
<h2>Why banning the tool doesn't fix what actually matters</h2>
<p>Here's the uncomfortable practical reality for an engineering leader reading this: developers reach for unsanctioned tools largely because the sanctioned path is slower, more restrictive, or just doesn't cover the workflow they're already in — a different IDE, a different agent, a preference that isn't going away because a policy memo went out. A governance approach that depends on every developer voluntarily routing every change through one approved tool is fighting the same battle every "approved software list" has fought for twenty years, and losing for the same reasons.</p>
<p>The more durable approach is to stop trying to govern the tool and start governing the output, regardless of which tool produced it. That's a different question: not "was this written by an approved assistant," but "did this change get checked before it shipped, and can I actually show that" — a question you can answer the same way whether the diff came from a developer typing by hand, Copilot, Cursor, Claude Code, or a tool nobody in security has ever heard of.</p>
<h2>What actually gives you the inventory: an AI Bill of Materials</h2>
<p>Go back to those visibility stats for a second: most security leaders can't produce an inventory of AI tools in use, and most lack visibility into usage altogether. Both are really the same problem stated two ways: nobody can produce a list of what's actually in the codebase and where it came from.</p>
<p>That's the specific gap an AI Bill of Materials is built to close. For every product Dromeas covers, whether that product lives in a single repo or is spread across a dozen, it generates a machine-readable CycloneDX 1.6 ML-BOM and a human-readable report — covering detected AI/ML components, data flows, risk classification, and EU AI Act obligations — when a product actually invokes AI or machine-learning systems, refreshed with every release. For multi-repo products, the documentation process assembles evidence across the whole product rather than stopping at a single repo boundary.</p>
<p>Worth being precise about what this is and isn't. It's not a forensic attribution system for which coding assistant wrote each line, and it doesn't calculate a trustworthy percentage of AI-authored code — true authorship is genuinely hard to pin down from the outside, and we'd rather be honest about that limit than oversell it. What it is: a source-backed, regenerated-every-release record of what AI/ML is actually present in a given release, and what checks it passed — the release-level inventory those visibility stats say most organizations currently can't produce.</p>
<p>That reframes the shadow AI problem in a useful way. You may never get perfect visibility into every assistant a developer has installed on their laptop. You can get a lot closer to complete visibility into what shipped in your last release and whether it was reviewed, because that's a question about your repos, not about your developers' tool choices — and it stays answerable release after release, whether your product is one repo or many.</p>
<p>None of this replaces having an actual AI-usage policy; it just means the policy isn't the only thing standing between you and knowing what's in production. The same idea shows up on the model side too, in our post on <a href="https://dromeas.ai/blog/bring-your-own-model-self-hosted">running your own models</a> — governance has to travel with the code and the release, not with whichever tool or model happened to write it.</p>
<h2>A practical first step, before you write a bigger policy</h2>
<p>If your organization is somewhere in the "we know it's a problem, we haven't mapped it" stage, which the data above suggests is most organizations, the useful first move isn't a sweeping new AI-usage policy. It's mapping what's actually flowing through your repos today:</p>
<ol>
<li>Look at commit and PR patterns for signals, not just tool telemetry — unusual diff sizes, commit message patterns, or velocity spikes that don't match your known toolchain usage often surface shadow AI faster than trying to detect the tool itself.</li>
<li>Separate "which tool" from "was it checked." An AI BOM regenerated every release gives you the second half of that answer automatically, even before you've solved the first half.</li>
<li>Write the policy after you've done the mapping, not before. A policy built on a guess about usage patterns tends to under-cover the tools people are actually reaching for.</li>
<li>Treat this as a release-management problem, not just a security-tooling problem. The organizations with the least shadow-AI exposure aren't the ones with the strictest tool bans, they're the ones where nothing reaches production without passing the same check and leaving a record behind, no matter where it came from.</li>
</ol>
<p>Shadow AI isn't going away just because it got a name. It'll go away, if it does, because governance stopped being about which tool a developer chose and started being about whether what they shipped was actually safe, and provable.</p>
<hr />
<p><em>Dromeas checks every pull request, trunk commit, and local diff — quality, security, compliance, docs, testing, instrumentation — through MCP, whichever coding agent your developers are already using, and rolls the result into an AI Bill of Materials for every product, per repo, refreshed with every release. Full piece, with more on the review layer and the CISO's guide to AI-generated code at scale, originally published at dromeas.ai: <a href="https://dromeas.ai/blog/shadow-ai-in-your-codebase">https://dromeas.ai/blog/shadow-ai-in-your-codebase</a></em></p>
]]></content:encoded></item><item><title><![CDATA[Which AI Model Writes the Most Secure Code? What the 2026 Data Actually Shows]]></title><description><![CDATA[Ask five engineering leaders which AI coding model is "safe," and you'll get five confident, contradictory answers — most of them based on a vendor's marketing page rather than an actual security test]]></description><link>https://dromeas.hashnode.dev/which-ai-model-writes-the-most-secure-code-what-the-2026-data-actually-shows</link><guid isPermaLink="true">https://dromeas.hashnode.dev/which-ai-model-writes-the-most-secure-code-what-the-2026-data-actually-shows</guid><category><![CDATA[AI]]></category><category><![CDATA[Security]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[webdev]]></category><dc:creator><![CDATA[Manos Saratsis]]></dc:creator><pubDate>Sun, 13 Sep 2026 13:35:54 GMT</pubDate><content:encoded><![CDATA[<p>Ask five engineering leaders which AI coding model is "safe," and you'll get five confident, contradictory answers — most of them based on a vendor's marketing page rather than an actual security test. <a href="https://www.veracode.com/blog/2026-genai-code-security-report-ai-risk/">Veracode's newly released 2026 GenAI Code Security Report</a> puts a real number on the question, and it's not a comfortable one: across the field, AI-generated code still fails a security test roughly 44% of the time. The overall pass rate sits at 56%, statistically flat versus last year's 55%, even though AI now writes an estimated half of all committed code in the organizations tested.</p>
<p>Flat security performance at double the volume isn't a wash, it's the same failure rate landing on twice as much production code. So let's actually dig into what the 2026 data shows — where the risk concentrates, why "just pick the best model" is a harder answer than it sounds, and what a sane response looks like for a team that can't stop shipping AI-assisted code while it figures this out.</p>
<h2>The plateau nobody's marketing slide mentions</h2>
<p>A separate <a href="https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report">CodeRabbit study</a> reported 2.74x more vulnerabilities in AI-co-authored pull requests than in human-only pull requests, from 470 open-source PRs (320 AI-co-authored, 150 human-only) — not from Veracode. CodeRabbit notes an important limitation: authorship was inferred from signals rather than confirmed ground truth. Different datasets, same direction: model capability and secure output don't rise on the same curve.</p>
<p>Veracode's 2026 data doesn't restate that exact multiple, but its own headline finding rhymes with it: pass rates haven't meaningfully moved year over year, from 55% to 56%, which means the popular narrative of "the models are getting safer as they get smarter" just doesn't hold up against Veracode's own test suite.</p>
<p>The practical takeaway: if your security posture assumes "next quarter's model upgrade fixes this," the data says don't hold your breath. The fix has to happen at the review layer, not the model layer.</p>
<h2>Not all vulnerabilities are equal — and that's the more useful finding</h2>
<p>Aggregate pass rates hide a much sharper pattern once you break results out by vulnerability class, per Veracode's 2026 report:</p>
<ul>
<li>SQL injection: 83% pass rate</li>
<li>Cryptographic implementation: 87% pass rate</li>
<li>Cross-site scripting (XSS): 15% pass rate</li>
<li>Log injection: 12% pass rate</li>
</ul>
<p>That's a nearly 75-point gap between the vulnerability classes models handle well and the ones they consistently miss. Models do comparatively well on classes with well-known, syntactically obvious fixes (parameterized queries, standard crypto libraries), and do badly on classes that require actually tracking how untrusted data flows through an application.</p>
<p>That's a dataflow problem, not a syntax problem — and exactly the kind of defect a single-pass, single-model review is least likely to catch.</p>
<h2>Model selection has become a security decision</h2>
<p><a href="https://www.veracode.com/blog/2026-genai-code-security-report-ai-risk/">Veracode's report</a> puts it plainly: model selection has become a security decision. The data backs it up:</p>
<ul>
<li>The best-performing model in the report, GPT-5.5, leads at a 68% pass rate.</li>
<li>More than half of all tested models cluster tightly in the 50–53% range.</li>
<li>Coding-specialized models averaged 51% on security tasks — general-purpose models averaged 52%.</li>
<li>Reasoning models scored a bit higher (56%) than non-reasoning variants (51%).</li>
<li>Raw model size showed little correlation with security outcomes.</li>
</ul>
<p>There's no dominant "most secure" model you can standardize on and call the problem solved. Even the frontrunner still fails security checks roughly a third of the time.</p>
<h2>Why cross-checking beats betting on one model</h2>
<p>If no single model clears the bar, the answer isn't to keep hunting for one that will — it's to stop treating any one model's verdict as the last word. Dromeas's review pipeline runs independent Security, Quality, Compliance, and Bug Tracing analysts across every diff, with findings merged and checked against changed files and lines, specifically to catch the case where one model's blind spot is another model's obvious flag. Given a 70-point swing in pass rate by vulnerability type, a review process that only sees through one model's eyes inherits exactly one model's blind spots — and per this data, XSS and log injection are close to universal ones.</p>
<h2>What to actually do with this data this quarter</h2>
<ol>
<li>Put security pass rates in the procurement conversation — ask for vulnerability-class-level data, not just an aggregate "safe" claim.</li>
<li>Verify security as code is written, not just before release — catching a dataflow bug at the PR or local-diff stage is cheaper and faster than a pre-release scan.</li>
<li>Prioritize by exploitability, not raw finding count.</li>
<li>Write down the policy — which models are approved for which kinds of code, what happens when a model's output fails a check, who signs off.</li>
</ol>
<p>The uncomfortable truth in this year's data is that there's no shortcut past review — not a smarter model, not a bigger one, not a specialized one. The AI coding era didn't remove the need for a second set of eyes on security, it just changed whose eyes those need to be, and how fast they need to look.</p>
<hr />
<p><em>Full piece, with more on Dromeas's multi-model review pipeline, originally published at dromeas.ai: <a href="https://dromeas.ai/blog/which-ai-model-writes-most-secure-code">https://dromeas.ai/blog/which-ai-model-writes-most-secure-code</a></em></p>
]]></content:encoded></item><item><title><![CDATA[MCP for Code Review: What It Is, and How to Add a Review Layer to Claude Code or Cursor]]></title><description><![CDATA[Originally published on the Dromeas blog.
Your coding agent writes code. MCP is how you give it a reviewer. Here's what the protocol actually does — and a working setup you can copy in minutes.
MCP, i]]></description><link>https://dromeas.hashnode.dev/mcp-for-code-review-what-it-is-and-how-to-add-a-review-layer-to-claude-code-or-cursor</link><guid isPermaLink="true">https://dromeas.hashnode.dev/mcp-for-code-review-what-it-is-and-how-to-add-a-review-layer-to-claude-code-or-cursor</guid><category><![CDATA[mcp]]></category><category><![CDATA[claude-code]]></category><dc:creator><![CDATA[Manos Saratsis]]></dc:creator><pubDate>Wed, 09 Sep 2026 13:32:04 GMT</pubDate><content:encoded><![CDATA[<p><em>Originally published on the <a href="https://dromeas.ai/blog/mcp-review-layer-guide">Dromeas blog</a>.</em></p>
<p>Your coding agent writes code. MCP is how you give it a reviewer. Here's what the protocol actually does — and a working setup you can copy in minutes.</p>
<h2>MCP, in one paragraph</h2>
<p>The Model Context Protocol is an open standard that lets an AI client — Claude Code, Cursor, Windsurf, Zed — call external tools through a uniform interface. Instead of pasting output between your terminal and a review dashboard, the agent invokes a tool like <code>review_pull_request</code> directly and receives the result as structured data it can reason about. Tool support, not copy-paste, is what makes an agentic loop possible: the agent can act, observe the verdict, and repair — the loop we described in <a href="https://dromeas.ai/blog/loop-engineering-with-dromeas">loop engineering</a>.</p>
<h2>Why a review layer belongs in the loop</h2>
<p>A coding agent with no review layer grades its own homework. It writes a diff, checks that it compiles, and declares victory. A review layer changes the economics: the same agent submits its diff, gets an independent multi-model verdict — security, quality, compliance — and fixes what it finds before a human spends a minute on it. Self-review before the PR is the single highest-leverage place to insert checking, because the cost of a fix there is one tool call, not a review round-trip.</p>
<h2>The setup, step by step</h2>
<ol>
<li><strong>Connect your repositories.</strong> Sign up at dromeas.ai and connect GitHub, GitLab or Bitbucket. Dromeas builds a typed code map of your repos so review verdicts come with real context.</li>
<li><strong>Get your MCP endpoint.</strong> Open Settings &gt; MCP in the Dromeas app and copy your workspace's MCP server URL and API key. One endpoint covers Claude Code, Cursor, Windsurf, Zed and any other MCP-capable client.</li>
<li><strong>Add the server to your client.</strong> In Claude Code run <code>claude mcp add</code> with the URL; in Cursor, add the server under Settings &gt; MCP. The review tools appear in your agent's tool list immediately.</li>
<li><strong>Ask your agent to self-review.</strong> Before opening a PR, ask your coding agent to submit the diff for review. It calls <code>review_pull_request</code> or <code>review_local_diff</code>, gets a multi-model verdict, and can fix what it finds before a human ever looks.</li>
</ol>
<h2>What your agent can actually do once connected</h2>
<p>The Dromeas MCP server exposes the full review surface as tools: <code>review_pull_request</code> and <code>review_local_diff</code> for verdicts, <code>code_finder_search</code> and the code-map tools for cheap context before edits, <code>get_findings</code> and the fix tools for acting on results, and release tools like <code>describe_release_state</code> for go/no-go decisions. The point is that review isn't a dashboard your agent can't see — it's a function it can call.</p>
<p>Because every tool call is scoped to your workspace and logged, you keep the audit trail that matters when agents start merging on their own. That's the same provenance property the <a href="https://dromeas.ai/blog/ciso-guide-ai-generated-code">CISO checklist</a> depends on.</p>
<hr />
<p>Connect a repo, copy your MCP endpoint, and your coding agent gets a six-agent review council it can call from the terminal. <a href="https://dromeas.ai/blog/loop-engineering-with-dromeas">Read about loop engineering</a> for the fuller picture of what the loop looks like end to end.</p>
]]></content:encoded></item><item><title><![CDATA[PR Review vs. Trunk Review: A Practical Guide to Choosing (or Combining) Both]]></title><description><![CDATA[Originally published on the Dromeas blog.
PR review and trunk-based review solve different problems, and most teams already run a hybrid without naming it. This guide is backed by data from 100,000+ r]]></description><link>https://dromeas.hashnode.dev/pr-review-vs-trunk-review-a-practical-guide-to-choosing-or-combining-both</link><guid isPermaLink="true">https://dromeas.hashnode.dev/pr-review-vs-trunk-review-a-practical-guide-to-choosing-or-combining-both</guid><category><![CDATA[code review]]></category><category><![CDATA[Artificial Intelligence]]></category><dc:creator><![CDATA[Manos Saratsis]]></dc:creator><pubDate>Wed, 09 Sep 2026 13:30:30 GMT</pubDate><content:encoded><![CDATA[<p><em>Originally published on the <a href="https://dromeas.ai/blog/pr-review-vs-trunk-review-guide">Dromeas blog</a>.</em></p>
<p>PR review and trunk-based review solve different problems, and most teams already run a hybrid without naming it. This guide is backed by data from <a href="https://dromeas.ai/blog/state-of-ai-coding-2026">100,000+ real pull requests and 24,000 trunk commits across 500+ open-source repos</a>.</p>
<p>When we looked at that data, one finding kept surfacing in different forms: the PR-versus-trunk question isn't binary. Teams that think they've picked one are usually already running a hybrid, they just haven't named it. Here's what the data actually supports about when each one earns its keep.</p>
<h2>What PR review is for</h2>
<p>PR review is built for the moments that benefit from a checkpoint: large or risky changes, anything that needs discussion before it merges. That matters because PR size in the wild is bimodal — plenty of small, quick changes, but a heavy tail of 1,000+ line PRs that are exactly where a pre-merge gate is worth the friction it adds. The PR is also the only place where a change is still cheap to reject: once it merges, every fix is a new change instead of a revision.</p>
<h2>What trunk review is for</h2>
<p>Trunk review is built for everything else. Direct-to-trunk commits in our sample averaged 63x smaller than PRs — fast, low-friction, and increasingly the shape of how coding agents actually commit. Waiting for a full PR gate on every micro-change doesn't match agent-speed workflows, and it doesn't need to. The mistake is assuming "no PR" means "no review" — trunk commits can be reviewed after the fact against the same bar, without blocking the commit path.</p>
<h2>The tail is where the time goes</h2>
<p>The part that costs teams the most time isn't the average case in either lane — it's the tail. Rejected PRs took 6x longer to resolve than approved ones in our data. That's where review capacity actually gets consumed: not on the easy approvals, but on the changes that go back and forth. Any review strategy that doesn't have an answer for the tail — smaller initial diffs, automated first-pass review, faster feedback — will feel slow no matter which lane you picked.</p>
<h2>A decision framework that matches reality</h2>
<p>Instead of "PRs for everything" or "trunk for everything," score each change on four axes and let the lane follow:</p>
<table>
<thead>
<tr>
<th>Signal</th>
<th>Lean PR gate</th>
<th>Lean trunk + post-merge review</th>
</tr>
</thead>
<tbody><tr>
<td>Change size</td>
<td>Hundreds of lines and up; the heavy tail</td>
<td>Tens of lines; single-purpose commits</td>
</tr>
<tr>
<td>Blast radius</td>
<td>Touches shared contracts, auth, billing, migrations</td>
<td>Local, reversible, behind a flag</td>
</tr>
<tr>
<td>Contributor type</td>
<td>New contributor, unfamiliar area, cross-team change</td>
<td>Owner of the area, or an agent doing a scoped fix</td>
</tr>
<tr>
<td>Discussion value</td>
<td>Design decisions others need to weigh in on</td>
<td>Mechanical or already-agreed work</td>
</tr>
</tbody></table>
<p>Most teams discover their real policy is already this table, applied informally. Making it explicit does two things: it stops the arguments about whether a given change "deserved" a PR, and it makes the post-merge lane a first-class citizen instead of an unreviewed loophole.</p>
<h2>Why most teams will end up running both</h2>
<p>The forces pushing the two lanes apart are getting stronger, not weaker. Risky changes are getting larger (more generated code per PR), and routine changes are getting smaller and more frequent (agents committing at agent speed). A single gate tuned for one of those shapes fails the other: too much friction on the small stuff, too little scrutiny on the big stuff.</p>
<p>The workable shape is both lanes with the same rigor on each: a real gate on the PR for the changes that benefit from one, and automatic, same-standard review on every trunk commit for the ones that don't. That's the model Dromeas implements — the same six-agent pipeline and multi-model council runs on PRs and on trunk commits alike, so the lane choice is a workflow decision, not a quality decision.</p>
<h2>The data behind this guide</h2>
<p>104,968 PRs and 23,964 trunk commits measured across 503 open-source repos — size distributions, review latency, and the slow-motion rejection problem in full.</p>
<p><a href="https://dromeas.ai/blog/state-of-ai-coding-2026">Read the State of AI Coding 2026</a> · <a href="https://dromeas.ai/blog/pr-vs-trunk-what-code-review-actually-looks-like">See the charts</a></p>
]]></content:encoded></item><item><title><![CDATA[Run code review and releases from a conversation]]></title><description><![CDATA[We just shipped chat as a first-class interface into Dromeas.
Instead of clicking through dashboards, you can now just ask: "review PR 247 on the payment service," "are we good to release v1.9.0?," "f]]></description><link>https://dromeas.hashnode.dev/run-code-review-and-releases-from-a-conversation</link><guid isPermaLink="true">https://dromeas.hashnode.dev/run-code-review-and-releases-from-a-conversation</guid><dc:creator><![CDATA[Manos Saratsis]]></dc:creator><pubDate>Tue, 01 Sep 2026 06:33:56 GMT</pubDate><content:encoded><![CDATA[<p>We just shipped chat as a first-class interface into Dromeas.</p>
<p>Instead of clicking through dashboards, you can now just ask: "review PR 247 on the payment service," "are we good to release v1.9.0?," "fix the top security finding and open a PR." Dromeas reads the diff, queries the Code Map for blast radius, runs the same quality/security/compliance agents that guard your trunk, and shows the result as a live status card — not a wall of text.</p>
<p>A few things worth calling out for anyone building similar agentic UX:</p>
<p>It's not a separate system. The chat calls the exact same MCP primitives (get_findings, code_map_search, run_finding_fix, approve_pull_request, etc.) that our IDE integrations for Claude, Cursor, and Copilot use. Start a release check in Cursor, see it finish in the chat.
Autonomy is a dial, not a toggle. Every workspace sets a default — manual, observe, assist, or auto — and you can override it per conversation. Manual shows a confirm card before anything ships; auto acts inside caps you set (file budget, severity threshold, model cost) and reports back after.
Structured over conversational-only. Long-running actions (a review, a fix, a doc run) return a live card with step, progress, and a deep link — updating in place instead of dumping another paragraph into the thread.</p>
<p>Video walkthrough: <a href="https://youtu.be/ypCt0d8sGis">https://youtu.be/ypCt0d8sGis</a>
Try it: <a href="https://dromeas.ai/chat">https://dromeas.ai/chat</a></p>
]]></content:encoded></item><item><title><![CDATA[Claude Code's ultrareview vs Dromeas Code Review with LLM council]]></title><description><![CDATA[Originally published at dromeas.ai

We heard about Claude Code's ultrareview and got excited — a cloud-run, multi-agent deep review sounded like exactly the kind of thing worth building a workflow aro]]></description><link>https://dromeas.hashnode.dev/claude-code-s-ultrareview-vs-dromeas-code-review-with-llm-council</link><guid isPermaLink="true">https://dromeas.hashnode.dev/claude-code-s-ultrareview-vs-dromeas-code-review-with-llm-council</guid><dc:creator><![CDATA[Manos Saratsis]]></dc:creator><pubDate>Thu, 20 Aug 2026 14:57:10 GMT</pubDate><content:encoded><![CDATA[<blockquote>
<p>Originally published at <a href="https://dromeas.ai/blog/claude-ultrareview-vs-dromeas-code-review">dromeas.ai</a></p>
</blockquote>
<p>We heard about Claude Code's ultrareview and got excited — a cloud-run, multi-agent deep review sounded like exactly the kind of thing worth building a workflow around.</p>
<p>So we pointed it at changes in our own repo and compared it against Dromeas code review: three analyzers (quality, security, compliance) cross-checked by an LLM council. Dromeas held up well in that first pass.</p>
<p>That result was interesting enough that we wanted a harder, more neutral test: a large, real, independently-approved pull request from a codebase neither tool had any stake in. So we picked openclaw/openclaw — a public, actively-developed agentic coding tool — and went looking for its biggest recently-merged, genuinely-reviewed PR. That led us to openclaw#124250, 31 files changed, approved by a human reviewer, and we ran the same head-to-head again.</p>
<h2>The PR</h2>
<p>"Preserve ClawHub external source identity and expose only supported actions" — merged, approved by a human reviewer (not a bot self-merge), XL size: 31 files changed, +1,064/−116 lines, spanning the Control UI, macOS, iOS, and Android clients plus the backend that serves them.</p>
<p>The bug it fixes: ClawHub's search API returns each result's source under a nested <code>install.reference</code> field, but the client code expected a flat <code>installRef</code>. Every external search result silently fell through to a synthesized <code>@owner/slug</code> reference — quietly pointing installs at a different publisher's skill than the one the operator actually picked. An identity-spoofing bug in a skill-installation flow, fixed across five client surfaces.</p>
<h2>What each tool found</h2>
<p><strong>ultrareview:</strong> 1 finding, nit severity — a duplicate test assertion in an Android test file, unrelated to the identity-spoofing bug the PR exists to fix.</p>
<p><strong>Dromeas's LLM council:</strong> 29 candidate findings raised, 17 kept after cross-verification. Three models (Opus 5, DeepSeek V4 Pro, GPT-5.6 Terra) independently analyzed the diff, then a decider cross-checked each finding. All 12 quality findings and all 5 security findings held up; 12 compliance findings were flagged as duplicates of already-caught security issues or dropped outright, with the report explaining why for each.</p>
<p>None of Dromeas's 17 kept findings overlap with ultrareview's one — not because ultrareview did a bad job reading the diff, but because questions like "is this credential field masked" or "does this action get an audit trail" were never in its scope. Full breakdown, cost comparison (~$5 for the full council run vs. $5–25 typical for ultrareview), and the four findings flagged for manual triage are in the full post →</p>
<p><a href="https://dromeas.ai/blog/claude-ultrareview-vs-dromeas-code-review">Read the full comparison</a></p>
]]></content:encoded></item></channel></rss>