Research & technology

We reviewed 50,000+ deploys. The AI coding boom doesn’t show a bug boom.

Laura Cressman
October 6, 2026

AI is helping teams write more code, and the assumption is that more bugs will follow. But that  isn't what we're seeing.

We followed the same group of QA Wolf customers from the start of 2025 through September 2026. They deployed about 30% more often, while the number of bugs we found fell. Bugs reported per deploy dropped by more than a third.

These teams have comprehensive end-to-end tests that check whether their applications work the way users expect. Each deploy that triggers those tests gives us a record of what shipped and which bugs our QA engineers verified and reported, mostly before release.

The data doesn't tell us whether AI caused the increase in release velocity. But it does show that, as AI becomes part of how software gets built, teams with robust testing can deploy more often without a corresponding rise in bugs. Here's what we found, what might explain it, and how we measured it.

AI is writing more of the code

The way software gets built has changed more in the last two years than in the two decades before. The shift is well documented. Google's 2025 DORA report, a survey of nearly 5,000 technology professionals, found that 90% now use AI at work and linked AI adoption to higher software delivery throughput (DORA).

We see it in our own data too. About a third of commits in August and September 2026 carried an explicit AI tool credit, such as "Co-authored-by: Claude" (more details on this in a future post).

Teams are deploying more often

We measured release frequency by counting deploys that triggered our test suites. For most teams that's a deploy to a test environment like staging, usually after code is merged.

Between Q1 2025 and Q3 2026, the teams in our group went from around 6,400 deploys tested per quarter to over 8,000, up 31%, with half of that increase coming after Q1 2026.

As you can see, the range is wide and some teams are even shipping less frequently than in Q1 2025 but nearly a third of teams more than doubled their release velocity.

Deploys rose overall, but customers varied widely
Deploys tested, indexed to Q1 2025 = 100. Combined, deploys rose 31%.
Source: QA Wolf platform data, same group of customers, Q1 2025 – Q3 2026.
Deploys rose overall, but customers varied widely
QuarterValueChange
Q1 2025All customers: 100; Middle half: 100–100Starting point (Q1 2025 = 100)
Q2 2025All customers: 117; Middle half: 84–126All customers +17% vs Q1 2025
Q3 2025All customers: 123; Middle half: 86–155All customers +23% vs Q1 2025
Q4 2025All customers: 113; Middle half: 82–158All customers +13% vs Q1 2025
Q1 2026All customers: 115; Middle half: 79–170All customers +15% vs Q1 2025
Q2 2026All customers: 131; Middle half: 80–171All customers +31% vs Q1 2025
Q3 2026All customers: 131; Middle half: 69–207All customers +31% vs Q1 2025

Did more deploys mean more bugs? No.

A quick note on what counts as a bug here, because it matters.

We build and maintain automated end-to-end tests for our customers, so they catch the kind of problems users would actually hit. When a test fails, one of our QA engineers investigates. Either the test needs updating, or it's a real bug, in which case they reproduce it and file a report.

So these aren't raw test failures or flaky tests. Every bug here was verified by a person before it was reported, so they're real issues that would impact users.

Across the same teams, reported bugs fell 20% between Q1 2025 and Q3 2026. The count bounced around in between, but it never followed deploys upward.

Reported bugs didn’t rise with deploys
Reported bugs per quarter, indexed to Q1 2025 = 100. Down 20% by Q3 2026.
Source: QA Wolf platform data, same group of customers, Q1 2025 – Q3 2026.
Reported bugs didn’t rise with deploys
QuarterValueChange
Q1 2025Index 100 (Q1 2025 = 100)Starting point
Q2 2025Index 94 (Q1 2025 = 100)−6% vs Q1 2025
Q3 2025Index 107 (Q1 2025 = 100)+7% vs Q1 2025
Q4 2025Index 90 (Q1 2025 = 100)−10% vs Q1 2025
Q1 2026Index 81 (Q1 2025 = 100)−19% vs Q1 2025
Q2 2026Index 86 (Q1 2025 = 100)−14% vs Q1 2025
Q3 2026Index 80 (Q1 2025 = 100)−20% vs Q1 2025

So each deploy had fewer bugs

Put the two together and the rate of bugs reported per deploy has fallen. In Q1 2025, we found about 14 bugs for every 100 deploys. By Q3 2026, it was about 9: a drop of more than a third.

Bugs per 100 deploys fell by more than a third
Reported bugs per 100 deploys tested. Q1 2025 → Q3 2026: 14.4 → 8.9.
Source: QA Wolf platform data, same group of customers, Q1 2025 – Q3 2026.
Bugs per 100 deploys fell by more than a third
QuarterValueChange
Q1 202514.4 bugs per 100 deploysStarting point
Q2 202511.6 bugs per 100 deploys−19% vs Q1 2025
Q3 202512.5 bugs per 100 deploys−13% vs Q1 2025
Q4 202511.5 bugs per 100 deploys−20% vs Q1 2025
Q1 202610.2 bugs per 100 deploys−29% vs Q1 2025
Q2 20269.5 bugs per 100 deploys−34% vs Q1 2025
Q3 20268.9 bugs per 100 deploys−38% vs Q1 2025

Why might that be?

We can't pinpoint the cause from this data. A few explanations seem likely, and more than one may be at work:

  • The AI question. Whether AI-written code is better or worse than human-written code is an open question. More on this in a future post.
  • Smaller changes. More deploys for the same amount of work means less change in each one, and fewer places for a bug to hide.
  • More testing earlier. Teams shipping faster are also adding checks earlier in the pipeline, like unit tests and AI code review, so more issues get fixed before a deploy reaches our end-to-end tests.

What we can conclude, and what we can’t

The usual assumption is that speed and quality trade off. Ship more and more things break.

For these teams, over these 21 months, that trade-off didn't show up. They deployed about 30% more often, while bugs reported per deploy fell by more than a third.

Every team in this data has comprehensive end-to-end coverage through QA Wolf. That testing is part of the lesson, not just a caveat. As AI helps teams write more code, these results show that faster releases and fewer bugs per deploy can go together when teams pair shipping faster with testing well.

We can't say that AI or testing caused the improvement, or what happened at teams without this coverage. And these are mostly bugs caught before release, not production incidents. But more speed doesn't have to mean more bugs.

Maybe there is a free lunch after all, at least for teams that test well.

How we measured

Data. QA Wolf platform data, Jan 1, 2025 – Sep 30, 2026. Comparisons run Q1 2025 → Q3 2026.

Which teams. The same group of customers every quarter, all with QA Wolf since before October 2024. We counted only the environments each customer tested every quarter, and left out customers who overhauled their release or testing setup.

Deploys tested. Distinct deployments to a customer's staging, QA or production environment that automatically triggered a QA Wolf test run through a deploy trigger. Not counted: scheduled runs, runs on temporary preview environments for branches and pull requests, and runs started by hand.

Bugs. Bugs a QA Wolf QA engineer investigated, verified and reported from a test run, excluding any later canceled, dated by when they were reported.

Robustness. The charts use totals, so a few large customers carry a lot of weight. When every customer counts equally, the results are similar.

Limits. No comparison group of teams without this testing. Only customers who stayed with us the whole period. Bugs found mostly in test environments, not production incidents. And correlation only.

‍

Want to see how this works on your team? QA Wolf builds, runs and maintains end-to-end tests that trigger on your deploys, and a QA engineer verifies every bug before it reaches you. Book a demo.

Try the AI testing platform that makes QA 12x faster.