AI is helping teams write more code, and the assumption is that more bugs will follow. But that isn't what we're seeing.
We followed the same group of QA Wolf customers from the start of 2025 through September 2026. They deployed about 30% more often, while the number of bugs we found fell. Bugs reported per deploy dropped by more than a third.
These teams have comprehensive end-to-end tests that check whether their applications work the way users expect. Each deploy that triggers those tests gives us a record of what shipped and which bugs our QA engineers verified and reported, mostly before release.
The data doesn't tell us whether AI caused the increase in release velocity. But it does show that, as AI becomes part of how software gets built, teams with robust testing can deploy more often without a corresponding rise in bugs. Here's what we found, what might explain it, and how we measured it.
AI is writing more of the code
The way software gets built has changed more in the last two years than in the two decades before. The shift is well documented. Google's 2025 DORA report, a survey of nearly 5,000 technology professionals, found that 90% now use AI at work and linked AI adoption to higher software delivery throughput (DORA).
We see it in our own data too. About a third of commits in August and September 2026 carried an explicit AI tool credit, such as "Co-authored-by: Claude" (more details on this in a future post).
Teams are deploying more often
We measured release frequency by counting deploys that triggered our test suites. For most teams that's a deploy to a test environment like staging, usually after code is merged.
Between Q1 2025 and Q3 2026, the teams in our group went from around 6,400 deploys tested per quarter to over 8,000, up 31%, with half of that increase coming after Q1 2026.
As you can see, the range is wide and some teams are even shipping less frequently than in Q1 2025 but nearly a third of teams more than doubled their release velocity.
Did more deploys mean more bugs? No.
A quick note on what counts as a bug here, because it matters.
We build and maintain automated end-to-end tests for our customers, so they catch the kind of problems users would actually hit. When a test fails, one of our QA engineers investigates. Either the test needs updating, or it's a real bug, in which case they reproduce it and file a report.
So these aren't raw test failures or flaky tests. Every bug here was verified by a person before it was reported, so they're real issues that would impact users.
Across the same teams, reported bugs fell 20% between Q1 2025 and Q3 2026. The count bounced around in between, but it never followed deploys upward.
So each deploy had fewer bugs
Put the two together and the rate of bugs reported per deploy has fallen. In Q1 2025, we found about 14 bugs for every 100 deploys. By Q3 2026, it was about 9: a drop of more than a third.
Why might that be?
We can't pinpoint the cause from this data. A few explanations seem likely, and more than one may be at work:
- The AI question. Whether AI-written code is better or worse than human-written code is an open question. More on this in a future post.
- Smaller changes. More deploys for the same amount of work means less change in each one, and fewer places for a bug to hide.
- More testing earlier. Teams shipping faster are also adding checks earlier in the pipeline, like unit tests and AI code review, so more issues get fixed before a deploy reaches our end-to-end tests.
What we can conclude, and what we can’t
The usual assumption is that speed and quality trade off. Ship more and more things break.
For these teams, over these 21 months, that trade-off didn't show up. They deployed about 30% more often, while bugs reported per deploy fell by more than a third.
Every team in this data has comprehensive end-to-end coverage through QA Wolf. That testing is part of the lesson, not just a caveat. As AI helps teams write more code, these results show that faster releases and fewer bugs per deploy can go together when teams pair shipping faster with testing well.
We can't say that AI or testing caused the improvement, or what happened at teams without this coverage. And these are mostly bugs caught before release, not production incidents. But more speed doesn't have to mean more bugs.
Maybe there is a free lunch after all, at least for teams that test well.
How we measured
Data. QA Wolf platform data, Jan 1, 2025 – Sep 30, 2026. Comparisons run Q1 2025 → Q3 2026.
Which teams. The same group of customers every quarter, all with QA Wolf since before October 2024. We counted only the environments each customer tested every quarter, and left out customers who overhauled their release or testing setup.
Deploys tested. Distinct deployments to a customer's staging, QA or production environment that automatically triggered a QA Wolf test run through a deploy trigger. Not counted: scheduled runs, runs on temporary preview environments for branches and pull requests, and runs started by hand.
Bugs. Bugs a QA Wolf QA engineer investigated, verified and reported from a test run, excluding any later canceled, dated by when they were reported.
Robustness. The charts use totals, so a few large customers carry a lot of weight. When every customer counts equally, the results are similar.
Limits. No comparison group of teams without this testing. Only customers who stayed with us the whole period. Bugs found mostly in test environments, not production incidents. And correlation only.
Want to see how this works on your team? QA Wolf builds, runs and maintains end-to-end tests that trigger on your deploys, and a QA engineer verifies every bug before it reaches you. Book a demo.