Research & technology

The Cost of AI Slop in Software Development Report

QA Wolf
September 9, 2026

Developers have long suspected AI slop was expensive. But nobody had put a number to it. We did — and that number might surprise you: they start around $1.17 TRILLION.

And that's only a partial accounting of the total costs. 

This piece adds up the true cost of AI-generated code that skipped proper verification, accumulated in production, and is now quietly compounding.

Wondering how we came to it? Read on. 

Turning AI slop into a dollar figure

We thought this project would take an afternoon. Instead, it took weeks. Not because the data about the impact of AI slop doesn't exist but because too much of it exists but it’s all fragmented and partial. 

Over the last couple of years, study after study has measured various small parts of the problem. There's research on how much AI-generated code gets thrown away almost as soon as it's written. There's also research on how often that code ships with security holes in it. And there's giant macro estimates for what bad software costs the economy as a whole, irrespective of whether it’s written by AI.

But none of these studies came with dollar figures attached. What follows is our attempt to quantify the costs in any way that we could. You might disagree with the way we calculated one figure or another but we’ve done what we could to be as rigorous as possible. We also showed our work so you can plug in your own values or estimates. 

In the end, we decided to look at the cost of AI slop across 6 different categories: 

  • Code generation
  • Verification of AI-generated code
  • Change failure rates/fixes in production
  • Incidents
  • AI slop debt
  • Lost revenue from customer churn

Money lost in code generation

EntelligenceAI recently released a study that looked at their customers’ engineering data and found that for every $1 of AI tokens: 

  • $0.44 was going to fixing bugs in generated content.
  • $0.27 was spent on rework due to bad generated code. 

That gave us a total of $0.71 per dollar of AI tokens or 71% of all direct AI spend that can be directly counted towards the costs of AI slop.

How much a company is actually spending on AI slop per engineer, however, can vary widely depending on what tool they’re using. For AI coding tools like GitHub Copilot that are sold via flat rate subscriptions, the cost of AI slop per engineer on their Pro+ plan that costs $39 per month would be 71% of that or $27.69 per engineer per month and $332.28 per engineer per year. 

On the other end of the spectrum are the AI slop costs of consumption based code generation tools. For example, take Anthropic’s average reported Claude Code costs per developer of $6 to $12 a day or $180 to $360 per month. Earlier this year, they listed that amount in their docs but have since removed it, likely because the average is now higher. But, as that’s the most recent data we have, we decided to use that for our calculations.

At that rate, AI slop makes up $127.80 to $255.60 per month in AI costs or $1,533.60 to $3,067.20 per year per developer. With many companies reporting power users racking up as much as $500 to $2,000+ per month, that would cost $355 to $1,420 per month or $4,260 to $17,040 per year. 

But just calculating the per engineer cost doesn’t provide a sense of what the total costs of AI slop in generation could be. According to Mordor Intelligence, global spending on AI code generation tools is expected to reach $9.35 billion annually by the end of 2026. Following the above ratio, you could say that 71% of that or $6.64 billion per year is wasted on AI slop. 

Generation costs:  $6.64 billion globally or $332.28 - $17,040+ per engineer annually.

Additional time it takes to validate and ship AI code

Generating code has never been faster. Verifying it has never been slower.

Reviewing AI-generated code takes more effort than reviewing human written code. Developers have to trace unfamiliar logic, hunt subtle bugs, and untangle large blocks of output nobody remembers writing. 

In Harness's 2026 State of Engineering Excellence Report, most developers said that since their teams adopted AI they now spend 30% more time reviewing code — and 28% reported their review time climbing by more than 30%.

Engineers were already spending up to five hours a week reviewing code on average. A 30% increase on five hours adds about an hour and a half a week, pushing review to roughly six and a half hours. 

Testing is likely the same story, however try as we might, we couldn't find studies quantifying how much extra time is spent testing and debugging AI code, so we can't quantify that for our analysis. We do know that around 45% of developers say debugging AI-generated code takes longer than fixing human-written code, according to Stack Overflow's 2025 Developer Survey. But since no studies exist on how long it previously took on average to debug code or how much extra time devs are spending, there's no way to objectively measure that so that remains a gap in our analysis.

According to Rippling, the average US software engineer earns about $148,000, or roughly $71 an hour. That extra hour and a half of verification a week costs about $107 per engineer a week, or roughly $426 a month. Over a year, that's about $5,538 per engineer. At a company with 50 engineers, that's roughly $277,000 every year.

Validation costs: ~$5,538 per engineer annually.

Fixes in production

Developers are spending more time fixing AI code after it's shipped but it’s hard to pinpoint exactly how much extra time they’re spending. 

Here’s what we know: 

  • A 2026 study found 60% of software leaders reported quality issues in the past year because code creation outpaced testing capacity. 
  • A study by GitClear found the percentage of new code requiring revision within two weeks grew from 5.5% in 2020 to 7.9% by 2024, highlighting the hidden time engineers spend fixing logically flawed AI implementations. 
  • A 2026 study by Cortex found a 30% increase in change failure rates.

However, none of those studies quantify the number of extra bugs or quality issues those percentages amount to and it’s impossible to definitively pinpoint it from the data provided. 

But we also know that the cost of a defect scales with how late it's caught. The most commonly cited data on this is research by the Systems Sciences Institute at IBM, where in a 1995 study, they found that the cost to fix an error in production was 100x more than fixing one caught during the design stage.

But other research has found much lower increases in time and cost. We decided it was better to be conservative. So, for that reason, we used a 2002 study by the National Institute of Standard Technology (NIST) for our calculations. It measured the cost of fixing bugs in production and found that it took 15 hours compared to five hours if it was found at the coding stage. That represents three times extra effort or 10 hours more per bug. 

Conservatively, we also estimated a company might see three to five extra bugs per month that are found in production due to AI slop, representing 30 to 50 hours of extra work per month spent fixing those bugs. Some companies might see more and some might see less. 

Using the average salary data for a software engineer from Rippling that we shared above, that would translate into $2,312 a month and $27,750 per year at the lower figure of 30 hours per month and $3,854 a month or $46,250 per year for the higher figure of 50 extra hours.

Fix costs: $27,750 - $46,250 per company annually.

Incidents

Over the last several years, there’s been a rise of AI code related incidents and that upwards trajectory is unlikely to change soon. In March 2026, Amazon held a mandatory all-hands after internal documents surfaced describing a "trend of incidents" with a "high blast radius" linked to "Gen-AI assisted changes." Amazon is far from alone. 

Here’s what we know: 

  • According to the ThousandEyes blog, global outages climbed from 1,382 in January 2025 to 1,595 in February, then spiked to 2,110 in March.
  • In a 2026 report, 100% of the technology leaders surveyed claimed their company had experienced AI-related downtime

So, what’s the cost of an incident? It varies depending on the size of the company and the nature of an incident (i.e. complete downtime vs a slowed application) but research shows it costs an average of $15,000 per minute or $900,000 per hour, according to Splunk and Cisco's Hidden Costs of Downtime 2026 report. 

It’s hard, however, to determine how much extra downtime companies are experiencing directly due to AI. So, we decided to be conservative and predict just one to five AI-related incidents a year at an incident length of 175 minutes, the average found in research by PagerDuty, at a cost of up to $15,000 per minute. 

If teams had just one extra incident per year, that could cost them up to $2,635,000 depending on the size of their company and the severity of the incident. If they had five extra incidents, that could cost the company up to $13,125,000 depending on the same factors.

Larger companies or companies that had more incidents due to AI code, could see costs significantly higher than that. 

Incident costs: Up to $2,635,000 - $13,125,000+ on average per company annually.

AI slop debt

AI slop debt is the new tech debt.

And technical debt is expensive. CISQ’s 2022 report pegged the cost of accumulated software debt at $1.52 trillion for U.S. codebases. That’s because technical debt forces teams to spend a large share of their time on maintenance and rework instead of new product features. 

Developers are already feeling the pain with a 2026 study by Sonar Source finding that 40% of developers believe AI has increased debt by generating unnecessary or duplicative code. 

Here’s what we know about what’s driving AI slop debt: 

  • Recent academic research comparing human-written code to AI-generated code found that large language models produce code with significantly lower lexical diversity but with much higher structural repetition. This repetitive, pattern-based uniformity severely increases long-term maintainability issues.
  • Gitclear found that refactored code fell from about 21% of changed lines in 2022 to 3.8% in 2026. What’s more, the number of codeblocks with five or more duplicated lines increased 8x. That’s because assistants make it easy to tab in a new block and are unlikely to suggest reusing an existing function, partly because of limited context size. But that has huge implications on maintenance. 
  • Security debt is also set to become a big issue. Veracode tested 100+ LLMs across 80 tasks and found AI introduced security vulnerabilities in 45% of cases, choosing the insecure option nearly half the time when given the choice.

The problem with all those stats is that they only tell us there’s a problem but don't quantify how big of a problem it is. One academic study that is often cited as finding increases in tech debt due to AI coding agents found more specifically that AI coding agents increased static analysis warnings by 30% and code complexity by 41%. While both arguably are potential technical debt, it doesn’t represent all of the types of technical debt that could be created so we don’t know what percentage of the total technical debt those figures represent. 

In order to get to some kind of figure, we decided to take those figures as representative of technical debt as a whole and multiply them by the 2022 level of technical debt from the CISQ report, that would mean AI increased the costs of U.S. technical debt by over $456 billion and then again by $623 billion, adding up to a total of $1.07 trillion if one were to measure them cumulatively. 

However, that’s a very big number and would make up the majority of our costs so we decided to be more conservative and assume technical debt increased by between $456 billion and $623 billion. 

AI slop debt costs: $456 billion to $623 billion (or potentially as high as $1.07 trillion) in the U.S. alone.

Customer churn due to poor quality applications

It wouldn’t be a surprise if companies were facing customer churn because of incidents and defects due to AI-written code but no studies exist that quantify how much churn can specifically be attributed back to AI slop. 

The reason why is likely because it would be particularly hard to measure in a study. You would have to know what percentage of customers churned due to application issues, the particular application issue that led to them churning, and whether those application issues were directly related to AI code. A customer survey could get you one half of that data but a private company is unlikely to divulge the contents of their postmortems. 

But, according to TopTal, 90% of app users reported they stop using an app due to poor performance, while 88% of online consumers are less likely to return to a site after a bad experience. Industry groups like CISQ have also pegged the cost of losses due to poor software quality and operational failures at over $2.41 trillion annually in the U.S in 2022. While neither figure directly addresses issues from AI-generated code, the poor quality of the output undoubtedly would impact customer churn the same way non-AI generated software quality does. 

One way to measure it would be to take the 2022 figure of $2.41 trillion in annual losses due to poor software and calculate what the proportional increase in software defects since then would translate to in additional costs. 

For this, we decided to use the 30% increase that Cortex found in change failure rates in their 2026 study. This is a conservative choice since change failures might have gone up even more than 30% between 2022 when the CISQ study was released and the year of the Coxtex study but Cortex hadn’t previously measured this in other reports so we don’t know. 

A 30% increase in change failure rates assuming that those failures caused issues for customers in production could translate into $723 billion in additional losses due to poor software quality. And that’s in the U.S. alone. 

Customer churn cost: $723 billion in the U.S. alone

Adding up the costs of AI slop

We couldn't collapse all of these figures into one number. A per-developer cost and a national or global economy cost are different units and so we’re presenting them that way. 

Per developer, per year: $5,870 to $22,578+ 

Broken down, here’s what AI slop costs for a single engineer:

  • Wasted code generation: $332.28 to $17,040+. 
  • Extra verification time: Extra verification time: ~$5,538. 

Per company, per year: $2.66 million to $13.17 million+

Here’s what lands on a single organization once slop reaches production:

  • Fixing production defects: $27,750 to $46,250. 
  • AI-related incidents: $2,635,000 to $13,125,000+.

Incidents are where the bill starts to matter. A single extra AI-driven outage can cost more than an entire engineering team's annual salaries. They sit on top of the per-developer costs, which every organization pays per head and the economy wide costs that we couldn’t break down to a company basis but which are paid by companies directly. 

To get a better idea of your exposure, take the per-developer range above — $5,870 to $22,578+ a year in wasted generation and verification time — and multiply it by your headcount. At 50 engineers, that's roughly $940,000 to $2.7 million a year before a single incident or production bug enters the picture. At 200, it's $3.8 million to $10.8 million.

The macro figure, per year: $1.17 trillion

This is the number with all the digits:

  • AI slop debt: Starting at $456 billion (but as high as $1.07 trillion).
  • Lost revenue from churn: $723 billion.
  • Wasted spend on AI code-generation tools: $6.64 billion.

But here's the crucial caveat: this is almost entirely a U.S. figure. Both the costs of slop debt and churn are derived from figures that measured only the costs for U.S. codebases. Only the $6.64 billion in wasted tool spend, is a global figure. Which means the real worldwide total isn't $1.8 trillion — it's some large multiple of it that nobody can cleanly calculate, because the national-scale studies that we used to calculate this simply don't exist outside the U.S.

And even this understates it. The macro figure captures some of the per-developer and the per-company costs but not all of them. 

So, $1.17 trillion is a floor, not a ceiling — the measurable slice of a much larger, mostly dark total.

Conclusion

Code generation got faster. Verification didn't. A problem caught early is cheap. The same problem caught in production is expensive. AI's failure mode is pushing more code into production faster, without the verification step that would have caught the problem while it was still cheap. 

The part teams keep getting wrong is they treat verification as a tax on velocity, the thing slowing them down. It's the opposite. Verification is what makes velocity safe. Skip it and you don't go faster — you just move the bill to production, where it's marked up 100x and paid in incidents, churn, and engineering hours nobody planned for.

The good news is this is a solved problem. Our agentic testing platform was built with learnings from 100 million test runs for companies like DoorDash, Cursor, Lovable, Drata, and more. It builds and maintains end-to-end test coverage that runs against every change, catches the regressions AI slop introduces, and flags them before they ship. Automated, parallelized, and fast enough to keep pace with AI-generated pull requests instead of becoming the bottleneck behind them. Our purpose-built QA platform handles the test creation and the maintenance so you can ship fast AND safe. 

The math in this piece mostly represents the cost of shipping unverified code. QA Wolf is how you stop paying it.

Try our agentic QA testing platform for free today.

Try the AI testing platform that makes QA 12x faster.