Technical Debt
Google tested 117 candidate metrics against what its engineers actually reported as debt. None of them worked. Here is what does — and why the problem is getting worse right now.
Google tested 117 candidate objective metrics against what its own engineers reported as technical debt. None of them predicted it. The models explained less than 1% of the variance.
That result deserves more attention than it gets, because an entire category of tooling is sold on the opposite premise. If you have been shopping for a technical debt dashboard that scores your estate in dollars, the best available research says the thing you are buying does not measure the thing you want.
Debt is the gap between the current state and the ideal state for your context. Two codebases with identical static-analysis profiles can carry completely different debt loads depending on what the business needs them to do next. Complexity you never touch is not debt. It's just complexity.
Google's operational answer is almost aggressively unglamorous: survey about a third of engineers quarterly, asking which categories of debt are actively slowing them down. Over five years this produced the largest positive shift of any productivity factor they measure. It costs nearly nothing to run.
Three independent-ish datasets point the same direction on AI-assisted development, and the direction is not good.
Merge rates for AI-generated pull requests versus unassisted ones, across 8.1 million PRs from 4,800 teams (LinearB). AI-assisted PRs are also about 2.5× larger and wait more than five times as long for review pickup. The vendor sells engineering metrics — but merge rate is a hard mechanical number, difficult to manufacture.
Alongside that: GitClear's analysis of 623 million code changes finds refactoring collapsed from around 21% of changed lines in 2022 to 3.8%, with code duplication up 81% since 2023. And Veracode's controlled benchmark across 100+ models found 45% of AI-generated code samples failed security tests — with newer, larger models no better. Functional correctness improved. Security performance stayed flat.
Both are vendors with an interest in the conclusion, and I'd hold the exact magnitudes loosely. But the direction is corroborated by DORA's independent finding that AI adoption raises throughput while degrading delivery stability. More code, less refactoring, larger reviews, more security findings. That is a debt accumulation curve.
Cloud spend became governable because FinOps gave it a unit of account. Technical debt has no equivalent and, per Google's research, may never have a fully automated one. So debt stays invisible on the P&L at exactly the moment AI is accelerating its accumulation.
McKinsey's framing is the one that works in a boardroom: CIOs estimate debt at 20–40% of the value of their entire technology estate, companies pay an additional 10–20% on project costs to work around it, and organizations in the worst quintile are 40% more likely to have cancelled or incomplete modernizations. Those are self-reported estimates, not measurements — say so when you use them. But the last one lands, because a cancelled modernization is a visible, expensive failure that a board has usually already lived through.
The Production AI Blueprint is a fixed-fee engagement that sorts out your data, governance, controls, and ownership — and hands you a sequenced roadmap you own outright.