Technical Debt

You cannot buy a technical debt dashboard

Google tested 117 candidate metrics against what its engineers actually reported as debt. None of them worked. Here is what does — and why the problem is getting worse right now.

Google tested 117 candidate objective metrics against what its own engineers reported as technical debt. None of them predicted it. The models explained less than 1% of the variance.

That result deserves more attention than it gets, because an entire category of tooling is sold on the opposite premise. If you have been shopping for a technical debt dashboard that scores your estate in dollars, the best available research says the thing you are buying does not measure the thing you want.

Why the metrics fail

Debt is the gap between the current state and the ideal state for your context. Two codebases with identical static-analysis profiles can carry completely different debt loads depending on what the business needs them to do next. Complexity you never touch is not debt. It's just complexity.

Google's operational answer is almost aggressively unglamorous: survey about a third of engineers quarterly, asking which categories of debt are actively slowing them down. Over five years this produced the largest positive shift of any productivity factor they measure. It costs nearly nothing to run.

Meanwhile the debt is accumulating faster

Three independent-ish datasets point the same direction on AI-assisted development, and the direction is not good.

32.7% vs 84.5%

Merge rates for AI-generated pull requests versus unassisted ones, across 8.1 million PRs from 4,800 teams (LinearB). AI-assisted PRs are also about 2.5× larger and wait more than five times as long for review pickup. The vendor sells engineering metrics — but merge rate is a hard mechanical number, difficult to manufacture.

Alongside that: GitClear's analysis of 623 million code changes finds refactoring collapsed from around 21% of changed lines in 2022 to 3.8%, with code duplication up 81% since 2023. And Veracode's controlled benchmark across 100+ models found 45% of AI-generated code samples failed security tests — with newer, larger models no better. Functional correctness improved. Security performance stayed flat.

Both are vendors with an interest in the conclusion, and I'd hold the exact magnitudes loosely. But the direction is corroborated by DORA's independent finding that AI adoption raises throughput while degrading delivery stability. More code, less refactoring, larger reviews, more security findings. That is a debt accumulation curve.

The governance asymmetry

Cloud spend became governable because FinOps gave it a unit of account. Technical debt has no equivalent and, per Google's research, may never have a fully automated one. So debt stays invisible on the P&L at exactly the moment AI is accelerating its accumulation.

McKinsey's framing is the one that works in a boardroom: CIOs estimate debt at 20–40% of the value of their entire technology estate, companies pay an additional 10–20% on project costs to work around it, and organizations in the worst quintile are 40% more likely to have cancelled or incomplete modernizations. Those are self-reported estimates, not measurements — say so when you use them. But the last one lands, because a cancelled modernization is a visible, expensive failure that a board has usually already lived through.

What to do about it

  1. Run the survey. Quarterly, a third of your engineers, one question: which categories of debt are slowing you down right now? This is the highest-return thing on the list and it is free.
  2. Track three mechanical metrics. PR merge rate, duplication rate, and refactoring as a share of changed lines. All three are cheap to instrument and all three are currently moving the wrong way industry-wide.
  3. Put debt on the board agenda as modernization risk, not as code quality. "40% more likely to have a cancelled modernization" is a sentence a board acts on. "High cyclomatic complexity" is not.
  4. Don't buy the dashboard. Spend the money on the review capacity that AI-assisted development is currently overwhelming.

Sources

  1. Jaspan & Green (Google), summarized with citations at How Google measures and manages tech debt
  2. LinearB, 8 million PRs (4,800 teams, 42 countries) — vendor-published
  3. GitClear, The Maintainability Gap (623M code changes) — vendor-published
  4. Veracode, 2025 GenAI Code Security Report — vendor-published controlled benchmark
  5. DORA, 2025 State of DevOps Report
  6. McKinsey, Breaking technical debt’s vicious cycle

Kurt Wysock

Fractional & interim CTO and board technology advisor. Currently interim CTO at a luxury destination-club operator, head of the Technology & Architecture Office at a global consulting firm, and founder of JSummit Consulting. 30+ years in enterprise architecture across energy, hospitality, manufacturing, financial services, retail, and healthcare. TOGAF® 9 Certified. More about Kurt · Get in touch

← Simplification funds AI

Want this applied to your systems?

The Production AI Blueprint is a fixed-fee engagement that sorts out your data, governance, controls, and ownership — and hands you a sequenced roadmap you own outright.