Engineering performance
DORA, SPACE, METR, rework, and verification cost. These essays cover how AI-assisted engineering is measured, why old metrics mislead, and what to track once agents do the work. Browse all insights.
State of Startups 2026: 61% of Startup Codebases Are Majority AI-Written
Supabase's State of Startups 2026 finds 61% of startups have more than half their codebase written by AI, and 40% sit at 76-100%. Review budgets do not scale with that.
DORA Metrics Are Necessary But Insufficient For Agentic Development
DORA measures delivery-system behavior and still matters, but it cannot see whether an autonomous engineering system stays architecturally coherent as autonomy rises. Governance metrics are the missing third layer.
Atlassian's Agentic Pivot: 88% Need a System of Record, 19% Have One
Atlassian's 2026 survey of 1,100+ engineers finds AI compressed implementation while review and governance lag. Context helps agents; decisions have to bind them.
Brooks's Law After AI: When More Coding Agents Make Software Harder to Control
Agents remove ramp-up time. They do not remove coordination cost. What Brooks's Law means when the next contributor is an AI agent.
How to Close the Agentic Adoption Gap: What “Enforce” Means for Software Engineering
McKinsey's adoption playbook ends with Enforce. For software engineering, accepted architectural decisions also need to reach and constrain agent execution.
You Can’t Govern What You Can’t See: The Visibility Gap in Agentic Engineering
Claude Code pricing and Copilot billing reveal a deeper problem: leaders have little visibility into an emerging agentic workforce. Costs are a symptom; governance starts with visibility.
The Strongest AI Engineering Teams Won’t Be Built From Bigger Agents
A DevOps study found collaborative team structures beat both silos and full integration. Applied to AI agents, the missing variable is a governance layer — with the metrics to prove it.
Stanford AI Index 2026: AI Coding Is Becoming Solved. Engineering Governance Has Not.
The Stanford AI Index 2026 shows coding benchmarks near saturation and adoption mainstream. If generation is solved, the unsolved layer is keeping agent output aligned with architectural decisions.
The SPACE Framework: Measuring GitHub Copilot’s Real Productivity Impact
Most Copilot ROI reports stop at accepted suggestions. The SPACE Framework — from GitHub, Microsoft Research, and University of Victoria — reveals the governance gap hiding beneath the activity gains.
METR's AI Productivity Studies: Why AI Coding Feels Fast but Measures Slow
Two METR studies say almost the opposite thing about AI's impact on developer productivity. Reconciling them shows what data teams should actually measure — and where governance pays back.
Why AI Coding Productivity Gains Often Lead to More Rework
Faros AI's 2026 telemetry shows throughput up 33.7% — and bugs per developer up 54%, code churn up 861%. Much of that churn is rework correcting AI changes that violated system-level decisions.
The Verification Tax of AI Coding Agents: Why Faster Code Creates More Review Work
Glean finds workers reclaim 11 hours a week from AI and hand 6.4 back in botsitting. For engineering teams the bill lands as a verification tax — and the refund is executable architectural constraints, not more review.
GitHub Just Turned AI Code Review Into a Budget Line Item
GitHub Copilot moved to usage-based AI Credits on June 1, 2026, and code review now burns Actions minutes per PR. When review is metered, ungoverned output becomes a finance problem — governance economics.
AI Engineering ROI: How to Measure the Real Return From AI Coding
AI spend is up 28x. Feature delivery is flat. Where the value leaks between generation and delivered software, and how to measure it in three layers instead of one.
AI Throughput Is Not Engineering Throughput
NBER research tracking 500,000+ GitHub developers found autonomous agents produce a 240% increase in coding activity, attenuating to 30% at the release stage. Coding activity was never the right unit.
The AI ROI Problem Is Not About Models. It Is About Systems.
Weak enterprise AI ROI is not evidence that AI fails to create value. It is evidence that organizations matured generation capability faster than the governance and verification infrastructure needed to operationalize it. Generation is commoditizing. Verification is not.
Deployment Quality Will Define the AI Era
The first AI era rewarded early adoption. The next rewards operational quality. KPMG research points to deployment quality as the new differentiator — and for engineering teams, that starts with governance.
AI Adoption Maturity Model: A Technical Analysis for Engineering Leaders
CMU SEI and Accenture's new maturity model defines five levels and eight dimensions, with Risk and Governance among them. The question the model leaves open is how governance decisions actually reach coding agents.
Datadog’s State of AI Engineering Report Quietly Confirms the Governance Crisis
1,000+ production orgs. 70% running 3+ models. One sentence buried in the data: “In practice, model churn becomes a governance problem.” Here is what the report actually says about where the industry is headed.