The metric nobody reports credibly
Every engineering org is now shipping AI-generated code, and almost none of them can say whether it is better, worse, or riskier than what their humans write. The numbers that get quoted - suggestion acceptance rate, lines generated, tool seat adoption - describe usage, not quality. They inflate easily and collapse under any real audit.
Measuring AI-generated code credibly means tagging provenance at the source and then following that code through its entire life: how much survives, how often it is reverted, how much review it needed, and whether it caused failures downstream. That is the difference between a slide and a signal. Our measurement product is built to produce the second kind.
Once provenance is recorded, AI code stops being a black box. You can finally answer the question your CTO is actually asking: is this making us faster without quietly making us more fragile? Compare the two populations directly with PR quality analytics.