The Agentic Maturity Model
Adoption is not maturity.
Maturity is where trust lives.
Every enterprise AI maturity model on the market measures adoption: seats, tokens, enthusiasm. A company can be a top-decile AI spender and a bottom-decile AI organization in the same fiscal year — and the adoption dashboard will not tell you which one you are looking at. What distinguishes maturity is not how much work AI produces, but what mechanism establishes that the work is correct, and whether that mechanism scales.
Audit-checkable, not sentiment-checkable
Each level is defined by a structural fact you could verify in an afternoon, not a survey answer.
Transitions are driven by pain, not aspiration
Each level ends because its trust mechanism hits a wall. Identify your pain, and you know your level.
Maturity is not monotone in AI usage
A Level 2 organization may consume more tokens than a Level 3 organization. Mature organizations spend deliberately.
The spine
Five levels of trust location
From prohibition to a compounding, self-improving system. Each level is named by where trust lives, ends at a wall, and carries an audit check you could run this afternoon.
Trust mechanism
Human review of all AI output
Where the modal enterprise is — because Level 2 feels like maturity. Licenses procured, policy written, copilots deployed, usage dashboards green. But the trust mechanism is unchanged from the pre-AI era: a human reads everything. AI has scaled production without scaling verification. Review queues balloon, rubber-stamping quietly becomes the norm, and defect escape rates rise while every adoption metric improves. This is the plateau. Most “AI transformation” programs park here permanently and call it success.
The question the organization is asking
“How do we review all of this?”
The wall that ends it
The review bottleneck
The centerpiece
The Diagonal Law
Your AI program has two numbers, and you are probably tracking one of them. Capability is what you have deployed. Verification is what establishes that its output is correct. The relationship between the two is the whole game.
“Capability above verification is risk; verification above capability is waste.”The Diagonal Law
None
Prompt
Skill
Agent
Orchestration
Hover or tap a cell. The diagonal is the five maturity levels; everything else is a pathology with a name.
The misalignment inventory
Nearly every enterprise AI failure is an off-diagonal state with a name. Four pathologies live on the grid; three live on the knowledge and observability tracks that run beneath it. Locate your organization in one of these cells in under a minute.
Shadow Fleet
C1–2 · V0
Policy says no; egress logs say yes. Usage exists — it is just invisible. Risk without governance, and no telemetry to even measure it.
The Review Bottleneck
C3 · V2
Agents produce task-scale output; humans still read every diff — or pretend to. The defining pathology of the era, and where most “successful” AI programs are parked.
Cowboy Autonomy
C3–4 · V1
Agents merging with neither human nor machine gates. Fast until it isn’t. Unpriced risk accumulating off the books, discovered by the incident rather than the dashboard.
The RAG Plateau
Knowledge track stuck at stage 2
“We implemented RAG” presented as an AI strategy. Runtime retrieval hoarding, with no compilation of stable knowledge into skills.
Dashboard Theater
Observability track stuck at stage 2
Seats, tokens, and acceptance rates reported as outcomes. No defect-escape or gate-calibration data exists. Acceptance rate is a sentiment metric in a lab coat.
Compliance Freeze
C1 · V3–4
Gates so heavy nothing ships. Trust infrastructure with nothing to trust — the above-diagonal failure: waste as a residence, not a staircase step.
Skill Rot
Knowledge stage 3 without observability stage 4
A skill library authored once in a burst of enthusiasm, never mined, never versioned — quietly decaying into misinformation.
The full argument — the review-bottleneck arithmetic, the truce between the CTO and the CISO, why you cannot leapfrog — is in the essay.
Beneath the grid
The five tracks
The spine says where trust lives. The tracks say what that implies for capability, verification, knowledge, observability, and economics — each a five-stage ladder aligned to the levels.
The unit of AI work
- C0None — shadow prompts only.
- C1Prompt — personal, ephemeral. Value dies with the author.
- C2Skill — the first organizational unit: versioned, shared, reviewable. The prompt is to the skill what the shell one-liner is to the committed script.
- C3Agent — autonomy at task scale. Only safe once verification has reached rails + adversarial review.
- C4Orchestration — multi-agent control flow tuned by cost, wall-clock, and accuracy — maturing into the self-improving lifecycle.
The capability ladder is the visible face of maturity, which is why every vendor sells it as the whole story.
The progression
The only safe path
Each boundary has one keystone unlock — the specific investment that dissolves the current wall. Capability rungs can be bought; trust relocation cannot. That asymmetry is why the risk triangle is crowded and the diagonal is not.
2 → 3 · The structural break
Adversarial review and frozen rails migrate trust from human attention to machine gates — plus the outcome telemetry to make the review bottleneck visible at all. Adversarial review is not a Level 4 luxury; it is the entry requirement for Level 3.
3 → 4 · The compounding turn
Skill mining and gate calibration start the distillation loop: recurring findings become deterministic controls, observed work becomes versioned capability, and cost per merged, verified change starts falling.
Terms of art
- Frozen rails
- Executable tests written and locked before implementation begins. The builder cannot argue with them; neither can the agent.
- Adversarial review
- Fresh reviewer contexts chartered to refute the work, not assess it, running until findings converge. The ADLC calls this prosecution.
- Planted defects
- Known bugs run through your gates to measure what they actually catch. The difference between “we have review” and a catch rate.
- Skill mining
- Converting observed, recurring work into versioned, shared, reviewable skills — capability that survives departures.
- Distillation
- Converting recurring findings into permanent deterministic controls. The system gets cheaper and stricter at once.
You cannot verify what you cannot observe, and you cannot distill what you do not record — every level transition is an observability upgrade before it is a tooling upgrade.
The reference implementation
The AMM tells you where you are. The ADLC shows what Level 4 looks like in practice.
The verification stages are defined by outcomes, not by any one methodology — any process that demonstrably achieves them qualifies. For software development, the Agentic Development Lifecycle is the reference implementation: written for the practitioner who will build Level 4, where the AMM is written for the leader who must get the organization there.
Visit agenticlifecycle.ai →The series
Eight essays, one argument
The model is published as a series at voodootikigod.com. Each essay takes one track or one transition and makes the case in full.
- 01
The Adoption Curve Is Not a Maturity Model
Trust location as the real axis; why every existing model gets Goodharted.
- 02
The Five Levels
Walls, unlocks, audit checks — your copilots made you feel mature; the review bottleneck says otherwise.
- 03
The Diagonal Law
The capability × verification grid and the misalignment inventory. Locate yourself in a minute.
- 04
RAG Is Runtime Knowledge, Skills Are Compiled Knowledge
The knowledge ladder; the RAG plateau; skill mining as the compiler.
- 05
You Cannot Distill What You Do Not Record
Observability as the transition-gating track; dashboard theater vs. outcome and learning telemetry.
- 06
From Review to Prosecution to Calibration
The verification ladder for an enterprise audience; adversarial review as the 2→3 keystone.
- 07
The Economics of Level 4
Cost per merged verified change; why maturity is not monotone in usage.
- 08
The Assessment
The audit-checkable diagnostic; sequencing the unlocks; the ADLC as the reference implementation.
Locate your organization in under a minute.
A handful of audit-checkable questions — no sentiment, no self-scoring. Run it per workstream, not per company: incidents start with your outlier teams, not your organization-wide average.
Take the assessment →Or tell us what you’re rolling out and we’ll be in touch as the series and new diagnostics ship.