A programme can be doing Phase 4 work while operating at Level 2. That sentence sounds contradictory only because adoption roadmaps and capability assessments are routinely collapsed into one maturity ladder.
Suppose a platform team is building system context for agents: service ownership, dependency maps, policy boundaries and production telemetry. That is advanced adoption work. Meanwhile the pilot repositories may still have unreliable tests, ambiguous acceptance criteria and no tested rollback path. The programme has moved forward. Its safe delegation boundary has not moved nearly as far.
This is a practical category error, not a dispute about labels. Roadmap progress answers what we are changing. Capability evidence answers what an agent may be trusted and authorized to do now.
Three concepts, three questions
I use three concepts in the AI-maturity series:
| Concept | Question it answers | Assessment key |
|---|---|---|
| Phase | What adoption or transformation work is under way? | Programme or workstream |
| Capability level | What can this scope perform repeatedly, with the required controls and evidence? | Bounded scope—person, repository, team, system or platform—plus change class, authority boundary and assessment time window |
| Readiness dimension | Which evidence justifies—or contradicts—the claimed level? | A property such as specification, verification, IAM or reversibility, evaluated for the same assessment key |
That is the distinction. The rest of the model exists to make it usable.
A phase is not necessarily completed once, everywhere, in order. Different workstreams can overlap, retreat or advance at different speeds. A capability level is not a reward for completing a phase. It is an assessment for an explicit scope, change class, authority boundary and period of time.
This matters because “we are Level 4” is almost content-free. “The payments API repositories can delegate low-risk dependency updates through staging, but not production promotion, because the listed checks and rollback controls passed during the last assessment window” can be examined. The second statement is less impressive at a town hall and much more useful during an incident.
Model boundary
The scheme below is the demming.dev working model, version 0.1. It is author synthesis for agentic software delivery, not an industry standard, a CMMI appraisal, or a claim of certification.
Within this series, maturity level is shorthand for this site-owned capability level. It is not the CMMI maturity-level term.
That boundary is deliberate. CMMI has its own defined capability and maturity levels, scopes and appraisal method. Its distinction between capability in individual practice areas and maturity across predefined sets of practice areas is useful context, but those definitions should not be silently repurposed for this model. See the CMMI Institute’s level definitions.
The model also should not turn governance into the final phase of an otherwise technical rollout. The NIST AI RMF Core describes governance as cross-cutting and says its actions are not necessarily an ordered set of steps. ISO/IEC 42001 similarly describes an organizational management system which is established, maintained and continually improved. Neither source defines the levels below. They support the narrower point that management, evidence and improvement continue across the lifecycle.
The adoption phases
The phases describe a plausible programme vocabulary, not a mandatory implementation sequence:
| Phase | Adoption work |
|---|---|
| Phase 0 | No organized AI-engineering practice; use, if any, is unmanaged or not visible as a programme. |
| Phase 1 | Individual assistance and discovery: people learn where models help, fail and create new review cost. |
| Phase 1.5 | A durable, human-supervised workflow connects accepted intent, agent execution, review and CI evidence. GitHub issues and pull requests are one implementation, not the definition. |
| Phase 2 | Repository-readiness work makes build paths, constraints, tests, ownership and stopping conditions legible. |
| Phase 3 | Team operating models standardize change classes, review, exception handling, metrics and learning. |
| Phase 4 | System-aware governance connects repositories to dependencies, service ownership, policy, environments and operational evidence. |
| Phase 5 | Bounded autonomy is introduced for selected change classes whose controls and reversibility justify it. |
Phase 1.5 remains in the vocabulary because it names a real bridge. A personal chat may be useful, but it leaves little durable organizational evidence. A human-supervised delivery record can preserve intent, the proposed change, verification results, review decisions and why execution stopped. It does not prove repo readiness by itself. It makes the missing readiness visible.
There is no requirement to wait for every Phase 2 activity before learning about team policy, nor to finish a single enterprise platform before trying a bounded automation. Programmes are messier than diagrams. The phase label is a coordination aid, nothing more.
The five capability levels
The levels describe demonstrated delivery capability:
| Level | Capability state | Evidence boundary |
|---|---|---|
| Level 1 — Individual AI Assistance | A person uses AI within their existing authority and remains the execution and review boundary. | Useful outcomes may exist, but there is no repeatable organizational delegation contract. |
| Level 2 — Repository-Ready Agentic Workflow | A human-supervised agent can perform approved change classes in selected repositories. | The repository exposes instructions, constraints, deterministic checks, ownership and stopping behaviour. |
| Level 3 — Team-Managed Agentic SDLC | A team-managed workflow applies shared risk classes, review rules, evidence and exceptions across its scope. | The practice survives beyond one expert, tool or unusually well-prepared repository. |
| Level 4 — System-Aware Governed Agents | Agents can operate across explicitly modelled system boundaries without receiving unrestricted authority or context. | Identity, least privilege, dependency and ownership context, policy enforcement, audit and environment controls are demonstrated. |
| Level 5 — Bounded Autonomous Delivery | Bounded autonomy is demonstrated when selected low-risk change classes can proceed without synchronous human execution or review at every step. | The approved envelope includes reliable verification, monitoring, promotion limits, rollback and accountable exception ownership. |
These are levels of the engineering and management system, not levels of model intelligence. A stronger model may improve execution inside a boundary. It does not create production authority, repair an untestable repository or decide which residual risk the organization accepts.
“Demonstrated” also means repeatable under ordinary conditions. A carefully staged demo can reveal possibilities; it cannot establish a capability level. The assessment needs representative work, recorded failures as well as successes, and evidence that the control path behaves when the agent proposes something it is not allowed to do.
Readiness is the evidence layer
The level claim is tested across readiness dimensions. The exact inventory depends on risk and context, but the recurring dimensions in this series are:
- intent and specification quality;
- repository and architecture legibility;
- verification and evaluation strength;
- identity, authorization and governance enforcement;
- system context and ownership;
- operational observability and reversibility; and
- measurement and review discipline.
They are deliberately not a vendor checklist. A team can implement durable work records in GitHub, GitLab or another controlled delivery system. Infrastructure boundaries may be expressed through Terraform, another declarative system, or controls outside infrastructure as code. The assessment concerns the property and its evidence, not whether a fashionable file or product name appears.
Nor should the dimensions be averaged into a vanity score. For a production database migration, weak rollback evidence may bound delegation regardless of excellent repository instructions. For a documentation correction, the same weakness may be irrelevant. This is why safe delegation capacity is assessed per change class and authority boundary.
A crosswalk, not a staircase
Phase work can create evidence for several levels, and a level gap can redirect the roadmap:
| Programme observation | Capability reading | Decision |
|---|---|---|
| Phase 1.5 trials produce clean pull requests but repeatedly misunderstand acceptance criteria. | Execution is promising; the specification dimension does not yet support Level 2 for that change class. | Improve issue discovery and acceptance evidence before widening delegation. |
| Phase 3 standardization exists on paper, but only one engineer can recover failed runs. | The team has adopted conventions without demonstrating a Level 3 operating capability. | Test failure handling, ownership transfer and exception paths. |
| Phase 4 platform work exposes dependencies and policy, while several repositories still cannot run deterministic tests. | System context has improved, but those repositories remain below Level 2 for code changes. | Keep authority narrow and fund repository readiness explicitly. |
| A low-risk Phase 5 workflow has reliable checks, promotion limits and rollback in one domain. | Level 5 may be justified for that bounded change class—not for the organization as a whole. | Record the envelope and resist generalizing the label. |
This crosswalk is where the model earns its keep. Later-phase work is not fraudulent because an earlier capability is weak; it may be exactly the work needed to fix that weakness. The error is converting activity into authority without the evidence in between.
So I would not ask an organization which phase it is in, full stop. I would ask which workstreams are active, which scope and change class are being assessed, what level is claimed there, what evidence supports it, and which failed test would cause the claim to be withdrawn.
If those questions cannot be answered, the number is decoration.