AAC and AHC, in one page
Two repositories, one boundary. Neither is complete alone, and merging them would make one of the two redundant.
They are separate because they answer different questions. An obligation nobody knows how to build is a principle; a component nobody knows how to check is a hope. Keeping them apart is what lets each stay short.
Sibling catalog
AAC — AI Assurance Catalog
States what must be true.
108 test obligations for AI applications, tagged by which architecture archetype owes them, which class of machinery can produce a verdict, at which lifecycle stage, and whether failure blocks a release.
- Unit — an obligation
- Asks — is it right?
- Answer is — a verdict
- Axes — mechanism × stage
- Fails when — a test does not pass
This catalog
AHC — AI Harness Catalog
States what must exist.
118 harness capabilities across 16 layers, tagged by which shapes need them, where in a system they can physically live, and — the substance — the design decisions their builder is forced to make.
- Unit — a capability
- Asks — what do I build?
- Answer is — a component
- Axes — layer × position
- Fails when — a component does not exist
A worked pair
The assurance catalog says, in AAC-0055: maximum steps, wall-clock and token budget are enforced by the harness — not requested in the prompt — and the stop path is tested including its partial-result behaviour.
Note “enforced by the harness.” The obligation cannot say where that counter lives, what the caller receives at cutoff, or whether the budget is enforced in the loop or at the gateway — those are construction questions, and answering them inside an assurance catalog would turn it into a framework manual. So a capability here states what the loop must contain, and cites discharges: [AAC-0055] back.
AAC states what must be true. AHC states what must exist. Properties, verified ⟷ components, built. Every “does this belong here?” question reduces to that sentence, and the linter enforces it mechanically: a requirement containing verification language — tested, asserted, measured, scored — is a build error, because that sentence belongs in the other repository.
What one capability looks like
Three fields carry the weight. failure_mode is what justifies the capability existing — a vague one means it is a principle, not a component. design_decisions is the substance; a decision with nothing genuinely traded away is not a decision. discharges is the join, and it runs both ways: an obligation no capability discharges is a hole in this catalog.
id: AHC-0004
level: MUST
layers: [L2] # model invocation
positions: [P3, P2] # harness library, gateway backstop
title: One choke point for every model call
requirement: >-
Every model invocation in the system passes through a single component.
Provider SDKs are constructed in that one place and nowhere else, and the
set of reachable models is enumerable from it.
failure_mode: >-
Once a second call path exists, every subsequent capability in this catalog
acquires a hole. Cost accounting undercounts, spans go missing, the policy
point is bypassed, and the model allow-list is advisory.
discharges: [AAC-0011, AAC-0100, AAC-0098, AAC-0094] # <- the join
design_decisions:
- question: In-process choke point, gateway, or both?
tension: >-
In-process sees types, caller identity and business context, but governs
only code that imports it. A gateway covers everything on the network
path and knows none of that context.
resolution: >-
Both, with the split made explicit. Write it down; the failure mode is
each side assuming the other did it.
Conceptual now, physical later
Everything published today is conceptual and portable: what must exist, what breaks without it, and what you must decide. Nothing here names a product in normative text, and nothing here ships a component.
The physical layers come later and deliberately so — realizations (how each capability actually gets built, the only place products may appear) and reference skeletons (runnable, copied rather than imported). Products are the fastest-rotting layer; authoring them alongside each capability means writing them two or three times before the catalog stabilises. See the roadmap.
AHC describes what a component must do and what breaks without it. It never ships the component. If anything under references/ is ever published as an importable dependency rather than a skeleton to copy, the boundary has been crossed — and every capability in the catalog becomes an advertisement for that library.
The dependency direction
One-directional: AHC → AAC. This catalog cites assurance identifiers; the assurance catalog contains no reference to this one and does not need to know it exists. That protects the neutral half — assurance guidance competes with nobody, and construction guidance competes with every framework's documentation.
Everything that is neither the model nor your business logic
The deterministic scaffolding that turns a model into a system. Sixteen layers, and that count is the argument.
Observability is L11 and the eval harness is L12. That they are two of sixteen is the point of this catalog: teams routinely build those two, and call the harness done. Layers are tags, not a tree — a capability frequently spans two, because context assembly is also a cost concern and tool dispatch is also an authorization concern.
| Layer | Name | What it covers | Owns |
|---|---|---|---|
| L1 | Context assembly | Turning typed inputs into the bytes sent to the model: templating, prompt versioning, retrieval placement, ordering, compaction, token budget. | What the model is told. |
| L2 | Model invocation | The call itself: client construction, parameters, streaming, timeouts, the choke point every call passes through, and which model is reachable at all. | How the call is made. |
| L3 | Tool layer | Tool definition and schema, dispatch, argument validation, sandboxing, result shaping and truncation, and what a tool is allowed to reach. | What the model can do. |
| L4 | Control loop | Iteration, step and budget accounting, termination conditions, planning and replanning, interruption and resumption. | When the system stops. |
| L5 | State and memory | Conversation state, working memory, checkpoints, long-term stores, and the durability boundary between what survives a crash and what does not. | What is remembered. |
| L6 | I/O contracts | The typed boundary between the harness and its callers in both directions: input validation, response parsing, and the declared shape of failure. | What the caller can rely on. |
| L7 | Policy enforcement | The positions at which a rule can actually be applied to traffic, what happens on detection, and the open-or-closed default. | Where a rule becomes real. |
| L8 | Concurrency and flow control | Parallel calls, fan-out limits, queueing, backpressure, and behaviour when a provider throttles rather than fails. | What happens under load. |
| L9 | Determinism and replay | Pinned parameters, recorded fixtures, mocked providers, and the ability to re-run a past call without the network. | Whether yesterday can be reproduced. |
| L10 | Failure handling | Retries, fallback, idempotency, compensation, partial results, and the declared degradation path when the model or a tool is unavailable. | What happens when something breaks. |
| L11 | Observability | Span emission, attribute population, trace-context propagation across out-of-process hops, and payload capture with its redaction point. | Whether a past run can be explained. |
| L12 | Eval harness | The seams that let an evaluation drive the system: a callable task entrypoint, dataset plumbing, injectable graders, and CI wiring. | Whether the system can be measured at all. |
| L13 | Cost accounting | Token and spend accounting attributed to a unit of work, attribution tags applied at the boundary, and ceilings that can actually stop a call. | What it costs and who pays. |
| L14 | Human-in-the-loop | Approval gates, escalation paths, the surface a reviewer sees, and what the system does while it waits. | Where a person can intervene. |
| L15 | Release and configuration | Model pins, prompt versions, feature flags, canary and rollback, and the resolution of all of it into a recorded, reproducible configuration. | Which version is running. |
| L16 | Identity and authorization | Whose identity a call carries, propagation of that identity to tools and downstream systems, authorization of actions, secret handling and egress. | Who the system is acting as. |
Ten shapes, and what each one adds
The archetype vocabulary is owned by the assurance catalog and pinned here, so the two catalogs join on the same ten shapes.
Boundaries are drawn by two questions only — who owns control flow, and what the output touches. Not by domain, sector or model size, because those do not change what you have to build. Every shape owes the core block in full; the count on each row is what is genuinely new about that shape, which is why this is 118 capabilities and not several hundred.
Where it lives, and who supplies it
Deliberately not the assurance catalog's axes. That one organises by mechanism and stage — verification axes. This one organises by position and approach.
Position is this catalog's spine. It is the axis architects actually argue about, and the same capability at a different position is a different system: a budget enforced in the loop and a budget enforced at the gateway catch different failures and miss different ones. On every capability, the first position listed is primary; the rest are backstops.
| Code | Position | What it can see, and what it cannot stop |
|---|---|---|
| P1 | Caller / edge | Before the harness is entered — the API surface, the UI, the trigger. Anything enforced here is invisible to any other entry point, which is why so little belongs here. |
| P2 | Gateway / proxy | Out of process, on the request path. The only position that covers calls made by code you do not own, and the only one that can enforce across languages. Also a hop, a dependency and a failure domain. |
| P3 | Harness library | In-process, wrapping the model call. Sees types, business context and the caller's intent; governs only code that imports it. |
| P4 | Control loop | The iteration itself. The only position that can see a trajectory — that these fourteen calls are one runaway task rather than fourteen tasks. |
| P5 | Tool boundary | Between the decision to call a tool and the tool running. The last place an action can be stopped while it is still cheap to stop. |
| P6 | State store | The durable substrate. Enforcement here survives process death, which is what makes it the right home for anything that must outlive a crash. |
| P7 | Offline / CI | Out of band, not on the request path. Unlimited time budget, no user waiting, and no ability to prevent anything in production. |
| P8 | Human surface | A queue, an approval, a console. Slow, expensive, and the only position that can exercise judgement the system does not have. |
Approach answers who supplies the machinery. Five of the six are the assurance catalog's vocabulary, kept identical so an adopter filters both catalogs the same way. framework is the one deliberate addition: “adopt an agent framework or build the loop” is the defining harness question and has no assurance-side equivalent.
| Code | Approach | What you get, and what you inherit |
|---|---|---|
| in-house | Build it yourself | Your own code, your own types. Most control and most portable; the strongest option for anything that encodes a decision specific to your domain, and the weakest use of time for anything commoditised. |
| framework | Adopt an agent or workflow framework | The loop, state and tool plumbing arrive built. You inherit its control flow model and its opinions, including the ones you have not read yet. The question is never "is it good" but "which of my decisions does it make for me". |
| open-source | Open-source library | A component you run yourself, doing one job. No vendor dependency; you own the upgrade treadmill. |
| platform | Hosted platform | Tracing, datasets, prompt registries and experiment history as a service. Fastest to useful, and where a growing share of your operational history ends up living — so the export path matters on day one. |
| gateway | AI gateway or proxy | Enforcement on the request path, out of process. The only approach that covers applications you did not write, and the only one that can prevent rather than detect. |
| cloud-native | Cloud provider managed service | Capability from the cloud you already run in. Low integration cost, procurement usually already done, opaque internals, hard to reproduce locally. |
Levels follow RFC 2119 usage and bind only within an adopting organisation.
| Level | Meaning |
|---|---|
| MUST | Inherent to the shape. A harness without it has a hole that no amount of testing closes, because there is nothing to test. Omitting requires a recorded, owned decision. |
| SHOULD | Expected in a production harness. Valid reasons to omit exist and belong in the design record. |
| MAY | Applies when the stated condition holds — regulated domain, untrusted input, real side effects, external users, more than one team. |
Sharpest line in the catalog. “A token budget is enforced by the assembly function and its application is recorded” — a capability. “The token budget is 100,000” — never. The number belongs to the adopting organisation, and any value published here would be wrong for almost everyone.
Where the harness meets what it does not own
A capability says what must exist. A port says where the harness meets something it did not build, and what must hold across that meeting whoever implements it.
The catalog is portable because it names no products, and the cost of that is a gap: a reader agrees that every model call needs one choke point and still has to invent the boundary between their code and a provider's library — and everyone invents a different one. Ports are that missing vocabulary, declared as data so an interface can be generated in any language without this repository ever shipping a package.
A port is not a partition of the catalog. Most capabilities are structural and cross no seam at all. Two fields carry the weight: kernel_owns, which states what stays on the harness side and is therefore not an implementation's to decide — a seam with nothing on the harness side is a client library, and the linter fails the build for it — and the substitution test, which is how you tell a port from one vendor's API with a wrapper on it.
Operations are stated by intent, never by signature — no types, no language, no error taxonomy. Those belong to a generated interface, which is downstream of this catalog. A trailing ? marks an operation an implementation may legitimately not offer, in which case the harness needs a declared path for its absence.
Capabilities, in plain English
Select a shape. Core is always owed; the archetype block is what that shape adds on top.
Each entry states what must exist, what breaks without it, where it can live, the obligations it makes verifiable, and — behind the disclosure — the decisions whoever builds it is forced to make.
Which shapes need which layers
Archetype deltas only; the core row is listed separately since it would otherwise appear in every column.
Read this as a design-review heat map. A dense cell is where that shape's harness work actually is. An empty cell for a layer you know you need means either the classification is wrong, or the requirement was already covered by core.
Read from the assurance side
The same relation inverted: for each assurance obligation, what has to exist before anyone can check it.
This is the table that makes the pair useful at design review. An obligation with several capabilities behind it is one where the test is cheap and the construction is not. The join is also a completeness check in both directions — an obligation no capability discharges is a hole here, and a capability citing nothing is a component with no stated assurance consequence.
Obligation identifiers link to the assurance catalog's source. They are owned there, permanent, and stable in both directions.
Turning the catalog into an architecture
- Classify the system. Decompose it into archetypes — most real systems are two or three. The classification is the assumption everything else rests on, and the thing most likely to be wrong.
- Take the union. Core, plus the deltas for each shape present. That list is the harness you owe, before anyone has argued about a framework.
- Pick a position for each. Loop, gateway, tool boundary, state store. This is where the real architecture argument lives, and writing the answer down is most of the value of the catalog.
- Answer the design decisions. Each capability names the tensions its builder cannot avoid. An unanswered one is not neutral — it gets answered by default, usually by whichever framework was adopted first.
- Record what you will not build. A skipped MUST becomes an explicit accepted risk with an owner and a review date. An honest gap beats a green diagram.
- Cite identifiers in code. One identifier in one place — a module docstring, an architecture decision record, a review checklist. That costs nothing and requires no buy-in to the rest of the catalog.
# AHC-0004 — all model calls route through this client, nothing constructs a # provider SDK directly. See docs/adr/0012-model-client.md class ModelClient: ...
Identifiers are flat and permanent. Layer, position and archetype are metadata on a capability, never part of its identity — so re-tagging never breaks a citation, nothing is ever deleted, and renumbering is never correct.
What is normative today, and what comes next
The normative catalog is complete. What remains is the physical half — how each capability actually gets built, and skeletons to copy.
Phase 4 is deliberately last. Products rot faster than anything else in the catalog, and authoring them alongside each capability means writing them two or three times before the specification stabilises.