Working draft 0.3.0Generated from capabilities/CC BY 4.0

AI Harness Catalog

What a harness for an AI application must contain — by architecture archetype, mapped to the assurance obligations each part makes verifiable.

Two catalogs sit either side of one line. The AI Assurance Catalog (AAC) publishes the test obligations an AI application owes: what must be true of it. This one, the AI Harness Catalog (AHC), publishes the machinery those obligations assume: what must exist in order for any of them to be checkable at all.

Read enough assurance obligations and a phrase starts recurring — “enforced by the harness”, “outside the model's control”. That phrase is a dangling reference. This catalog is the referent.

Capabilities
—
Core (all shapes)
—
Harness layers
—
Archetypes
—
Design tensions
—
Obligations discharged
—
Section 1 — Orientation

AAC and AHC, in one page

Two repositories, one boundary. Neither is complete alone, and merging them would make one of the two redundant.

They are separate because they answer different questions. An obligation nobody knows how to build is a principle; a component nobody knows how to check is a hope. Keeping them apart is what lets each stay short.

Sibling catalog

AAC — AI Assurance Catalog

States what must be true.

108 test obligations for AI applications, tagged by which architecture archetype owes them, which class of machinery can produce a verdict, at which lifecycle stage, and whether failure blocks a release.

  • Unit — an obligation
  • Asks — is it right?
  • Answer is — a verdict
  • Axes — mechanism × stage
  • Fails when — a test does not pass

github.com/dataagentsai/ai-assurance-catalog

This catalog

AHC — AI Harness Catalog

States what must exist.

118 harness capabilities across 16 layers, tagged by which shapes need them, where in a system they can physically live, and — the substance — the design decisions their builder is forced to make.

  • Unit — a capability
  • Asks — what do I build?
  • Answer is — a component
  • Axes — layer × position
  • Fails when — a component does not exist

github.com/dataagentsai/ai-harness-catalog

A worked pair

The assurance catalog says, in AAC-0055: maximum steps, wall-clock and token budget are enforced by the harness — not requested in the prompt — and the stop path is tested including its partial-result behaviour.

Note “enforced by the harness.” The obligation cannot say where that counter lives, what the caller receives at cutoff, or whether the budget is enforced in the loop or at the gateway — those are construction questions, and answering them inside an assurance catalog would turn it into a framework manual. So a capability here states what the loop must contain, and cites discharges: [AAC-0055] back.

The rule, in one line

AAC states what must be true. AHC states what must exist. Properties, verified ⟷ components, built. Every “does this belong here?” question reduces to that sentence, and the linter enforces it mechanically: a requirement containing verification language — tested, asserted, measured, scored — is a build error, because that sentence belongs in the other repository.

What one capability looks like

Three fields carry the weight. failure_mode is what justifies the capability existing — a vague one means it is a principle, not a component. design_decisions is the substance; a decision with nothing genuinely traded away is not a decision. discharges is the join, and it runs both ways: an obligation no capability discharges is a hole in this catalog.

id: AHC-0004
level: MUST
layers: [L2]              # model invocation
positions: [P3, P2]       # harness library, gateway backstop
title: One choke point for every model call
requirement: >-
  Every model invocation in the system passes through a single component.
  Provider SDKs are constructed in that one place and nowhere else, and the
  set of reachable models is enumerable from it.
failure_mode: >-
  Once a second call path exists, every subsequent capability in this catalog
  acquires a hole. Cost accounting undercounts, spans go missing, the policy
  point is bypassed, and the model allow-list is advisory.
discharges: [AAC-0011, AAC-0100, AAC-0098, AAC-0094]     # <- the join
design_decisions:
  - question: In-process choke point, gateway, or both?
    tension: >-
      In-process sees types, caller identity and business context, but governs
      only code that imports it. A gateway covers everything on the network
      path and knows none of that context.
    resolution: >-
      Both, with the split made explicit. Write it down; the failure mode is
      each side assuming the other did it.

Conceptual now, physical later

Everything published today is conceptual and portable: what must exist, what breaks without it, and what you must decide. Nothing here names a product in normative text, and nothing here ships a component.

The physical layers come later and deliberately so — realizations (how each capability actually gets built, the only place products may appear) and reference skeletons (runnable, copied rather than imported). Products are the fastest-rotting layer; authoring them alongside each capability means writing them two or three times before the catalog stabilises. See the roadmap.

The tripwire

AHC describes what a component must do and what breaks without it. It never ships the component. If anything under references/ is ever published as an importable dependency rather than a skeleton to copy, the boundary has been crossed — and every capability in the catalog becomes an advertisement for that library.

The dependency direction

One-directional: AHC → AAC. This catalog cites assurance identifiers; the assurance catalog contains no reference to this one and does not need to know it exists. That protects the neutral half — assurance guidance competes with nobody, and construction guidance competes with every framework's documentation.

Section 2 — The harness

Everything that is neither the model nor your business logic

The deterministic scaffolding that turns a model into a system. Sixteen layers, and that count is the argument.

Observability is L11 and the eval harness is L12. That they are two of sixteen is the point of this catalog: teams routinely build those two, and call the harness done. Layers are tags, not a tree — a capability frequently spans two, because context assembly is also a cost concern and tool dispatch is also an authorization concern.

LayerNameWhat it coversOwns
L1Context assemblyTurning typed inputs into the bytes sent to the model: templating, prompt versioning, retrieval placement, ordering, compaction, token budget.What the model is told.
L2Model invocationThe call itself: client construction, parameters, streaming, timeouts, the choke point every call passes through, and which model is reachable at all.How the call is made.
L3Tool layerTool definition and schema, dispatch, argument validation, sandboxing, result shaping and truncation, and what a tool is allowed to reach.What the model can do.
L4Control loopIteration, step and budget accounting, termination conditions, planning and replanning, interruption and resumption.When the system stops.
L5State and memoryConversation state, working memory, checkpoints, long-term stores, and the durability boundary between what survives a crash and what does not.What is remembered.
L6I/O contractsThe typed boundary between the harness and its callers in both directions: input validation, response parsing, and the declared shape of failure.What the caller can rely on.
L7Policy enforcementThe positions at which a rule can actually be applied to traffic, what happens on detection, and the open-or-closed default.Where a rule becomes real.
L8Concurrency and flow controlParallel calls, fan-out limits, queueing, backpressure, and behaviour when a provider throttles rather than fails.What happens under load.
L9Determinism and replayPinned parameters, recorded fixtures, mocked providers, and the ability to re-run a past call without the network.Whether yesterday can be reproduced.
L10Failure handlingRetries, fallback, idempotency, compensation, partial results, and the declared degradation path when the model or a tool is unavailable.What happens when something breaks.
L11ObservabilitySpan emission, attribute population, trace-context propagation across out-of-process hops, and payload capture with its redaction point.Whether a past run can be explained.
L12Eval harnessThe seams that let an evaluation drive the system: a callable task entrypoint, dataset plumbing, injectable graders, and CI wiring.Whether the system can be measured at all.
L13Cost accountingToken and spend accounting attributed to a unit of work, attribution tags applied at the boundary, and ceilings that can actually stop a call.What it costs and who pays.
L14Human-in-the-loopApproval gates, escalation paths, the surface a reviewer sees, and what the system does while it waits.Where a person can intervene.
L15Release and configurationModel pins, prompt versions, feature flags, canary and rollback, and the resolution of all of it into a recorded, reproducible configuration.Which version is running.
L16Identity and authorizationWhose identity a call carries, propagation of that identity to tools and downstream systems, authorization of actions, secret handling and egress.Who the system is acting as.
Section 3 — Taxonomy

Ten shapes, and what each one adds

The archetype vocabulary is owned by the assurance catalog and pinned here, so the two catalogs join on the same ten shapes.

Boundaries are drawn by two questions only — who owns control flow, and what the output touches. Not by domain, sector or model size, because those do not change what you have to build. Every shape owes the core block in full; the count on each row is what is genuinely new about that shape, which is why this is 118 capabilities and not several hundred.

Section 4 — Construction axes

Where it lives, and who supplies it

Deliberately not the assurance catalog's axes. That one organises by mechanism and stage — verification axes. This one organises by position and approach.

Position is this catalog's spine. It is the axis architects actually argue about, and the same capability at a different position is a different system: a budget enforced in the loop and a budget enforced at the gateway catch different failures and miss different ones. On every capability, the first position listed is primary; the rest are backstops.

CodePositionWhat it can see, and what it cannot stop
P1Caller / edgeBefore the harness is entered — the API surface, the UI, the trigger. Anything enforced here is invisible to any other entry point, which is why so little belongs here.
P2Gateway / proxyOut of process, on the request path. The only position that covers calls made by code you do not own, and the only one that can enforce across languages. Also a hop, a dependency and a failure domain.
P3Harness libraryIn-process, wrapping the model call. Sees types, business context and the caller's intent; governs only code that imports it.
P4Control loopThe iteration itself. The only position that can see a trajectory — that these fourteen calls are one runaway task rather than fourteen tasks.
P5Tool boundaryBetween the decision to call a tool and the tool running. The last place an action can be stopped while it is still cheap to stop.
P6State storeThe durable substrate. Enforcement here survives process death, which is what makes it the right home for anything that must outlive a crash.
P7Offline / CIOut of band, not on the request path. Unlimited time budget, no user waiting, and no ability to prevent anything in production.
P8Human surfaceA queue, an approval, a console. Slow, expensive, and the only position that can exercise judgement the system does not have.

Approach answers who supplies the machinery. Five of the six are the assurance catalog's vocabulary, kept identical so an adopter filters both catalogs the same way. framework is the one deliberate addition: “adopt an agent framework or build the loop” is the defining harness question and has no assurance-side equivalent.

CodeApproachWhat you get, and what you inherit
in-houseBuild it yourselfYour own code, your own types. Most control and most portable; the strongest option for anything that encodes a decision specific to your domain, and the weakest use of time for anything commoditised.
frameworkAdopt an agent or workflow frameworkThe loop, state and tool plumbing arrive built. You inherit its control flow model and its opinions, including the ones you have not read yet. The question is never "is it good" but "which of my decisions does it make for me".
open-sourceOpen-source libraryA component you run yourself, doing one job. No vendor dependency; you own the upgrade treadmill.
platformHosted platformTracing, datasets, prompt registries and experiment history as a service. Fastest to useful, and where a growing share of your operational history ends up living — so the export path matters on day one.
gatewayAI gateway or proxyEnforcement on the request path, out of process. The only approach that covers applications you did not write, and the only one that can prevent rather than detect.
cloud-nativeCloud provider managed serviceCapability from the cloud you already run in. Low integration cost, procurement usually already done, opaque internals, hard to reproduce locally.

Levels follow RFC 2119 usage and bind only within an adopting organisation.

LevelMeaning
MUSTInherent to the shape. A harness without it has a hole that no amount of testing closes, because there is nothing to test. Omitting requires a recorded, owned decision.
SHOULDExpected in a production harness. Valid reasons to omit exist and belong in the design record.
MAYApplies when the stated condition holds — regulated domain, untrusted input, real side effects, external users, more than one team.
The threshold boundary

Sharpest line in the catalog. “A token budget is enforced by the assembly function and its application is recorded” — a capability. “The token budget is 100,000” — never. The number belongs to the adopting organisation, and any value published here would be wrong for almost everyone.

Section 5 — The seams

Where the harness meets what it does not own

A capability says what must exist. A port says where the harness meets something it did not build, and what must hold across that meeting whoever implements it.

The catalog is portable because it names no products, and the cost of that is a gap: a reader agrees that every model call needs one choke point and still has to invent the boundary between their code and a provider's library — and everyone invents a different one. Ports are that missing vocabulary, declared as data so an interface can be generated in any language without this repository ever shipping a package.

A port is not a partition of the catalog. Most capabilities are structural and cross no seam at all. Two fields carry the weight: kernel_owns, which states what stays on the harness side and is therefore not an implementation's to decide — a seam with nothing on the harness side is a client library, and the linter fails the build for it — and the substitution test, which is how you tell a port from one vendor's API with a wrapper on it.

Ports
—
Core — every harness
—
Capabilities crossing a seam
—
Structural, no seam
—

Operations are stated by intent, never by signature — no types, no language, no error taxonomy. Those belong to a generated interface, which is downstream of this catalog. A trailing ? marks an operation an implementation may legitimately not offer, in which case the harness needs a declared path for its absence.

Section 6 — The catalog

Capabilities, in plain English

Select a shape. Core is always owed; the archetype block is what that shape adds on top.

Each entry states what must exist, what breaks without it, where it can live, the obligations it makes verifiable, and — behind the disclosure — the decisions whoever builds it is forced to make.

Section 7 — Coverage matrix

Which shapes need which layers

Archetype deltas only; the core row is listed separately since it would otherwise appear in every column.

Read this as a design-review heat map. A dense cell is where that shape's harness work actually is. An empty cell for a layer you know you need means either the classification is wrong, or the requirement was already covered by core.

Section 8 — The join

Read from the assurance side

The same relation inverted: for each assurance obligation, what has to exist before anyone can check it.

This is the table that makes the pair useful at design review. An obligation with several capabilities behind it is one where the test is cheap and the construction is not. The join is also a completeness check in both directions — an obligation no capability discharges is a hole here, and a capability citing nothing is a component with no stated assurance consequence.

Obligations reached
—
Discharge references
—
Capabilities citing none
—

Obligation identifiers link to the assurance catalog's source. They are owned there, permanent, and stable in both directions.

Section 9 — Using it

Turning the catalog into an architecture

  • Classify the system. Decompose it into archetypes — most real systems are two or three. The classification is the assumption everything else rests on, and the thing most likely to be wrong.
  • Take the union. Core, plus the deltas for each shape present. That list is the harness you owe, before anyone has argued about a framework.
  • Pick a position for each. Loop, gateway, tool boundary, state store. This is where the real architecture argument lives, and writing the answer down is most of the value of the catalog.
  • Answer the design decisions. Each capability names the tensions its builder cannot avoid. An unanswered one is not neutral — it gets answered by default, usually by whichever framework was adopted first.
  • Record what you will not build. A skipped MUST becomes an explicit accepted risk with an owner and a review date. An honest gap beats a green diagram.
  • Cite identifiers in code. One identifier in one place — a module docstring, an architecture decision record, a review checklist. That costs nothing and requires no buy-in to the rest of the catalog.
# AHC-0004 — all model calls route through this client, nothing constructs a
# provider SDK directly. See docs/adr/0012-model-client.md
class ModelClient: ...

Identifiers are flat and permanent. Layer, position and archetype are metadata on a capability, never part of its identity — so re-tagging never breaks a citation, nothing is ever deleted, and renumbering is never correct.

Section 10 — Roadmap

What is normative today, and what comes next

The normative catalog is complete. What remains is the physical half — how each capability actually gets built, and skeletons to copy.

Phase 0Taxonomy, schema, and a linter carrying both boundary rules.done
Phase 1The core layer — the capabilities every shape owes regardless of architecture.done
Phase 2Archetype deltas for all ten shapes — what is genuinely new about each.done
Phase 3Blueprints — requirements, architecture and design assembled per archetype into ten pages.next
Phase 4Realizations — how each capability gets built, authored in one dated pass. The only layer where products may be named.planned
Phase 5Reference skeletons — three, one per control-flow tier. Copied, never imported.planned
Phase 6Bidirectional coverage against the assurance catalog — every obligation discharged, every capability cited.done

Phase 4 is deliberately last. Products rot faster than anything else in the catalog, and authoring them alongside each capability means writing them two or three times before the specification stabilises.