Working draft 0.16.0Generated from catalog/CC BY 4.0

AI Assurance Catalog

Test obligations for AI applications, organised by architecture archetype, mapped to the machinery that can discharge them.

Governance frameworks say a system must be validated. Threat taxonomies say what can go wrong. Telemetry conventions say how to record a call. Metric libraries say how to score one property. None of them tell an architect, at design review: you are building this shape, therefore you owe these tests. That binding is what this catalog is.

Archetypes
—
Obligations
—
Core (all shapes)
—
Release gates
—
Build options
—
Section 1 — Scope

What this is, and what it refuses to be

A catalog of obligations organised by application shape, each mapped to how it can be physically realised.

It is not a benchmark, a metric implementation, a control set, or a vendor comparison. It does not tell you what score is good enough — thresholds belong to you, because they depend on your risk appetite and your users. It tells you which questions you are obliged to have an answer to, and which class of machinery can produce that answer.

Levels follow RFC 2119 usage and bind only within an adopting organisation.

LevelMeaning
MUSTInherent to the shape. A system without it cannot claim conformance. Skipping requires a recorded accepted risk with a named owner.
SHOULDExpected in production systems. Valid reasons to omit exist and must be stated in the risk register.
MAYApplies when the stated condition holds — regulated domain, high blast radius, external users.
The inheritance rule

Obligations compose downward. Every archetype owes the core block in full, so each archetype section lists only what is new about its risk surface. That is why this is 118 obligations rather than several hundred.

Identifiers

Identifiers are flat and permanent — archetype is metadata, never identity, so re-tagging a case never breaks a citation. Draft 0.1's archetype-scoped identifiers are shown as was A6-01 for traceability and are not citable.

Section 2 — Taxonomy

Ten shapes, and the compositions of them

Archetypes are distinguished by who owns control flow and what the output touches — not by domain or sector.

Those two questions are what actually change the test surface. A legal summariser and a medical summariser owe identical tests; a summariser and a tool-using agent do not. Real systems are usually two or three archetypes composed: classify each component, take the union, then test the seams between them.

Section 3 — Realization axes

From English to machinery

Every obligation carries three tags: the mechanism that computes the verdict, the stage where it runs, and whether failure blocks.

Keeping these orthogonal is the point. The same obligation — output must not leak PII — is a deterministic detector in CI, a blocking gateway filter at runtime, and a sampled model-graded check in production. One sentence, three realizations, different machinery. A catalog that fused them would force you to pick a vendor before you had finished thinking.

CodeMechanismGood for, and where it breaks
M1Deterministic assertionExact match, regex, JSON-schema validation, parser or compiler exit code. Cheap, fast, zero ambiguity. Cannot judge quality — only conformance.
M2Programmatic metricField-level F1, recall@k, p95 latency, cost per successful task. Needs labelled data or instrumentation; yields a number you can trend and threshold.
M3Model-graded, reference-freeA model scores the output against a written rubric. The only practical mechanism for open-ended quality, and it must itself be validated — the judge obligations in `owes` apply to it, whatever the archetype of the system it grades.
M4Model-graded, reference-basedCompare against a gold answer, or pairwise against the previous release. More stable than M3 because the judge has an anchor. Requires curated references.
M5Trace assertionAssert over emitted spans: structure, ordering, attribute presence, step count, token totals. The only way to test HOW a system reached an answer rather than what it said.
M6Human reviewAnnotation queue or SME sign-off. Ground truth for everything else; the labels it produces are what validate M3 and M4. Does not scale, so spend it on calibration.
M7Adversarial suiteGenerated attacks, input mutation, fault injection, fuzzing. Coverage grows over time; every production incident should become a new case here.
CodeStageTrigger, budget, and who sees a failure
S1Dev loopRun by hand while iterating. Seconds. Nobody but the author sees a failure.
S2CI, pre-mergeEvery change to prompt, schema, tool definition or model pin. Minutes, cheap and deterministic — mock the model where the test is about plumbing.
S3Pre-release evalFull offline dataset against the release candidate. Real model calls, real cost. This is where release gates live.
S4Runtime gatewayInline on the live request path, blocking. A failure is seen by the end user as a refusal, redaction or fallback.
S5Production online evalAsynchronous scoring of sampled live traces. Raises an alert, never blocks. The only stage that sees your real input distribution.
S6Scheduled regressionPeriodic re-run of frozen cases to catch provider-side model drift and corpus staleness. Nothing changed on your side — that is the point.

Tool classes name a class, never a product. Products date the moment the market moves, and naming one turns a neutral catalog into a recommendation — so the linter fails the build if a product name appears in an obligation.

Section 4 — The catalog

Obligations, in plain English

Select a shape. Core is always owed; the archetype block is what that shape adds on top.

Section 5 — Coverage matrix

Which shapes carry which risks

Archetype deltas only; the core row is listed separately since it would otherwise appear in every column.

Read this as a design-review heat map. A dense column for your shape is where the test budget belongs. An empty cell for a risk you know you have means the classification is wrong.

Crosswalk — eu-ai-act

What an existing framework maps to

Informative. 18 entries, mapped to obligations across the archetypes that actually owe them; 22 more recorded below as unmapped, with the reason.

The external framework is organised by its own axis — threat, risk-management outcome, control or article; this catalog is organised by application shape. The mapping is many-to-many by construction, and that is the point — a single external item lands on several obligations across several archetypes, which is exactly what the framework on its own cannot tell you.

Crosswalks release out of band: a revision upstream must never force a version bump in the catalog. relation records how tight each mapping is, because claiming a test obligation is equivalent to a governance control is the fastest way to have a crosswalk dismissed.

EntryNameRelationObligations
Art. 9(6) Testing to identify risk management measures and confirm consistent performance evidence-for AAC-0001AAC-0005AAC-0010AAC-0098
"Perform consistently for their intended purpose": accuracy on a frozen set, the safety corpus, the run-to-run noise floor, and every route the system can serve from. Any covered row in a coverage report is testing in this sense; these are the ones that speak to consistency.
Art. 9(8) Testing before placing on the market, against prior defined metrics and probabilistic thresholds partial AAC-0001AAC-0013AAC-0102
The catalog obliges a frozen set, a regression comparison and a cost gate run before promotion. The metrics and probabilistic thresholds the article requires are the provider's to define — the catalog deliberately sets none (docs/NON-GOALS.md) — so the mapping is partial by design.
Art. 10(3) Testing data sets relevant, representative, free of errors, complete partial AAC-0001AAC-0031
Under Art. 10(6), a system that does not train models owes Art. 10(2) to (5) for its testing data sets only, which is the usual position of an LLM application. AAC-0001 requires the curated, version-controlled set to exist; AAC-0031 requires the unanswerable cases a set curated from answerable questions never contains, which is completeness in practice. Neither tests representativeness or error rate of the set itself.
Art. 12(1) Automatic recording of events (logs) over the lifetime of the system evidence-for AAC-0011AAC-0060AAC-0080
AAC-0011 is asserted in CI, so a log field that stops being written fails a build instead of being discovered when it is needed.
Art. 12(2) Logs enable identifying risk situations and substantial modification, post-market monitoring, and deployer monitoring evidence-for AAC-0012AAC-0100AAC-0080AAC-0114
Point (a), substantial modification, is the one teams miss: recording the model version, prompt revision and serving route per request (AAC-0012, AAC-0100) is what lets a provider show from the logs that what ran is what was assessed. Points (b) and (c) are served by the same records.
Art. 13(3)(b)(ii) Instructions for use state the level of accuracy, robustness and cybersecurity tested and validated partial AAC-0001AAC-0017AAC-0018
These produce the measured numbers the instructions must state, sliced per segment (AAC-0017). Writing the instructions for use is the provider's; the article's robustness and cybersecurity levels draw on the cases under Art. 15(4) and 15(5).
Art. 14(1) Designed so natural persons can effectively oversee the system partial AAC-0078AAC-0043AAC-0056
Tests that the oversight paths built into the design work — approval before irreversible action, including when the facts changed during the pause; handoff to a human with complete context; confirmation before destructive tools. Whether the oversight is effective in the article's sense is not a test result.
Art. 14(4)(a) Overseer can monitor operation and detect anomalies partial AAC-0114AAC-0079
Anomalies in behaviour rates alert a person (AAC-0114), and failure and silence are both loud (AAC-0079). Understanding the system's capacities and limitations, the first half of point (a), is not covered.
Art. 14(4)(d) Overseer can disregard, override or reverse the output partial AAC-0081
Shaky: AAC-0081 makes each side-effecting action reversible, which is the precondition for "reverse" when the output is an action. Disregarding or overriding a non-action output is a user-interface matter not covered.
Art. 14(4)(e) Overseer can intervene or interrupt through a stop procedure to a safe state partial AAC-0055AAC-0077
Shaky: both cases test that a stop enforced outside the agent's own logic reaches a defined state with its partial result handled. They are triggered by budgets, not by a person; a human-initiated stop reaching the same path is not separately obliged.
Art. 15(1) Appropriate accuracy, robustness and cybersecurity, performing consistently throughout the lifecycle evidence-for AAC-0001AAC-0013AAC-0016AAC-0082AAC-0014AAC-0098
"Throughout their lifecycle" is the half the scheduled re-runs carry: a hosted model changes with nothing changed on the provider's side (AAC-0016, AAC-0082), and live traffic is scored on the same rubrics as offline (AAC-0014). What level is appropriate is the provider's.
Art. 15(3) Accuracy levels and metrics declared in instructions for use partial AAC-0001AAC-0017
Produces the measurements to declare. The declaration is the provider's.
Art. 15(4) Resilience to errors, faults and inconsistencies; fail-safe plans; feedback loops partial AAC-0009AAC-0015AAC-0019AAC-0046AAC-0053AAC-0065AAC-0091
The first subparagraph is well covered: provider faults, degenerate input, partial failure in a graph, tool errors, a contained sub-agent failure, and a guardrail outage failing closed — "backup or fail-safe plans" tested rather than written down. The feedback-loop subparagraph for systems that continue to learn after placement is not covered; long-term memory (AAC-0041) is related but is not learning in the article's sense, and is not claimed.
Art. 15(5) Resilience against attempts by unauthorised third parties to alter use, outputs or performance partial AAC-0004AAC-0036AAC-0058AAC-0106AAC-0107AAC-0108AAC-0057AAC-0111
Adversarial inputs, confidentiality attacks and tampered pre-trained components are covered at application level. Training-data and model poisoning, both named in the article, are model-level and out of scope; AAC-0108 covers poisoning of a retrieval corpus, which the article does not name but which has the same effect on outputs.
Art. 17(1)(d) QMS includes examination, test and validation procedures and their frequency evidence-for AAC-0013AAC-0034AAC-0090AAC-0101AAC-0016
Evidence that test procedures are wired to the events that need them — every release, index rebuild, judge change and routing change — and run on a schedule. The quality management system that records them is the provider's.
Art. 26(5) Deployer monitors operation on the basis of the instructions for use evidence-for AAC-0114AAC-0116AAC-0079
Art. 72(2) Post-market monitoring collects and analyses performance data throughout the lifetime evidence-for AAC-0014AAC-0114AAC-0115AAC-0116AAC-0082
AAC-0115 is the piece most monitoring lacks: later outcomes joined to the run that produced them, so performance is measured on what happened rather than only on what the system said.
Art. 50(1) Persons informed they are interacting with an AI system partial AAC-0003
Shaky. Where the disclosure is made in the model's own output, it is a required-disclaimer constraint and AAC-0003 tests it like any other. Most disclosures are made by the interface, which no obligation covers.

Unmapped. No obligation provides evidence for these. Recorded rather than padded: a crosswalk that maps everything is claiming more than a test catalog can show.

EntryNameWhy not
Art. 5 Prohibited AI practices A question of what the system is for, decided before any test runs.
Art. 9(1)-(5), 9(7), 9(9)-(10) Risk management system — establishment, steps, measures, real-world testing, vulnerable groups, integration The risk management process itself. Testing is mapped at 9(6) and 9(8); real-world testing under Arts. 57 and 60 is a regulatory procedure.
Art. 10(2) Data governance and management practices Governance practices for data, including the bias examination in points (f) and (g), which is out of scope per docs/NON-GOALS.md.
Art. 10(5) Processing special categories of personal data for bias detection Fairness and bias detection is out of scope per docs/NON-GOALS.md.
Art. 11 Technical documentation (Annex IV) Documentation. A dated series of coverage reports is a candidate input to Annex IV point 2(g) on validation and testing, but no obligation discharges the documentation duty.
Art. 12(3) Minimum logging for remote biometric identification systems Specific to Annex III point 1(a) systems; no obligation is written for that use.
Art. 13(1) Operation sufficiently transparent for deployers to interpret output A design quality judged by the deployer. Citation support (AAC-0030) aids interpretation of a grounded answer but is not claimed as transparency in the article's sense.
Art. 13(3) other points Instructions for use — remaining content Documentation content; only point (b)(ii) is evidenced by test results.
Art. 14(4)(b) Overseer aware of automation bias Awareness of people; not a system property a test can show.
Art. 14(4)(c) Overseer can correctly interpret the output Interpretation by people; see Art. 13(1).
Art. 14(5) Separate verification by two natural persons for remote biometric identification Specific to Annex III point 1(a) systems.
Art. 15(2) Commission encourages benchmarks and measurement methodologies An obligation on the Commission, not on a system.
Art. 17(1) other points Quality management system — remaining elements Organisational elements of the QMS; only point (d) turns on test procedures.
Art. 19(1) Provider keeps automatically generated logs A retention duty. Note the interaction with AAC-0095, which obliges bounded log retention: the bound an adopter sets must also satisfy the minimum period in this article and in Art. 26(6). The catalog sets no period.
Art. 26(1)-(4), 26(6)-(12) Deployer obligations other than monitoring Use per instructions, assigned oversight, input data relevance, log retention, informing workers and affected persons — organisational duties. Monitoring is mapped at 26(5).
Art. 50(2) Synthetic output marked in a machine-readable format and detectable No obligation covers marking or watermarking generated content — a candidate gap, shared with NIST AI 600-1 §2.8.
Art. 50(3) Informing persons exposed to emotion recognition or biometric categorisation Specific to those system types; no obligation is written for them.
Art. 50(4) Disclosure of deep fakes and of generated text on matters of public interest A deployer's disclosure duty for published content.
Art. 72(1), 72(3) Post-market monitoring system established and based on a plan Establishing and documenting the system and its plan; its data collection is mapped at 72(2).
Art. 73 Reporting of serious incidents Reporting to authorities is a procedure.
Art. 53 Obligations for providers of general-purpose AI models Model-provider obligations; this catalog is application-level.
Art. 55 Obligations for providers of general-purpose AI models with systemic risk Includes adversarial testing of the model (point (1)(a)), which is model-level and out of scope per docs/NON-GOALS.md.
Crosswalk — iso-iec-42001

What an existing framework maps to

Informative. 6 entries, mapped to obligations across the archetypes that actually owe them; 32 more recorded below as unmapped, with the reason.

The external framework is organised by its own axis — threat, risk-management outcome, control or article; this catalog is organised by application shape. The mapping is many-to-many by construction, and that is the point — a single external item lands on several obligations across several archetypes, which is exactly what the framework on its own cannot tell you.

Crosswalks release out of band: a revision upstream must never force a version bump in the catalog. relation records how tight each mapping is, because claiming a test obligation is equivalent to a governance control is the fastest way to have a crosswalk dismissed.

EntryNameRelationObligations
A.6.2.4 AI system verification and validation evidence-for AAC-0001AAC-0002AAC-0003AAC-0013AAC-0098AAC-0099
Nearly every covered row of an adopter's coverage report is evidence for this control; listed are the core obligations that frame verification against the declared contract (AAC-0002, AAC-0003, AAC-0099) and validation against intended use (AAC-0001, AAC-0013, AAC-0098). A dated series of reports from CI is the evidence form a surveillance audit prefers — see docs/REPORT.md.
A.6.2.5 AI system deployment evidence-for AAC-0107AAC-0012AAC-0101AAC-0102AAC-0034
Release-path evidence: what serves is what was evaluated (AAC-0107), rollback is exercised (AAC-0012), and the changes that conventionally bypass evaluation — routing configuration, cost, index rebuilds — are gated like a model change.
A.6.2.6 AI system operation and monitoring evidence-for AAC-0014AAC-0114AAC-0116AAC-0016AAC-0082AAC-0079AAC-0115
A.6.2.8 AI system recording of event logs evidence-for AAC-0011AAC-0060AAC-0080AAC-0100AAC-0095
AAC-0095 belongs here as the counterweight: logs are redacted at write time and retention is bounded, and the two obligations hold together rather than trading off.
A.7.5 Data provenance partial AAC-0108
Retrieval-corpus provenance only. Provenance of training and fine-tuning data is model-level and out of scope per docs/NON-GOALS.md. Mapped on the reading that Annex A's data controls reach data the system uses in operation, not only data it was trained on; confirm against your copy.
A.10.3 Suppliers evidence-for AAC-0094AAC-0107AAC-0016AAC-0098
The model provider is the supplier that matters most to an LLM application. These show the supplier's product is restricted to an approved set, pinned, watched for silent change, and evaluated on every route. The supplier-management process is the organisation's.

Unmapped. No obligation provides evidence for these. Recorded rather than padded: a crosswalk that maps everything is claiming more than a test catalog can show.

EntryNameWhy not
A.2.2 AI policy Organisational policy.
A.2.3 Alignment with other organisational policies Organisational policy.
A.2.4 Review of the AI policy Organisational policy.
A.3.2 AI roles and responsibilities Organisational roles.
A.3.3 Reporting of concerns Organisational process for people to raise concerns.
A.4.2 Resource documentation Documentation.
A.4.3 Data resources Documentation of resources.
A.4.4 Tooling resources Documentation of resources.
A.4.5 System and computing resources Documentation of resources.
A.4.6 Human resources People and competence.
A.5.2 Impact assessment process Impact assessment is an organisational process.
A.5.3 Documentation of impact assessments Documentation.
A.5.4 Impacts on individuals or groups Impact assessment; fairness measurement is out of scope per docs/NON-GOALS.md.
A.5.5 Societal impacts Impact assessment.
A.6.1.2 Objectives for responsible development Organisational objectives.
A.6.1.3 Processes for responsible design and development Process definition; its operation is evidenced under A.6.2.4 and A.6.2.5.
A.6.2.2 Requirements and specification Specification is the adopter's. The catalog's obligations are not a system's requirements and are not claimed as its specification.
A.6.2.3 Documentation of design and development Documentation.
A.6.2.7 Technical documentation Documentation.
A.7.2 Data for development and enhancement Training and development data are model-level and out of scope per docs/NON-GOALS.md.
A.7.3 Acquisition of data Acquisition process; ingestion controls on a retrieval corpus are under A.7.5.
A.7.4 Quality of data Data-quality requirements are set by the organisation. A.7.5 carries the nearest evidence; mapping AAC-0108 here as well would claim more than a provenance check shows.
A.7.6 Data preparation Training-data preparation is out of scope per docs/NON-GOALS.md.
A.8.2 System documentation and information for users Documentation.
A.8.3 External reporting Organisational reporting process.
A.8.4 Communication of incidents Organisational communication process.
A.8.5 Information for interested parties Organisational communication.
A.9.2 Processes for responsible use Organisational process for the organisation's own use of AI.
A.9.3 Objectives for responsible use Organisational objectives.
A.9.4 Intended use Concerns how the organisation uses the system. Scope-of-refusal testing (AAC-0003) shows the system stays inside a declared scope, which is related but not the same claim.
A.10.2 Allocation of responsibilities Contractual and organisational.
A.10.4 Customers Customer relationship.
Crosswalk — nist-ai-rmf

What an existing framework maps to

Informative. 29 entries, mapped to obligations across the archetypes that actually owe them; 55 more recorded below as unmapped, with the reason.

The external framework is organised by its own axis — threat, risk-management outcome, control or article; this catalog is organised by application shape. The mapping is many-to-many by construction, and that is the point — a single external item lands on several obligations across several archetypes, which is exactly what the framework on its own cannot tell you.

Crosswalks release out of band: a revision upstream must never force a version bump in the catalog. relation records how tight each mapping is, because claiming a test obligation is equivalent to a governance control is the fastest way to have a crosswalk dismissed.

EntryNameRelationObligations
GOVERN 1.6 Mechanisms to inventory AI systems evidence-for AAC-0094AAC-0107
AAC-0094 is the one case that names this outright: an inventory of models in use records intent unless the egress point refuses everything not on it. AAC-0107 makes the inventory precise to the digest. Neither supplies the organisational inventory itself, or its resourcing.
GOVERN 6.2 Contingency for failures in third-party data or AI systems evidence-for AAC-0009AAC-0091AAC-0098
The contingency process is organisational; these show that its technical half works. AAC-0009 injects provider failure, AAC-0091 takes a third-party guardrail down, and AAC-0098 ensures the fallback the contingency relies on was itself evaluated.
MAP 3.5 Human oversight processes defined, assessed, documented partial AAC-0078AAC-0043AAC-0056
Covers "assessed" only — each case tests that an oversight path works (approval before irreversible action, handoff to a human, confirmation before destructive tools). Defining and documenting the oversight process is the organisation's.
MEASURE 1.2 Appropriateness of metrics and effectiveness of controls regularly assessed partial AAC-0084AAC-0088AAC-0090
Two halves are evidenced: a model-graded metric is checked against human labels and re-checked when it changes (AAC-0084, AAC-0090), and a guardrail control's effectiveness is measured in both directions (AAC-0088). Error reports and impacts on affected communities are not.
MEASURE 2.1 Test sets, metrics and TEVV tools documented partial AAC-0001
AAC-0001 requires the test set to exist and be version-controlled. The coverage report (docs/REPORT.md) records which mechanism and which tool produced each verdict, which is closer to this item than any single obligation, but a report is not an obligation and is not mapped.
MEASURE 2.3 Performance measured for conditions similar to deployment evidence-for AAC-0001AAC-0007AAC-0017AAC-0098AAC-0014
"Similar to deployment" is where these differ from an offline score: latency under realistic concurrency (AAC-0007), every route the system can actually serve from (AAC-0098), and live sampled traffic (AAC-0014). The criteria themselves are the adopter's.
MEASURE 2.4 Functionality and behaviour monitored in production evidence-for AAC-0014AAC-0114AAC-0116AAC-0079AAC-0011
AAC-0114 and AAC-0116 exist because traffic-based monitoring alone reads a broken deployment as a quiet hour; AAC-0011 is the telemetry everything else in this row depends on.
MEASURE 2.5 Valid and reliable; limits of generalisability documented partial AAC-0001AAC-0010AAC-0013AAC-0018AAC-0098
Validity and reliability are evidenced: accuracy on a frozen set, its noise floor, regression against the incumbent, stability under paraphrase, and every reachable model. Documenting the limits of generalisability is not an obligation in the catalog.
MEASURE 2.6 Evaluated for safety; fails safely, including beyond knowledge limits evidence-for AAC-0005AAC-0091AAC-0055AAC-0077AAC-0056AAC-0072AAC-0031AAC-0019
"Fail safely beyond its knowledge limits" lands on two cases teams rarely connect to safety: abstention when the corpus cannot answer (AAC-0031) and a tested strategy at the context limit (AAC-0019). Whether residual risk is within tolerance is a threshold, and belongs to the adopter.
MEASURE 2.7 Security and resilience evaluated evidence-for AAC-0004AAC-0036AAC-0058AAC-0057AAC-0106AAC-0107AAC-0108AAC-0111AAC-0071
See crosswalks/owasp-llm.yaml for the threat-by-threat split of the same cases.
MEASURE 2.8 Transparency and accountability risks examined partial AAC-0080AAC-0011AAC-0060AAC-0100
Accountability through traceability only: any action, step or route choice can be reconstructed afterwards. Transparency towards users and affected people is not covered.
MEASURE 2.10 Privacy risk examined evidence-for AAC-0006AAC-0032AAC-0040AAC-0095AAC-0096AAC-0097AAC-0117
MEASURE 2.13 Effectiveness of TEVV metrics and processes evaluated evidence-for AAC-0084AAC-0085AAC-0086AAC-0087AAC-0090AAC-0010
The closest fit in the RMF for the A10 judge obligations: a model-graded metric is itself a TEVV instrument, and these cases validate it — human agreement, bias probes, stability, calibration, re-validation on change. AAC-0010 supplies the noise floor without which no metric's movement can be read.
MEASURE 3.1 Existing, unanticipated and emergent risks tracked from deployed performance evidence-for AAC-0114AAC-0115AAC-0014
MEASURE 3.3 End-user feedback and appeal integrated into evaluation metrics partial AAC-0115
AAC-0115 joins later outcomes, including explicit feedback, to the run that produced them and counts them beside rubric scores — the "integrated into evaluation metrics" half. The appeal process is not covered.
MEASURE 4.3 Performance improvements or declines identified from field data evidence-for AAC-0014AAC-0114AAC-0115
Field-data cases only. AAC-0013 also identifies declines, but on an offline set, so it sits under MEASURE 2.5 instead. Consultation with affected communities is not covered.
MANAGE 1.1 Determination whether the system achieves its purpose and should proceed evidence-for AAC-0001AAC-0013AAC-0102
The `gate: true` field on an obligation is the catalog's input to this go/no-go. These three are the release gates for quality, regression and cost; the determination itself, and the thresholds behind it, are the adopter's.
MANAGE 2.4 Supersede, disengage or deactivate systems performing inconsistently partial AAC-0012AAC-0077
AAC-0012 exercises reverting to the previous model-and-prompt pair, which is supersession tested rather than assumed. AAC-0077 stops a run, not a system. Assigning responsibility for the decision is not covered.
MANAGE 3.1 Third-party risks monitored and controls applied evidence-for AAC-0016AAC-0094AAC-0107AAC-0009
MANAGE 3.2 Pre-trained models monitored as part of regular monitoring evidence-for AAC-0016AAC-0082AAC-0012AAC-0107
The strongest Core fit in this file. A hosted model changes under you with nothing changed on your side; AAC-0016 and AAC-0082 re-run frozen inputs on a schedule to catch it, and AAC-0107 detects a substituted artifact at the release gate.
MANAGE 4.1 Post-deployment monitoring plans implemented partial AAC-0114AAC-0115AAC-0116AAC-0101AAC-0012
Covers monitoring, capture of user input (AAC-0115), recovery (AAC-0012) and change management for routing (AAC-0101). Appeal and override, decommissioning and incident response are not covered.
MANAGE 4.3 Incidents and errors communicated; tracking and recovery processes followed partial AAC-0080AAC-0079
Shaky and deliberately marked partial: AAC-0079 makes failures reach a human and AAC-0080 makes any action explainable weeks later, which an incident process needs. Communicating incidents to affected parties is organisational and not covered.
AI 600-1 §2.2 Confabulation tests-for AAC-0029AAC-0030AAC-0031AAC-0024AAC-0110AAC-0112
Includes two forms of confabulation the profile's examples do not separate: a claim about the system's own actions that its tool results do not support (AAC-0110), and a reply that reports a completed outcome for work the system has only promised (AAC-0112).
AI 600-1 §2.3 Dangerous, Violent, or Hateful Content partial AAC-0005AAC-0088AAC-0091AAC-0092
The catalog obliges testing against the adopter's declared policy categories and does not name categories itself. Evidence for this risk only where the adopter's policy includes it.
AI 600-1 §2.4 Data Privacy partial AAC-0006AAC-0032AAC-0040AAC-0095AAC-0096AAC-0117
Application-level leakage paths only. Training-data memorisation and inference of sensitive attributes are model-level and out of scope per docs/NON-GOALS.md.
AI 600-1 §2.7 Human-AI Configuration partial AAC-0078AAC-0043AAC-0056
Shaky: these test that a human is in the loop where the design says one is. Automation bias, over-reliance, anthropomorphisation and emotional entanglement — the substance of this risk — are not covered.
AI 600-1 §2.9 Information Security partial AAC-0004AAC-0036AAC-0058AAC-0106AAC-0057AAC-0071AAC-0108AAC-0111
The profile names two halves. The attack surface of the GAI system itself (direct and indirect prompt injection, poisoning) is covered; the lowered barrier to offensive cyber capability is a model-capability question and is not.
AI 600-1 §2.11 Obscene, Degrading, and/or Abusive Content partial AAC-0005AAC-0088AAC-0091AAC-0092
As §2.3: evidence only where the adopter's declared policy categories include it.
AI 600-1 §2.12 Value Chain and Component Integration evidence-for AAC-0107AAC-0094AAC-0012AAC-0016AAC-0098
Covers the deployed value chain — which components serve, whether they are the ones evaluated, and whether a provider-side change is noticed. Vetting procured training data is out of scope.

Unmapped. No obligation provides evidence for these. Recorded rather than padded: a crosswalk that maps everything is claiming more than a test catalog can show.

EntryNameWhy not
GOVERN 1.1 Legal and regulatory requirements understood and documented Organisational process; no system property evidences it.
GOVERN 1.2 Trustworthy characteristics integrated into policy Organisational policy.
GOVERN 1.3 Level of risk management set by risk tolerance Risk tolerance is the adopter's threshold, which the catalog never sets.
GOVERN 1.4 Risk management process established through transparent policy Organisational policy.
GOVERN 1.5 Ongoing monitoring and periodic review of the risk management process Reviews the process, not the system; system monitoring is under MEASURE 2.4.
GOVERN 1.7 Decommissioning and phasing out safely No obligation covers retiring a system.
GOVERN 2.1 Roles, responsibilities and lines of communication documented Organisational roles.
GOVERN 2.2 Personnel and partners trained Training of people.
GOVERN 2.3 Executive leadership responsible for AI risk decisions Organisational accountability.
GOVERN 3.1 Decisions informed by a diverse team Team composition.
GOVERN 3.2 Roles for human-AI configurations and oversight defined Role definition; the tested oversight paths are under MAP 3.5.
GOVERN 4.1 Safety-first mindset fostered Organisational culture.
GOVERN 4.2 Teams document and communicate risks and impacts Organisational documentation practice.
GOVERN 4.3 Practices enable testing, incident identification and information sharing An organisational practice. An adopter's series of coverage reports is evidence for it as a whole; no single obligation is.
GOVERN 5.1 Feedback from those external to the team collected and integrated Organisational engagement process.
GOVERN 5.2 Adjudicated feedback incorporated into design Organisational engagement process.
GOVERN 6.1 Policies for third-party risks, including IP Organisational policy; the technical controls are under MANAGE 3.1.
MAP 1.1 Intended purposes, context and settings understood and documented Context-setting documentation.
MAP 1.2 Interdisciplinary actors and competencies Team composition.
MAP 1.3 Mission and goals for AI documented Organisational documentation.
MAP 1.4 Business value or context of use defined Business decision.
MAP 1.5 Organisational risk tolerances determined Risk tolerance is the adopter's threshold, which the catalog never sets.
MAP 1.6 System requirements elicited and understood Requirements elicitation is a process; obligations are not requirements for a specific system.
MAP 2.1 Tasks and methods defined (classifier, generative model, recommender) The catalog's archetype classification (taxonomy/archetypes.yaml) is a natural artifact for this, but it is a classification, not an obligation, so nothing maps.
MAP 2.2 Knowledge limits and use of output by humans documented Documentation; the tested behaviour at knowledge limits is under MEASURE 2.6.
MAP 2.3 Scientific integrity and TEVV considerations identified Planning-stage documentation; the measurement-side evidence is under MEASURE 2.13.
MAP 3.1 Potential benefits examined Benefit analysis.
MAP 3.2 Potential costs of errors examined against risk tolerance "Costs" here are harms from errors, judged against risk tolerance. The catalog's cost obligations measure spend, which is a different thing.
MAP 3.3 Targeted application scope specified Scope decision and documentation.
MAP 3.4 Operator and practitioner proficiency processes Training and certification of people.
MAP 4.1 Technology and legal risks of components mapped, including third-party Risk-mapping process.
MAP 4.2 Internal risk controls for components identified and documented Documentation of controls; the controls' operation is under MANAGE 3.1.
MAP 5.1 Likelihood and magnitude of impacts identified Impact assessment.
MAP 5.2 Engagement with relevant AI actors on impacts Organisational engagement process.
MEASURE 1.1 Metrics selected; unmeasured risks documented Selection is a decision. The coverage report's not-covered and accepted-risk rows are the natural record of what is not measured, but that is report format, not an obligation.
MEASURE 1.3 Independent assessors involved Who assesses, not what is true of the system.
MEASURE 2.2 Evaluations involving human subjects meet requirements Research-ethics requirement on the evaluation, not the system.
MEASURE 2.9 Model explained and validated; output interpreted in context Model explainability is model-level (docs/NON-GOALS.md). Citation support (AAC-0030) helps a reader interpret an answer but is not an explanation.
MEASURE 2.11 Fairness and bias evaluated Out of scope per docs/NON-GOALS.md: it needs domain-specific protected attributes and legal context. Per-slice accuracy (AAC-0017) is not a fairness measure and is not claimed as one.
MEASURE 2.12 Environmental impact and sustainability assessed No obligation measures energy or emissions. The cost obligations measure tokens and currency, which are not a proxy the catalog will claim.
MEASURE 3.2 Risk tracking where measurement is not yet possible Process for unmeasurable risks; by construction no test obligation reaches it.
MEASURE 4.1 Measurement approaches informed by domain experts and end users Consultation process.
MEASURE 4.2 Trustworthiness results validated with domain experts Consultation process.
MANAGE 1.2 Risk treatment prioritised Prioritisation decision.
MANAGE 1.3 Responses to high-priority risks planned Planning. Where the response is acceptance, an accepted-risk declaration in the coverage report records it, but that is not an obligation.
MANAGE 1.4 Negative residual risks documented Documentation of residual risk.
MANAGE 2.1 Resources and non-AI alternatives considered Resourcing decision.
MANAGE 2.2 Mechanisms to sustain the value of deployed systems Too general for a test obligation to evidence specifically; the drift and monitoring cases are mapped under MEASURE 2.4 and MANAGE 3.2.
MANAGE 2.3 Response to previously unknown risks Incident procedure.
MANAGE 4.2 Continual improvement integrated into updates Improvement process.
AI 600-1 §2.1 CBRN Information or Capabilities A model-capability uplift question, evaluated on the model. An application whose policy categories include it gets AAC-0005 as under §2.3.
AI 600-1 §2.5 Environmental Impacts No obligation measures energy or emissions.
AI 600-1 §2.6 Harmful Bias and Homogenization Fairness is out of scope per docs/NON-GOALS.md. Judge bias (AAC-0085) is bias of an evaluator, a different sense of the word.
AI 600-1 §2.8 Information Integrity Provenance and authentication of generated content is not an obligation in the catalog — a candidate gap, shared with EU AI Act Art. 50(2). Misinformation from confabulation is mapped under §2.2.
AI 600-1 §2.10 Intellectual Property Training-data and legal question; no application-level test obligation.
Crosswalk — owasp-llm-top-10

What an existing framework maps to

Informative. 10 entries, mapped to obligations across the archetypes that actually owe them.

The external framework is organised by its own axis — threat, risk-management outcome, control or article; this catalog is organised by application shape. The mapping is many-to-many by construction, and that is the point — a single external item lands on several obligations across several archetypes, which is exactly what the framework on its own cannot tell you.

Crosswalks release out of band: a revision upstream must never force a version bump in the catalog. relation records how tight each mapping is, because claiming a test obligation is equivalent to a governance control is the fastest way to have a crosswalk dismissed.

EntryNameRelationObligations
LLM01 Prompt Injection tests-for AAC-0004AAC-0036AAC-0058AAC-0106
Three distinct vectors, deliberately separate cases. AAC-0004 is the direct channel, AAC-0036 is injection carried in retrieved documents, and AAC-0058 is injection arriving through tool output. Systems that test the front door thoroughly routinely miss the latter two, which is precisely what the archetype axis makes visible.
LLM02 Sensitive Information Disclosure tests-for AAC-0006AAC-0032AAC-0040AAC-0095AAC-0096
Includes two disclosure paths that are not prompt-related at all: AAC-0032 permission-scoped retrieval, and AAC-0096 cache keys that omit a permission dimension. Both leak one principal's data to another without any model behaviour being at fault.
LLM03 Supply Chain partial AAC-0012AAC-0107AAC-0094
AAC-0107 was written to close this gap, found by the first pass of this crosswalk against draft 0.1. Coverage remains partial: the catalog addresses artifact provenance and version pinning but not the security of the training or fine-tuning pipeline, which is out of scope as a model-level rather than application-level concern.
LLM04 Data and Model Poisoning partial AAC-0108AAC-0041AAC-0034
AAC-0108 was written to close this gap. Retrieval-corpus and long-term- memory poisoning are covered; training-data poisoning is deliberately out of scope per docs/NON-GOALS.md.
LLM05 Improper Output Handling tests-for AAC-0022AAC-0027AAC-0068AAC-0071AAC-0072
The strongest-covered entry in the list, because A8 can verify its output by executing it. AAC-0071 treats generated artifacts as untrusted contributor code, which is what they are.
LLM06 Excessive Agency tests-for AAC-0056AAC-0057AAC-0078AAC-0081AAC-0055
AAC-0057 carries the load: authorisation enforced server-side by the tool rather than by prompt instruction. Most excessive-agency incidents are not model failures — they are permission failures that a cooperative model happened to be masking.
LLM07 System Prompt Leakage tests-for AAC-0106
Had no obligation at all until this crosswalk found the hole. AAC-0106 deliberately carries two halves: extraction is tested, and separately no capability may depend on the prompt staying hidden. Testing extraction alone measures the wrong thing.
LLM08 Vector and Embedding Weaknesses tests-for AAC-0032AAC-0034AAC-0036AAC-0108AAC-0028
AAC-0028 evaluates retrieval independently of generation. Most reported failures attributed to this entry are retrieval-quality failures wearing a generation costume, and they cannot be diagnosed without separating the two.
LLM09 Misinformation tests-for AAC-0029AAC-0030AAC-0031AAC-0084AAC-0001
AAC-0031 abstention is the one teams omit: a dataset curated from answerable questions never contains the unanswerable ones, so the obligation has to be built deliberately or it is silently untested.
LLM10 Unbounded Consumption tests-for AAC-0055AAC-0064AAC-0077AAC-0093AAC-0102
OWASP frames this as availability and denial of service. The catalog also carries it as an economic obligation — AAC-0102 gates cost regression at release — which is the framing the security-oriented entry does not supply. See patterns/cost.yaml for the failure shapes.
Section 6 — How to build it

From the obligation to the thing you actually write

Informative. Every obligation is reachable several ways, at very different cost, effort and lock-in — so each is listed with the trade-off it makes rather than a recommendation.

Approach answers who provides the machinery, and is orthogonal to mechanism and stage. Two of the six are worth knowing about before you read the rest: gateway is the only approach that can prevent rather than detect, and the only one that covers calls your application code forgot to route through the wrapper. in-house is chronically under-considered and frequently strongest — several entries below are a dozen lines and beat anything purchasable, because they assert what you meant rather than what a product happens to measure.

Verify every product claim before relying on it. These describe the kind of capability a category of product typically offered as of the date on each file. Names, features and pricing change without notice, and some entries are already wrong.

Section 7 — Patterns

Why these obligations exist

Informative. 56 known failure shapes, each terminating in the obligations that would surface it — or declaring itself a gap.

A pattern is a diagnosis, not an obligation. Keeping the two apart is what lets a case statement stay short: the pattern carries the war story so the obligation does not have to. Nothing in this section is required for conformance.

The rule that earns the section: a pattern must terminate in a case identifier or state what is missing. An uncaught pattern is a hole in the catalog, not an omission in the pattern — so this doubles as coverage validation, and the linter reports every one of them rather than passing silently.

IDPatternSymptomCaught by
AACP-0013 Prompt-only guardrails
agent · A6 A7 A9
Limits hold in testing and fail under unusual inputs. Post-incident review finds the constraint was documented and never enforced. AAC-0055AAC-0056AAC-0057
AACP-0014 Tool-error ping-pong
agent · A6 A7
Runs that terminate on step budget rather than completion, with a trace full of near-identical calls. AAC-0053AAC-0055AAC-0054
AACP-0015 Instructions arriving through tool output
agent · A3 A6 A7 A9
The agent takes an action nobody requested, traceable to content returned by a tool or retrieved document rather than to user input. AAC-0058AAC-0036AAC-0004
AACP-0016 Context lost at handoff
agent · A7
A sub-agent produces confident, well-formed work that answers a subtly different question. No error appears anywhere in the trace. AAC-0062AAC-0061
AACP-0017 Supervisor ping-pong
agent · A7
Cost and latency spike on a minority of requests; the trace shows the same task handed between two agents repeatedly. AAC-0063AAC-0064
AACP-0018 Trajectory judged only by its answer
agent · A6 A7
Quality metrics are strong while cost and latency drift upward with no identified cause. AAC-0054AAC-0060AAC-0100
AACP-0001 Retry amplification
cost · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
Spend rises while request volume is flat. Per-call cost dashboards look healthy and the invoice does not. AAC-0008AAC-0102AAC-0104
AACP-0002 Context accretion
cost · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
Input tokens per task climb steadily across releases with no single change responsible. AAC-0103
AACP-0003 Conversation tail cost
cost · A4 A6 A7
Long sessions cost far more than session count suggests. Cost per turn rises with turn index. AAC-0042AAC-0103
AACP-0004 Tool-output dumping
cost · A6 A7
Agent runs that touch data-returning tools cost many times those that do not, out of proportion to their step count. AAC-0105AAC-0054
AACP-0005 Frontier-model default
cost · A1 A2 A5 A6 A7 A10
The most capable available model serves every call, including classification, routing and formatting steps. AAC-0089AAC-0102
AACP-0006 Cancellation is not propagated
cost · A1 A3 A4 A6 A8
Spend attributable to sessions the user abandoned. Generation continues after the client has disconnected. gap — no case
AACP-0007 Volatile prompt prefix defeats caching
cost · A1 A2 A3 A4 A5 A6 A7 A8 A10
Prompt-cache hit rate near zero despite a large, stable system prompt. Cost per call does not fall as expected after enabling caching. gap — no case
AACP-0008 Unbounded fan-out under a supervisor
cost · A7
Occasional runs cost orders of magnitude more than the median with no obvious difference in the request. AAC-0064AAC-0067
AACP-0009 Ungoverned direct provider access
cost · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
Provider invoices exceed the total of all gateway-attributed spend. Inventory of models in use cannot be reconciled with what is billed. AAC-0094AAC-0104
AACP-0010 Judge cost exceeds the system it judges
cost · A10
Evaluation spend approaches or exceeds serving spend. Online scoring is the largest line item. AAC-0014AAC-0104
AACP-0011 Unevaluated cheap route
cost · A1 A2 A3 A4 A5 A6 A7 A8 A10
Quality complaints that cannot be reproduced, clustering around periods of high load or provider degradation. AAC-0098AAC-0099AAC-0100
AACP-0012 Cost per call, not per task
cost · A2 A3 A4 A5 A6 A7 A8 A9
Unit cost falls release over release while total spend rises. Optimisation work shows gains that never appear on the invoice. AAC-0008AAC-0050
AACP-0019 Retry after the far end said no
operations · A6 A7 A9
The same write appears two or three times in one trace, each refused, and the reply says it could not be done — or, worse, that it was. AAC-0053AAC-0047
AACP-0020 Retry storm on a failing dependency
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
One dependency degrades and the system's traffic to it multiplies, so the dependency degrades further and the bill rises with it. AAC-0053AAC-0009
AACP-0021 Write without a read
operations · A6 A7 A9
Irreversible actions taken on records the system never looked at in that conversation — cancellations of orders it had not fetched, refunds against totals it had not read. AAC-0113AAC-0056
AACP-0022 Invented argument
operations · A6 A7 A9
Tool calls that fail with not-found, or succeed on the wrong record, for identifiers the user never gave. AAC-0052
AACP-0023 A tool that is not offered
operations · A6 A7 A9
Unknown-tool errors in the trace, often after a model or prompt change, for tools with plausible names that do not exist. AAC-0051
AACP-0024 Arguments that fail their schema, again
operations · A6 A7 A9
Validation errors on tool arguments that repeat across units of work for the same tool, rising after a tool's schema changes. AAC-0052
AACP-0025 An answer built on a truncated result
operations · A2 A6 A7 A9
Confident answers about "all" of something — every order, every row — that miss the part a result bound cut off. AAC-0105
AACP-0026 A dead tool on the surface
operations · A6 A7 A9
Tool definitions that cost context on every call and are never used, or used only by mistake. AAC-0103AAC-0051
AACP-0027 Tool latency drift
operations · A2 A6 A7 A9
End-to-end latency climbs with no change in the system, and the model calls are as fast as ever. AAC-0007
AACP-0028 Reply contradicts a tool result
operations · A2 A4 A6 A7
Users told something the system's own lookup said was otherwise: an order reported shipped that the store said was pending. AAC-0110AAC-0113
AACP-0029 A claimed action without its effect
operations · A6 A7 A9
Users told something was cancelled, refunded or booked, and nothing happened. Found when they come back. AAC-0110
AACP-0030 An identifier nobody gave
operations · A2 A4 A6 A7
Replies that cite order numbers, reference codes or amounts that exist nowhere in the system. AAC-0110AAC-0024
AACP-0031 Internal text in a reply
operations · A1 A2 A4 A6 A7
Users shown tool names, JSON, stack traces or fragments of the instructions. AAC-0106AAC-0002
AACP-0032 Asks for what it was already given
operations · A4 A6 A7
"Could you give me your order number?" in reply to a message containing one. Users repeat themselves and leave. AAC-0038
AACP-0033 Personal data in a reply
operations · A1 A2 A4 A6 A7
Card numbers, phone numbers or another person's details shown to a user. AAC-0006
AACP-0034 A completion with nothing in it
operations · A1 A4 A6 A7
Units of work counted as completed whose reply is empty, a fragment, or a promise to look. AAC-0112
AACP-0035 Apology without progress
operations · A4 A6 A7
Conversations of several turns in which each reply apologises and none moves anything. AAC-0037
AACP-0036 Ended by a limit, not by an answer
operations · A6 A7 A9
A share of units of work end on a step budget, a cost ceiling or a loop detector, and the reply is a hand-off sentence. AAC-0055AAC-0077
AACP-0037 A longer path than the task needs
operations · A6 A7 A9
Tool calls per unit of work rise for the same kinds of request, with latency and cost rising behind them. AAC-0054
AACP-0038 Unreadable model output
operations · A1 A2 A3 A5 A6 A7 A8 A9 A10
A rise in failed units of work after a model or prompt change, with the provider reporting success. AAC-0002
AACP-0039 Answered over a failed tool
operations · A2 A6 A7 A9
Confident replies in units of work where the lookup failed. AAC-0053
AACP-0040 The user repeats themselves
operations · A4 A6 A7
The same message, or nearly, twice in one conversation. AAC-0039AAC-0037
AACP-0041 Back within a day
operations · A4 A6 A7
Resolution rates that look good while contact volume per user does not fall. AAC-0115AAC-0042
AACP-0042 A person asked for right after an answer
operations · A4 A6 A7
Escalations whose conversation shows a completed answer immediately before the request for a person. AAC-0115
AACP-0043 Turns to resolution creeping up
operations · A4 A6 A7
Conversations get longer for the same requests, and late turns cost more than early ones. AAC-0042
AACP-0044 Escalations nobody needed
operations · A4 A6 A7
People on the desk close hand-offs the system could have handled, and the share is concentrated in one rule. AAC-0020AAC-0043
AACP-0045 Waits that lapse with nobody
operations · A6 A7 A9
Approvals expire and hand-offs time out, and the users behind them are told nothing more. AAC-0043AAC-0078
AACP-0046 One rule refusing too much
operations · A1 A4 A6 A7
A sudden rise in refusals, all under one rule, after a change to rules or prompt. Users experience the product as broken. AAC-0020AAC-0088
AACP-0047 Authority reached for
operations · A6 A7 A9
The far end refuses calls for want of authority — no approval, a scope the session does not hold — rather than for the state of the record. AAC-0057AAC-0111
AACP-0048 A promise nothing is keeping
operations · A4 A6 A7
"I'll get back to you" in replies, and nobody does. AAC-0112
AACP-0049 Healthy status codes, broken behaviour
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
Error rate flat, latency normal, and users complaining. The system is up and not doing its job. AAC-0114
AACP-0050 A quiet deployment
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
A drop in traffic read as a quiet hour. The deployment was broken and nobody could reach it. AAC-0116AAC-0079
AACP-0051 Latency tail over budget
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
Average latency fine, and a tenth of users waiting half a minute. AAC-0007
AACP-0052 Provider degradation
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
A burst of failed units of work with the same cause, from the model provider or the gateway. AAC-0009
AACP-0053 Spend outrunning its budget
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
The month's budget is gone by the middle of it, or one unit of work costs what a hundred usually do. AAC-0008AAC-0093AAC-0077
AACP-0054 The served model is not the pinned one
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
Behaviour shifts with no release, and the configuration names the same model it always did. AAC-0012AAC-0094
AACP-0055 A release that moved the rates
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
A deploy, then a slow change in escalations, refusals or cost that nobody connects to it. AAC-0013AAC-0101
AACP-0056 Telemetry that silently drops
operations · A1 A2 A3 A4 A5 A6 A7 A8 A9 A10
A trace store with fewer units of work than the system served, and online scores computed over the ones that happened to arrive. AAC-0011AAC-0079
Section 8 — Using it

Turning the catalog into a test plan

  • Classify. Decompose the system into archetypes — most are two or three. Write the classification down; it is the assumption everything rests on and the thing most likely to be wrong.
  • Take the union. Core, plus the deltas for each archetype present, plus seam cases for every boundary between them.
  • Realize each obligation. Name the mechanism, the stage, the tool and whether it gates. An obligation with no named tool is not planned, it is aspirational.
  • Register what you will not do. Any MUST you skip becomes an explicit accepted risk with an owner and a review date. An honest gap beats a green dashboard — and under a management-system audit, an identified, owned, dated gap is conformant while an undiscovered one is a finding.
  • Bind identifiers into code. Put the case identifier in the test name and in a span attribute. Once coverage is queryable from your traces, this stops being a document and becomes a control.
  • Feed incidents back. Every production failure becomes a new frozen case tagged to the archetype that produced it.