
A vendor sends you an AI Bill of Materials. It is a well-formed JSON file, a few hundred lines, and it validates against the schema. It names the base model, the fine-tuning dataset, the inference runtime, three MCP servers, and a guardrail policy pack. Everything your procurement checklist asked for is in there.
Now answer a harder question: is any of it true?
That is not a rhetorical jab at vendors. It is the practical problem with a document whose entire content is a set of claims by the party being assessed. An AI-BOM is produced by the supplier, about the supplier's own system, and handed to you. Unless something in it can be independently re-derived, you have not acquired evidence. You have acquired a very tidy assertion.
An AI Bill of Materials is a structured, machine-readable inventory of everything an AI system is built from — models, datasets, agents and tools, prompts and guardrails, and the runtime it executes in — together with the provenance and integrity evidence that supports each entry. It is the AI equivalent of a software bill of materials, extended to cover the artifacts that make an AI system behave the way it does.
The OWASP AIBOM project's Foundations Guide frames it in four parts: a header declaring what the document covers and how complete it claims to be, a component list, the directed data flows between those components, and the evidence and provenance references attached to the claims. Every AI-BOM has all four.
The word doing the most work in that sentence is evidence. The same guide draws the line plainly: "A component asserted without any evidence reference is a bare assertion; the same component carrying a hash, a signature, and a link to its model card is a claim a consumer can independently verify."
That distinction is the whole subject of this post.
Because in 2026 the guidance stopped being aspirational and started naming fields. Three documents did most of that work, and they are worth separating because they say different things.
In May 2026, seven national cyber agencies — the US Cybersecurity and Infrastructure Security Agency, Germany's BSI, France's ANSSI, the UK's NCSC, Canada's CSE, Italy's ACN and Japan's NCO, in collaboration with the European Commission — jointly published Software Bill of Materials for AI — Minimum Elements. It is the first G7 guidance on the subject. Its framing is deliberately unglamorous: "AI systems are also software systems. Therefore, SBOMs still remain valid for AI systems. The minimum elements in an SBOM for AI are in addition to the general SBOM minimum elements." It is also explicit that the elements "are not mandatory; do not create requirements, standards, or legislation; and are open to further refinements."
In July 2026, CISA and a long list of international partners published the 2026 Minimum Elements for a Software Bill of Materials, updating the 2021 NTIA minimum elements from seven data fields to seventeen. Two of the new ones matter here more than the rest: Component Hash Value and Component Hash Algorithm. CISA's stated rationale for the second is worth keeping: "Identifying the algorithm used to generate the hash value is critical for making component hashes useful." The same document declines to add AI-specific fields, pointing at the G7 guidance instead.
And the NIST AI Risk Management Framework has asked for the underlying capability since January 2023, in a single subcategory — GOVERN 1.6: "Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities." Its MAP 4.2 goes on to ask that "Internal risk controls for components of the AI system, including third-party AI technologies, are identified and documented," which is the subcategory an AI-BOM most directly serves.
One scoping note, because getting it wrong is common. The EU AI Act's Annex IV technical documentation — which does require a component-and-dependency description, third-party pre-trained model disclosure, dataset provenance, and signed test reports — is not yet in force for high-risk systems. Following the Digital Omnibus on AI — Regulation (EU) 2026/1744 of 8 July 2026 — the European Commission's own implementation timeline puts Annex III high-risk obligations at 2 December 2027 and Annex I product-embedded high-risk at 2 August 2028. What is in force today is most of the Act: the majority of its rules applied from 2 August 2026 and enforcement has started, on top of the general-purpose AI obligations that began 2 August 2025, whose Annex XI and XII documentation duties already require "the type and provenance of data and curation methodologies" from model providers. Cite the part that applies.
The G7 minimum elements are the most useful answer available from a government source, because they are organized as seven clusters of named fields rather than as a principle. Paraphrased at the cluster level:
Metadata — who authored the document, in what format and version, with which tool, at what time, and critically, the SBOM author signature. System-level properties — the system's name, version, producer, components, and its data flows and data usage. Models — name, identifier, version, producer, license, training properties, and the model hash value together with the model hash algorithm. Dataset properties — name, content, identifier, license, sensitivity, statistical properties, dataset hash and dataset provenance. Infrastructure — the software and hardware it runs on, which for a self-hosted model is a materially larger surface than for a hosted API. Security properties — controls, compliance posture, vulnerability references. Key performance indicators — security metrics and operational performance.
For the file itself there are two mature formats, and neither is AI-specific bolted on afterwards. CycloneDX reached version 1.7 in October 2025 and is standardized as ECMA-424, second edition, December 2025; machine-learning support has been in the schema since 1.5, in the form of a modelCard object on a component of type machine-learning-model, carrying modelParameters (including datasets, architectureFamily, modelArchitecture), quantitativeAnalysis and considerations. SPDX 3.0 (current release 3.0.1) introduced a first-class ai profile and a dataset profile; its AIPackage class carries hyperparameter, informationAboutTraining, limitation, metric, safetyRiskAssessment and typeOfModel, and its DatasetPackage carries dataCollectionProcess, datasetNoise, knownBias and hasSensitivePersonalInformation.
Both formats also carry an integrity slot — verifiedUsing in SPDX, hashes in CycloneDX. That slot is where a document becomes testable.
A caution on maturity, since the space moves fast: the OWASP AIBOM project is listed as an OWASP Incubator project, and its Foundations Guide — v1.0, September 2026 — targets CycloneDX 2.0, which was announced in August 2026 with release "expected in fall 2026, with Ecma ratification in December" and has not shipped as of this writing; cyclonedx.org still lists 1.7 as the current version. Treat the guide as a well-argued direction of travel, not a settled standard.

Consider a mid-market insurance platform. The example below is a synthetic composite — an illustration assembled from patterns, not a real company or client.
The platform's claims-triage assistant — an agentic system with tool access, not a chatbot — has a genuinely good AI-BOM. It was generated inside the training pipeline, it validates, it lists the base model with a SHA-256 digest, the fine-tuning dataset with a digest, four tools the agent may call, and the guardrail policy pack version. Procurement is satisfied. The security team files it.
Four things then happen over eleven weeks, none of which is a breach, and none of which the AI-BOM reflects.
The model provider updates the alias the inference code points at, so the weights actually serving traffic no longer match the digest in the document. A platform engineer adds a fifth tool — an internal customer-lookup API — through an MCP server the agent discovers at runtime, requiring no application redeployment. That tool call runs under a non-human identity nobody onboarded. The guardrail policy pack is edited to reduce false positives on a legitimate claim category. And the fine-tuning dataset is refreshed with three months of new claims, including a subset that was never scrubbed for policyholder identifiers — the kind of path AI data leakage testing exists to find, and which no inventory field would have flagged.
Every one of those changes is a regeneration trigger in the OWASP guide's own list, which singles out the first as the dangerous class: "Supplier-side changes deserve particular attention, because they are the trigger class that fires without any local release."
Nobody lied. The document was accurate the day it was produced. It is now, in the guide's blunt phrasing, "worse than none, because it invites confidence it cannot support."
Five tests, each of which produces evidence rather than a reassurance. Four of the five can be run against a deployed system; the fifth needs visibility into the pipeline that emits the document.
One: re-derive the model digest. Hash the model artifact actually loaded in the target environment and compare it to the Model hash value in the document, using the declared algorithm. This is what the G7 elements exist to enable — the hash algorithm field is there, in their words, "to allow validation of the integrity of the target component." Where a model is pulled from a public hub, the same check runs upstream: a repository commit and per-file digest pin the claim to something re-derivable. MITRE ATLAS carries this as mitigation AML.M0014, "Verify AI Artifacts": "Verify the cryptographic checksum of all AI artifacts to verify that the file was not modified by an attacker." ATLAS also names the reason the check has to repeat, in technique AML.T0109, "AI Supply Chain Rug Pull": adversaries "may publish legitimate AI components or software, gain user adoption, then push an update with a malicious variant, leading to AI Supply Chain Compromise." Note that ATLAS now ships monthly content updates, so date-stamp the version you worked from — the release current at the time of writing is content version 2026.09, dated 2026-09-15.
Two: verify the document's own signature and provenance chain. An unsigned AI-BOM tells you nothing about who produced it or whether it was edited in transit, which is why SBOM author signature is a G7 metadata element and SBOM Author Signature is a CISA 2026 minimum element. Where the supplier uses in-toto attestations, the BOM becomes a signed predicate bound to a specific artifact digest, and both CycloneDX and SPDX are vetted predicate types. Sigstore adds transparency-log inclusion, so a tester can establish when and under which identity a document was signed rather than only that a signature validates. For model weights specifically, OpenSSF Model Signing hashes every file in the bundle and signs the manifest — which is also the honest answer to a real gap, since both the CISA and G7 hash definitions describe hashing an executable component artifact, and a sharded checkpoint or a LoRA adapter is not an executable. That reading is ours, not the standards'.
Three: enumerate what is running and diff it against what is declared. This is the test most AI-BOM programs skip. CISA's 2026 document puts accuracy, coverage and completeness outside the minimum elements' own scope, but points at exactly this method to get them: organizations "may use sources such as open source software repositories or binary analysis tools to detect components not listed in the SBOM." For an agentic system the equivalent is observing the tool-invocation and egress surface under load and comparing it to the declared services and their trust boundaries. CycloneDX already models the gap — its formulation object distinguishes declared from observed, and its lifecycle phases include discovery as a first-class alternative to design.
Four: read the evidence metadata, not just the values. CycloneDX 1.7 records, for each method supporting an identity claim, the technique used to establish it — the enum includes hash-comparison, binary-analysis, instrumentation, dynamic-analysis, attestation and, at the weak end, filename — each with a confidence score from 0 to 1. A component identified by filename at confidence 0.3 and one identified by hash comparison at 1.0 are both "in the BOM." They are not both evidence. A tester who reads only concludedValue has thrown away the most useful field in the schema.
Five: test the freshness claim, not just the content. Pick the regeneration triggers that apply to the system — a supplier alias change, an adapter swap, a dataset refresh, a guardrail edit — and establish whether each one actually causes the document to be regenerated. The measurable property is the lag between a trigger event and the corresponding update. A program that cannot answer this has a point-in-time artifact describing a system that moves — the same reason model drift is a security concern and not only a quality one, and the same reason continuous testing beats an annual snapshot on systems that change without releases.

This is the question that decides how much an AI-BOM is worth to you, and almost nothing written about the format addresses it.
Sort every field into two piles. In the first: hashes, signatures, digests, version pins, repository commits, tool definition hashes, license identifiers, declared endpoints. These are checkable — a tester can re-derive them independently and get a yes or a no.
In the second: dataCollectionProcess, knownBias, intendedUse, dataPreprocessing, ethicalConsiderations, the G7 Dataset provenance narrative describing web crawling and labeling steps. These are unfalsifiable from outside. They may be entirely true. They may be carefully worded. No test you can run against the deployed system distinguishes the two cases, because the evidence for them lives in a training pipeline you cannot see.
That asymmetry is not a flaw in the standards — it reflects what is genuinely verifiable downstream. But it should change how you read a document and what you ask for. When a supplier's AI-BOM is rich in narrative provenance and thin on digests, you are holding a well-written account rather than an inventory you can audit. The G7 guidance anticipates the honest version of this: where the author cannot access the component artifact, "the SBOM author should indicate the value is unknown." An explicit unknown is more useful than a confident paragraph.
There is a second lever, and it is the header. The OWASP guide requires a completeness claim, and explains exactly why it carries weight: "If an AIBOM does not list a particular dataset and does not claim to be complete, the absence proves nothing… If the AIBOM does claim to be complete for the declared scope, however, the absence becomes meaningful: it is now a producer assertion that no such dataset participates within scope." A completeness claim converts silence into a statement someone is accountable for. Ask for it.

An AI-BOM earns its place when it stops being a deliverable and becomes an input to a control. Three shifts do that.
It is generated by the pipeline, not written by a human — emitted on a successful build, so the document and the artifact share an origin. It is enforced at a gate, so a declared digest becomes an admission decision: the Sigstore Model Validation Operator is one published example of the mechanism, where verification runs in an init container and, if it fails, the workload pod does not start. And it is re-validated on triggers rather than on a calendar, with supplier-side changes treated as the priority class because they fire without any local release.
This is the same argument ioSENTRIX makes about every AI control, and the same reason we treat proving an autonomous action is genuine as the harder problem than controlling access to it. Governance asks whether the artifact exists. Assurance re-derives what it claims. An AI-BOM is an unusually good candidate for adversarial testing precisely because so much of it is arithmetic: a hash either matches or it does not, a signature either validates or it does not, an undeclared egress destination either appears in the traffic or it does not.
The OWASP guide, to its credit, refuses to oversell its own artifact: "It is not a verdict. It is not a statement that a system is compliant with any particular regulation… It is not a trust score, an assurance level, or a quality rating." And, more sharply: "An AIBOM that no one consumes is an engineering byproduct wearing a governance label."
Independent adversarial testing is what turns it into something else. That position produces the evidence whoever performs it; the point is only that the party making the claim should not be the party checking it.
Is an AI-BOM the same thing as an SBOM? No, but it builds on one. The G7 guidance is explicit that AI systems are software systems, so ordinary SBOM elements still apply, and the AI-specific elements are in addition to them. In practice an AI-BOM adds models, datasets, agents and tools, prompts and guardrails to a conventional component inventory, and it adds provenance and integrity evidence for each.
Which format should we ask suppliers for? Either CycloneDX or SPDX is defensible today. CycloneDX 1.7 is standardized as ECMA-424 and has carried machine-learning model support since 1.5; SPDX 3.0 has dedicated ai and dataset profiles. What matters more than the choice is that whichever format you receive is signed and carries hashes, because those are the only parts you can independently check.
Does an AI-BOM satisfy the EU AI Act? No. No document satisfies a regulation by itself, and the Annex IV technical-documentation requirement for high-risk systems is not yet applicable — it arrives in December 2027 for Annex III systems and August 2028 for Annex I product-embedded systems. A good AI-BOM covers several Annex IV items, particularly the architecture and third-party component descriptions, but Annex IV also requires dated and signed test logs and reports, which an inventory does not produce.
How often should an AI-BOM be regenerated? On triggers, not on a schedule. The useful trigger list includes retraining or dataset refresh, adapter or fine-tuning changes, dependency upgrades in the inference path, guardrail and policy-pack changes, disclosed vulnerabilities in a tracked component, and — the one most programs miss — upstream model or alias changes made by a supplier, which occur without any release on your side.
Can an AI-BOM be tested, or is it just documentation? The checkable portion can be tested directly: re-derive model and dataset digests against the deployed artifacts, validate the document's signature and transparency-log entry, enumerate running components and declared egress and diff both against the document, and read the per-claim evidence technique and confidence rather than only the concluded values. The narrative fields — collection process, known bias, intended use — cannot be verified from outside the training pipeline, and a document weighted heavily toward those is correspondingly weaker evidence.
ioSENTRIX is a CREST-accredited, ISO/IEC 27001 certified offensive security firm. Our AI and ML penetration testing engagements treat an AI-BOM the way we treat any other control document: as a set of claims to be re-derived rather than reviewed. That means hashing the artifacts actually serving traffic against the declared digests, validating signatures and attestation chains, enumerating the tool and egress surface an agentic system uses in practice and diffing it against what the document declares, and testing whether the supplier-side and pipeline triggers that should regenerate the record actually do. We assure the stack you have chosen; we do not sell a competing platform.
If you want a cheaper starting point than an engagement, run the two-pile exercise on the most recent AI-BOM you were given. Sort its fields into what you could re-derive today and what you would simply have to believe. The ratio is your answer.