
Your AI vendor questionnaire came back complete. Every answer is "Yes." The vendor holds a certification, publishes a trust page, and has a security whitepaper with a diagram in it. Your procurement process is satisfied.
Now answer a harder question: which of those answers have you seen evidence for, and which have you simply received?
That distinction is the whole subject of this post. There is no shortage of AI vendor question lists — most of them are fine, and most of them produce answers a vendor can give truthfully while the underlying control still fails in practice. What is scarce is a view of which answers break when somebody tests them. That is the view we have, because testing them is the work.

An AI vendor security questionnaire is a structured set of questions a buyer sends a supplier of AI capability, asking it to describe how it handles model, data and integration risk. It collects assertions. On its own it produces no evidence that any described control works.
That is not a criticism of questionnaires — it is a description of what they are for. The Cloud Security Alliance is explicit about the distinction in its own programme. Its AI Controls Matrix (AICM v1.1) carries 247 control objectives across 18 domains and is a superset of the Cloud Controls Matrix v4.1; the AI-CAIQ is the matching question set of 320 questions, which CSA describes as guiding organizations "in performing a self-assessment of their AI safety controls or an evaluation of third-party vendors."
Look at how those answers are then graded. Under STAR for AI, Level 1 is earned by completing and submitting the AI-CAIQ self-assessment. A second Level 1 tier, Valid-AI-ted, runs the submitted questionnaire through automated scoring — it validates the answer document. Only Level 2 brings in a third party, and CSA defines it as requiring both a third-party ISO/IEC 42001 certification and a Valid-AI-ted AI-CAIQ.
That ladder is worth reading carefully, because it is the clearest published statement of the thing buyers get wrong: a Level 1 entry in a public registry is a vendor's description of itself, and only Level 2 brings a third party into it. Check which tier a listing represents before you treat it as assurance. This is the same problem we have written about in general third-party risk assessment, sharpened by the fact that AI systems change underneath their own documentation.
Below are the answers that, in our experience, most often survive a questionnaire and not an engagement. We are describing patterns across assessments, not any identifiable customer, and every one of them comes with the test that settles it.
"Our guardrails prevent that." Guardrails are usually instructions in a system prompt. OWASP is blunt about what that is worth. Its 2025 entry on system prompt leakage — an entry the 2026 edition has since reorganized — states that "the system prompt should not be considered a secret, nor should it be used as a security control," and that privilege separation and authorization bounds checks "need to occur in a deterministic, auditable manner, and LLMs are not (currently) conducive to this." The test: adversarial input against the deployed endpoint, counting refusals and bypasses with reproduction steps, not a review of the prompt text.
"We don't train on your data." This is frequently true and almost never verifiable from outside. It is a contractual and policy property, not a technical one you can observe. OWASP's 2025 guidance on sensitive information disclosure frames it exactly that way — terms of use allowing opt-out, plus data sanitization on the provider's side. The test: read the contract and the subprocessor list, and separately test what is observable — whether your data crosses a tenant boundary. OWASP names the adjacent failure directly, at the user level: a user receiving a response containing another user's personal data through inadequate data sanitization. That is a leakage path you can actually probe.
"The model is isolated per tenant." Isolation failures are observable when they occur and invisible when they do not, so "we found nothing" is a weaker statement than it sounds. The test: cross-tenant retrieval attempts against shared indexes, embedding stores and caches, with the boundary defined before testing starts, not after.
"We log everything." Logging is detection, not prevention. OWASP's 2025 excessive agency entry says so plainly — logging and monitoring do not stop the behavior, they limit the damage after the fact. The test: trigger an action that should be blocked, then check whether it was blocked or merely recorded. Those are different findings.
"We use a well-known foundation model, so it's secure." The model's provenance is not your system's security posture, and provenance itself is thinner than most buyers assume. OWASP's 2025 supply chain entry states that "currently there are no strong provenance assurances in published models," and recommends third-party model integrity checks with signing and file hashes to compensate. The test: model and component integrity verification, plus an AI bill of materials — which is why AI supply chain security is a separate workstream from model selection.
"A human reviews every action." Ask what the human sees, how long they have, and what the default is when they do nothing. A review step that approves by timeout is not oversight. The test: exercise the approval path under realistic volume, and check the fail-open behavior — the question set for agentic systems is different from the one for a chatbot.

Ask questions whose answers are artifacts rather than adjectives. The shape that works is: what will you give me, and what may I do to verify it?
Ask for the boundary, in writing. Which model versions, hosted where, with which subprocessors, and what changes without notice. A vendor that cannot draw its own system boundary cannot be assessed, and you cannot scope a test of your own integration without it.
Ask what the model is allowed to do, not what it is told to do. Tool permissions, outbound network reach, write access, and the identity each action runs under. This is where non-human identity becomes the real access-control question.
Ask for evidence that is dated and scoped. A penetration test report with a date, a version and a scope statement is evidence. A badge is not. If the report is old relative to the model version in production, say so out loud.
Ask what their third-party assurance actually covers. ISO/IEC 42001:2023 certifies an AI management system — how the organization governs AI. It is meaningful, and it is not a statement that a given model held under attack. ISO published ISO/IEC 42006:2025 specifically to set requirements for the bodies doing that certifying, which tells you the certification's value depends on the auditor.
Ask whether you may test. The most informative answer in the whole process is the response to "may we run an assessment against our own tenant, and under what rules." NIST's AI RMF Playbook points the same way: among the documentation prompts it offers under GOVERN 6.1, the third-party risk subcategory, is "Did you ensure that the AI system can be audited by independent third parties?" A vendor that welcomes scoped testing is telling you something; so is one that refuses it.
A composite example, assembled from patterns rather than any single organization: a healthcare SaaS buyer runs a long AI assessment questionnaire on a document-summarization vendor. Everything comes back green. A scoped test against the buyer's own tenant finds that uploaded files are chunked into a shared vector index, and a crafted query returns a passage from another customer's document. The vendor's answers were not dishonest — no question in the set had asked how retrieval was partitioned, and no answer had been tested.
They help you ask completely, and they do not substitute for evidence.
CSA's AICM and its 320-question AI-CAIQ give the widest AI-specific question coverage we can point to, and the STAR tiers tell you what a given registry entry does and does not mean. The NIST AI Risk Management Framework's GOVERN function carries the dedicated third-party subcategories — GOVERN 6.1 on policies for third-party AI risk, 6.2 on contingency for failures in third-party data or AI systems deemed high-risk. The OWASP Top 10 for LLM Applications gives the risk vocabulary; a 2026 edition was published in August 2026, in which excessive agency ranks third, so check which edition a report or a vendor is citing. For agentic vendors, OWASP's Top 10 for Agentic Applications uses its own ASI01–ASI10 identifiers and is the more relevant list.
For a general-purpose model provider, there is one genuinely external artifact worth demanding: under the EU AI Act, providers must draw up and make publicly available a summary of the content used for training, following the AI Office's template. Ask for the date as well as the document — the obligation has applied to newly placed models since August 2025, while providers of models already on the market when it took effect have until August 2027, so its absence is not automatically a red flag. It is something you can read without the vendor's permission, which makes it rarer than it sounds.

Is there a standard AI vendor security questionnaire? CSA's AI-CAIQ, which maps to the AI Controls Matrix, is the broadest AI-specific vendor question set published by a major security body that we can point to. NIST and ISO publish frameworks and management-system standards rather than questionnaires; OWASP publishes risk taxonomies, plus narrower vendor evaluation criteria aimed at AI red-teaming providers. The EU's model contractual clauses for AI procurement are contract terms, not a question set.
Does a CSA STAR listing mean a vendor has been audited? It depends on the level, and the difference is substantial. STAR Level 1 is a self-assessment the vendor submits; the Valid-AI-ted tier applies automated scoring to that submission. Level 2 is the tier that involves a third party. Check which one a listing represents before treating it as assurance.
Our vendor is ISO/IEC 42001 certified. Is that enough? It is good evidence that AI governance exists as a managed system, and it is not evidence that a specific control held under attack. Treat it as a reason to ask better technical questions, not as a reason to stop asking.
Can we penetration test a SaaS AI vendor? Often yes, against your own tenant and under agreed rules of engagement — but it has to be negotiated, ideally before contract signature when you still have leverage. Testing scope, notification and permitted techniques belong in the contract, not in an email thread after an incident.
How often should AI vendors be reassessed? Tie it to change rather than the calendar. A model version change, a new tool permission, a new subprocessor or a change to data handling are all events that can invalidate a prior answer, and any of them can land without a release note.
ioSENTRIX is a CREST-accredited, ISO/IEC 27001 certified offensive security firm. We work the gap this post describes: taking the answers your AI vendors have given you and testing the ones that can be tested, against your own tenant, under scoped rules of engagement. That is AI and ML penetration testing applied to third-party risk, and it produces a dated report with reproduction steps rather than a rating. We assure the stack you chose — we do not sell a competing platform, and we have no interest in telling you a vendor is bad when the honest finding is that a question was never asked.
If you have a questionnaire sitting completed in a folder, the useful next step is small: pick the five answers you would least like to be wrong, and find out which of them are testable. Talk to us about scoping that.