AI vendor security questions should reveal the exact system path: which data enters, who can access it, which models and subprocessors receive it, what the system may do, how the release was tested, and who responds when it fails.

Do not accept a product-wide “yes” where the control depends on a plan, region, configuration, connector, or contract. Ask for the scope, evidence, and consequence of the answer.

The 25 questions below are designed for a production AI application or agent. They complement a standard software-security review. They do not replace legal advice, a privacy assessment, or controls required by your organization.

A useful answer has scope, evidence, and a consequence

A questionnaire often stops at “Do you encrypt data?” The vendor answers yes, and the cell turns green.

The better review has four fields:

FieldWhat it should contain
AnswerThe vendor's direct response
ScopeProduct, plan, region, data path, users, and actions covered
EvidenceDiagram, contract, policy, configuration, test, log, report, or current certification
DecisionAccept, require a condition, narrow the release, or stop

This structure respects both sides. A small vendor may have a sound control without a formal certification. A certified provider may have an audit whose scope excludes the feature you plan to use. The evidence decides what the answer means for this system.

NIST's Generative AI Profile recommends updating acquisition and supplier-risk processes for privacy, security, intellectual property, model libraries, tools, APIs, and ongoing third-party monitoring. The use-case boundary remains central.

Define the system being assessed

  1. What exact business function and deployment are you proposing? Ask the vendor to name the product, plan, region, model path, integrations, and action capabilities.
  2. Which responsibilities belong to you, us, and other providers? Look for a clear responsibility map.
  3. Which controls are unavailable in the proposed configuration? A candid limitation is more useful than a universal claim.

The first three questions prevent later answers from drifting between a marketing site, an enterprise tier you are not buying, and a custom configuration that has not been built.

Ask how data moves and persists

  1. Which data categories will the system receive, retrieve, generate, or infer?
  2. Can you provide the complete data-flow and storage diagram? Include prompts, retrieved context, tool calls, logs, caches, approval views, outputs, and backups.
  3. Which model and service providers receive our data? Name the exact service and routing path.
  4. Which subprocessors participate, in which regions, and under which agreements?
  5. Is customer content used for provider or shared-model training? Ask for the contract and configuration that determine the answer.
  6. What are the retention periods for each data surface? “We do not store documents” does not answer prompt, trace, cache, or backup retention.
  7. How do export, return, deletion, and deletion verification work? Request a test or procedure alongside the policy sentence.

Region deserves precision. European primary storage does not prove that every model request, connector call, email, or support path remains in Europe.

Test identity, tenants, and retrieval

  1. How are users and service identities authenticated? Include administrative and background-job identities.
  2. How is customer or workspace isolation enforced and tested? Ask for the boundary and a denied-access case.
  3. How does retrieval enforce the current user's source permissions? Authorization should occur before content enters model context.
  4. What happens when a user's access changes after content was indexed or cached?
  5. Which credentials can bypass normal row, role, or tenant policies? Ask how those credentials are stored, monitored, and restricted.

Row-level security can be a valuable isolation control. It does not automatically constrain a service-role process or an application search tool. The vendor should explain how database policy and application authorization work together.

Examine model, prompt, and output controls

  1. How do you treat instructions inside retrieved emails, files, web pages, and tool results?
  2. Which prompt-injection and data-exfiltration cases do you test? Request examples and expected behavior.
  3. How are model outputs validated before storage, display, or execution?
  4. How do you record source evidence, uncertainty, and model or prompt version where the use case requires them?
  5. What changes when a model, prompt, retrieval method, or provider is updated? Ask which changes trigger renewed testing or customer notice.

OWASP's LLM application risk work is a useful threat-model input. A vendor should translate those risk classes into tests for the proposed application.

Separate tool access from action authority

  1. Which tools and operations can the AI request in our configuration? A connector name is too broad; ask for read, draft, send, change, delete, and administrative operations.
  2. Which actions are permitted, approval-gated, or prohibited? Ask where the policy is enforced.
  3. How is approval authenticated and bound to the exact action? The approver should see material fields and the destination.
  4. How do you prevent duplicate or replayed side effects? Request the idempotency and partial-failure behavior.
  5. Can we inspect requests, refusals, approvals, execution results, and operator actions? Ask what is logged, who can access it, and how long it remains.

These questions expose a common gap. The model may request an action, but the application needs deterministic code to decide whether that request can reach a real system.

Review deployment, incidents, and continuity

The 25 questions above describe the main control surfaces. The evidence pack should also show how the vendor operates them.

Ask for the release process, environment separation, production promotion, rollback, backup and restore tests, monitoring, incident severity model, notification commitments, continuity dependencies, support hours, and credential ownership if the relationship ends.

The UK NCSC secure AI development guidelines span design through operation and maintenance. That framing matters for procurement because the vendor you select will keep changing the system after the initial review.

For a consequential path, ask to see a system-specific release record. It should identify the version, data, providers, users, permissions, actions, tests, fallback, approval, and known limits.

Score evidence without rewarding confident prose

Use a simple four-level rubric for each material answer.

LevelEvidence qualityBuyer response
0No answer or a product sloganTreat as unknown; do not release the dependent path
1Policy or verbal statementSeek configuration or operational evidence
2Current diagram, contract, configuration, or procedureValidate that scope matches the proposed system
3Tested evidence for the exact path and versionAccept with stated residual limits and monitoring

Do not add the scores into a false universal grade. Some controls are release blockers, while others can be accepted as limitations. Weight them by the data and consequence of the use case.

Use gaps to shape the contract or release

An unanswered question creates a decision point, and abandoning the supplier is only one possible result.

The contract can prohibit specific data, require a named region, set incident-notification terms, constrain subprocessors, establish deletion duties, or require approval for an action class. The technical release can use a dedicated source folder, remove a sending tool, rely on fictional data, or remain a controlled evaluation until evidence exists.

Compsia prepares a trust pack for the exact production system, mapping every control to its tested path and evidence. Buyers should expect that level of specificity from every supplier.

The secure enterprise AI solutions guide explains how this evidence fits into the wider buying decision.

Procurement questions about AI vendors

Is SOC 2 enough for an AI vendor?

SOC 2 can provide useful independent evidence when the report scope matches the service. It may not answer application-specific retrieval, model routing, prompt injection, tool permissions, or action approval. Review both.

Should we require zero data retention?

Require the retention pattern your use case and obligations need. Verify which surfaces the claim covers. Application logs, connector traces, approval records, and backups may follow different rules from the model provider.

What should disqualify an AI vendor?

Disqualifiers depend on the use case. Examples include an unauthorized provider path, inability to enforce a required access boundary, no incident owner, direct consequential writes without policy, or refusal to state material limitations.

Primary references

  1. Generative AI Profilenvlpubs.nist.gov
  2. LLM application risk workowasp.org
  3. UK NCSC secure AI development guidelinesncsc.gov.uk
  4. NIST SP 800-218csrc.nist.gov

Continue reading: Secure Enterprise AI Solutions: What Buyers Should Compare.