AI Vendor Risk Assessment: A Questionnaire That Works

VERITY — AI Vendor Risk

AI Vendor Risk Assessment: A Questionnaire That Works

A structured, 22-question assessment for any third party whose product touches your data via large language models, machine learning systems, or autonomous agent functionality — with notes on what a defensible vendor answer actually looks like.

What It Is

Why Generic TPRM Isn’t Enough

An AI vendor risk assessment is a structured review of a third party whose product touches your data via large language models, machine learning systems, or autonomous agent functionality. It extends — but does not replace — the third-party risk management (TPRM) workflow your organization is already running.

The reason a generic TPRM questionnaire is no longer sufficient is straightforward: the questions that matter for AI vendors — training data handling, model isolation, prompt logging, retention of inferred data, fine-tuning rights, sub-processor disclosure for the underlying foundation model — are not in standard TPRM templates. The result is an assessment process that approves AI vendors with the same depth of review as a static SaaS app, and absorbs the same level of risk the Observability Gap describes.

This piece publishes the 22-question assessment Armorstack issues on behalf of clients, with notes on what a defensible answer looks like.

Why It Matters

Three Reasons AI Vendor Risk Warrants Its Own Track

AI Vendors Handle Data Differently

Prompts are a new class of data the vendor receives. Retrieval payloads are another. The vendor’s retention, logging, and use of those payloads — including for model training — needs to be reviewed before the contract closes, not after.

The Vendor’s Dependencies Are Now Your Supply Chain

Most AI features are built on a foundation model from a small number of providers. Your vendor’s posture is, in part, a function of their upstream provider’s posture. The questionnaire has to surface that chain.

Existing Certifications Don’t Cover the AI Surface

A SOC 2 Type II report or an ISO 27001 certification does not, by itself, demonstrate that a vendor’s AI features run with appropriate isolation, logging, or training-data controls. Necessary, but not sufficient, for AI vendors.

Escalation Triggers

When to Issue the Assessment

Three triggers should escalate a vendor into the AI track.

A New LLM or GenAI Feature

The vendor has launched any LLM or generative AI feature that touches your data, prompts, or documents — even if the feature is “optional.”

A Stale Contract or DPA

The vendor’s contract or DPA was last reviewed before mid-2024 — the period in which most SaaS platforms began shipping AI features. Existing agreements likely don’t contemplate model-layer use.

Regulated Data Is in Scope

The vendor handles regulated data (PHI, PCI, CUI, FERPA, GLBA) and has any AI feature in scope. Sector-specific obligations make the AI assessment non-optional.

The output of a Shadow AI Discovery typically produces the working list of vendors to run through this track.

The Assessment

The 22-Question AI Vendor Risk Assessment

We organize the questionnaire into five sections. Each question carries a tier. A defensible mid-market vendor should answer all Critical questions clearly and provide documentation on High questions.

Critical — disqualifying if unanswered or weak
High — documentation expected
Medium — good-practice indicator
Section 1 of 5 — Questions 1–6

Data Handling

Critical

1. What customer data does the AI feature receive?

Look for: explicit list — prompts, attached documents, retrieved records, inferred metadata, conversation history. A vendor that cannot enumerate this list is not ready for assessment.

Critical

2. Is customer data ever used to train, fine-tune, or improve the underlying model?

Look for: a clear “no” by default, with an opt-in path documented if applicable. The vendor should distinguish between aggregated telemetry, fine-tuning data, and reinforcement-learning feedback data.

Critical

3. What is the retention period for prompts, outputs, and any associated metadata?

Look for: stated retention period (e.g., 30 days, 90 days), with the ability to negotiate shorter retention contractually.

High

4. Is data segregated between customers at the model layer?

Look for: confirmation that one customer’s prompts cannot influence another customer’s outputs (no shared context windows, no shared retrieval indexes).

High

5. Where is data processed and stored, and is regional residency configurable?

Look for: explicit list of regions, including foundation-model provider regions, and documented residency options for regulated workloads.

Medium

6. What encryption is applied in transit and at rest, including for prompts and embeddings?

Look for: TLS 1.2+ in transit, AES-256 at rest, customer-managed key (CMK) options for sensitive deployments.

Section 2 of 5 — Questions 7–11

Model and Architecture

Critical

7. Which foundation model(s) power the AI feature, and which provider hosts the model?

Look for: explicit named models (e.g., Claude 4.x, GPT-4o, Gemini 1.5) and named hosting providers (Anthropic, OpenAI, AWS Bedrock, Azure OpenAI, etc.). The chain should be transparent.

High

8. What is the vendor’s process for model updates, and how is customer behavior tested before rollout?

Look for: documented model-update governance, customer notification on material changes, regression-testing process.

High

9. Is the AI feature deterministic or stochastic, and how is variability managed for high-stakes use cases?

Look for: acknowledgement of stochasticity, documented temperature/sampling controls, deterministic-mode availability if relevant.

Medium

10. Are there documented controls for prompt injection and indirect prompt injection?

Look for: explicit references to layered controls — input filtering, retrieval-source allowlisting, output filtering, action gating. See Prompt Injection Prevention for the eight-layer taxonomy.

Medium

11. For agent or tool-use features, what is the action authorization model?

Look for: per-action authorization, allowlisted action set, audit logs, customer-controlled approval gates for high-impact actions.

Section 3 of 5 — Questions 12–15

Security and Incident Response

Critical

12. Will the vendor provide their SOC 2 Type II report, ISO 27001 certificate, and any AI-specific attestations?

Look for: current report (within 12 months), no major exceptions, AI-relevant control coverage.

Critical

13. What is the AI-specific incident response process?

Look for: notification SLA (typically 72 hours or better), explicit handling of model-layer incidents (data leakage via inference, prompt injection compromise, supply chain compromise of the foundation model).

High

14. Is there a published vulnerability disclosure program covering the AI feature?

Look for: VDP scope explicitly including the AI surface, named contact, response SLA.

High

15. How is access to prompts, outputs, and model logs by vendor employees controlled and logged?

Look for: least-privilege access, MFA, audit logs of vendor employee access to customer data.

Section 4 of 5 — Questions 16–19

Compliance and Governance

Critical

16. Does the vendor support contractual no-train terms, BAA, and DPA addenda for AI processing?

Look for: standard willingness to execute these documents, with AI-specific terms (no training, retention, sub-processor disclosure).

Critical

17. Are sub-processors — including the foundation model provider — explicitly disclosed and contractually bound?

Look for: complete sub-processor list, disclosed in contract or DPA, with customer notification on additions.

High

18. Is the vendor aligned with NIST AI RMF, EU AI Act risk tiers, or comparable frameworks?

Look for: explicit framework mapping, model cards, governance documentation. See NIST AI RMF Implementation.

Medium

19. For regulated data (PHI, PCI, CUI, FERPA), are AI-specific compliance controls in place?

Look for: sector-specific certifications, BAA terms that cover AI processing, documented evidence of regulator-acceptable handling.

Section 5 of 5 — Questions 20–22

Operational

High

20. What logging and audit-trail data is available to the customer for the AI feature?

Look for: customer-accessible logs of prompts, outputs, retrievals, agent actions. Without this, Layer 2 of the Observability Gap cannot be closed for that vendor.

Medium

21. What is the customer’s ability to disable, restrict, or audit the AI feature post-purchase?

Look for: feature-level disable, role-based access controls, ability to restrict the feature to specific user populations.

Medium

22. What is the vendor’s roadmap and end-of-life policy for the AI feature?

Look for: stable product positioning, customer-protective end-of-life policy, data export rights on sunset.

Scoring

How to Score and Decide

We score the questionnaire on three axes.

Critical Pass Rate

All Critical questions answered clearly with defensible posture. Anything less is a hold.

High Coverage Rate

Target 80%+ on High questions. Below 60% triggers an executive escalation.

Documentation Completeness

SOC 2, DPA, BAA, sub-processor list, model card, and AI-specific addendum all available.

The output is a tiered recommendation: approve, approve with conditions (e.g., contractual addenda, configuration restrictions, sunset clause), or reject. We then feed the result into the client’s existing TPRM register so AI risk integrates with the broader vendor risk view rather than running in a parallel spreadsheet.

FAQ

Common Questions

Can we just bolt these questions onto our existing TPRM questionnaire?

Yes — and we encourage that. The point is not a separate process. It is an extended set of questions that get asked of any vendor in the AI track, integrated into the same TPRM workflow, scored against the same risk register.

What if the vendor refuses to answer a Critical question?

Treat the refusal as the answer. A vendor who will not disclose foundation-model provider, training-data use, or sub-processor list is a vendor whose risk you cannot quantify. Hold the deal pending escalation.

Our vendors keep saying “we use AWS Bedrock so we’re fine.” Is that sufficient?

Bedrock answers some questions (hosting, residency, encryption, no-train by default for foundation models). It does not answer the questions about the vendor’s own application-layer handling of prompts, outputs, retention, or agent actions. The vendor still owes you the application-layer answers.

How often should we re-assess?

Annually as a baseline. Trigger an out-of-cycle reassessment on (a) any material model change, (b) any sub-processor change, (c) any AI-related incident affecting the vendor or its foundation-model provider.

We have hundreds of vendors. Where do we start?

Risk-rank by data sensitivity, then by AI-feature criticality. Top decile gets the full 22-question assessment in cycle 1. Mid-tier gets a 10-question subset (the Critical questions). Tail-tier gets noted but deferred.

Where does Armorstack help?

Two engagement types. (1) A questionnaire-as-a-service offering — Armorstack issues the assessment on the client’s behalf, scores the responses, and feeds results into their TPRM. (2) A vCISO-led TPRM-AI integration — embedding the AI track into the client’s existing TPRM operating model. Both run from the VERITY portfolio.

Get Help Running This Assessment

If you want this 22-question assessment delivered against a specific vendor list, Armorstack runs the engagement as a fixed-scope service — we issue, score, and integrate the results into your TPRM register.

Or call 877-890-5508

Last reviewed: 2026-05-01. Authored by Dale Boehm, CEO Armorstack. CISA + CDPP.