You are viewing the current published version.
Business Advanced ChatGPT

Evidence-Grounded AI Vendor Evaluation and Procurement Risk Scorecard

Evaluate an AI vendor using evidence-qualified scoring, data-flow analysis, security and privacy controls, commercial risk, decision gates, and a safeguarded pilot plan.

View all versions
Best foranalysis
ToolChatGPT
DifficultyAdvanced
Full Prompt
Evaluate the proposed AI vendor for procurement or renewal using only the supplied materials and any sources ChatGPT can actually access in the current session. Produce decision support for authorized business, procurement, security, privacy, legal, and compliance reviewers; do not present the analysis itself as organizational approval.

## Evaluation context
Business context: [Business context]
AI tool or vendor name: [AI tool or vendor name]
Vendor website or product summary: [Vendor website or product summary]
Intended use case: [Intended use case]
Departments or users: [Departments or users]
Data the tool will access: [Data the tool will access]
Data the tool will store or process: [Data the tool will store or process]
Integrations required: [Integrations required]
Compliance requirements: [Compliance requirements]
Security requirements: [Security requirements]
Budget or pricing information: [Budget or pricing information]
Contract or procurement constraints: [Contract or procurement constraints]
Existing alternatives: [Existing alternatives]
Risk tolerance: [Risk tolerance]
Definition of done: [Definition of done]

## Input and access rules
Treat the vendor identity, intended use, affected users, expected data categories, material integrations, applicable requirements, risk tolerance, and decision objective as minimum inputs. If any are absent or materially ambiguous, ask a concise set of blocking questions before issuing a recommendation. You may still create a clearly labeled preliminary assessment when bounded progress is safe.

Pricing details, contracts, data processing terms, security reports, subprocessors, architecture diagrams, retention schedules, incident history, accessibility reports, support terms, and exit documentation are useful supporting inputs. Mark their absence as an evidence gap rather than inventing their contents.

Use ChatGPT to organize supplied evidence, compare claims with requirements, expose conflicts, calculate transparent scores, and draft questions and controls. Do not imply that ChatGPT accessed a website, contract, private system, vendor portal, configuration, certification report, or live integration unless that access occurred in the current session. A URL alone is not evidence of its contents. If browsing is available and used, identify the pages consulted and retrieval date; otherwise request pasted or uploaded source material.

Do not contact the vendor, accept terms, approve expenditure, sign a contract, change configurations, connect an integration, provision users, upload business data, run a pilot, delete records, or certify compliance. Those actions require authorized humans and, where applicable, security, privacy, legal, compliance, procurement, and system-owner approval.

Do not request passwords, API secrets, private keys, authentication tokens, or unnecessary personal data. Recommend redacted documents and synthetic or de-identified pilot data where practical.

## Evidence discipline
Create an evidence ledger and assign each source a stable identifier. For every material conclusion, score, risk, and decision condition, cite one or more evidence identifiers or mark it unsupported.

Classify information as one of the following:
- Supplied fact: directly stated in business-provided material.
- Vendor claim: stated by the vendor but not independently or contractually verified.
- Independent evidence: supported by a credible third party, with scope and date recorded.
- Contractual commitment: contained in applicable executed or proposed terms, with the relevant clause identified.
- Observation: directly visible from an artifact or authorized demonstration available in the session.
- Assumption: a declared premise used for bounded analysis.
- Unknown: no adequate evidence is available.
- Conflict: sources disagree or apply to different products, plans, regions, or dates.

Never convert a vendor claim into a verified control. Check whether evidence applies to the correct product, service tier, hosting region, deployment model, integration, legal entity, and evaluation date. Flag stale, marketing-only, partial, inaccessible, or scope-mismatched evidence. Preserve conflicting claims and state what would reconcile them.

## Evaluation workflow

### 1. Establish scope and data flow
Summarize the business outcome, users, administrators, deployment model, integrations, and alternatives. Map the expected flow of prompts, files, recordings, metadata, outputs, telemetry, support data, and account data through collection, transmission, storage, model processing, human review, subprocessors, export, retention, deletion, and backup expiry.

Classify likely data sensitivity, including public, internal, confidential, personal, special-category, financial, health, customer, employee, authentication, intellectual-property, or regulated data. Do not assume a data category is permitted merely because the tool can process it.

### 2. Inventory and qualify evidence
Build the evidence ledger before scoring. Record source, owner or publisher, document date, retrieval date when applicable, product and plan scope, region, evidence class, relevant claim, limitations, and identifier. List missing critical artifacts such as a data processing agreement, security report, subprocessor list, retention policy, model-training terms, deletion procedure, breach notice terms, service levels, pricing schedule, and export documentation.

### 3. Apply critical decision gates
Evaluate these gates before calculating an overall result:
- Prohibited or unapproved data would enter the service.
- The vendor cannot explain training, retention, deletion, residency, or subprocessors for material data.
- Required SSO, MFA, role-based access, audit logging, encryption, tenant isolation, or administrator controls are absent or unverified.
- Applicable legal, regulatory, contractual, records-management, intellectual-property, accessibility, or data-transfer requirements cannot be met.
- Material contract terms conflict with mandatory requirements or accepted risk.
- The use case creates high-impact automated decisions without appropriate human oversight, testing, appeal, and accountability.
- No feasible containment, offboarding, export, access revocation, or incident-response path exists.

A failed gate should normally lead to Reject, Defer pending information, or a tightly restricted Pilot first recommendation. Explain any exception and identify the human risk owner authorized to accept it.

### 4. Score the vendor
Score each category from 1 to 5 only when enough evidence exists:
1 = confirmed material failure or unacceptable exposure
2 = major gaps requiring substantial remediation
3 = partially meets requirements with enforceable conditions
4 = meets documented requirements
5 = exceeds requirements with strong, applicable evidence

Use NE for not enough evidence and NA only when a category is genuinely irrelevant. Do not treat NE as zero or silently exclude it. Report evidence coverage as the percentage of applicable categories with a supported score. Do not issue a numeric overall score when coverage is below 70 percent or when security, privacy, compliance, or data-governance gates remain unresolved.

Assess:
- Business fit and measurable value
- Usability, accessibility, and adoption fit
- Security architecture and assurance
- Privacy and data protection
- Compliance and contractual readiness
- Identity, roles, and administrator controls
- Auditability, logging, and monitoring
- Integration and technical fit
- Pricing transparency and total cost exposure
- Vendor maturity, resilience, and incident response
- Support and service management
- Portability, termination, and lock-in risk

For every score, provide the rationale, evidence identifiers, confidence level, requirement gap, and condition needed to improve the score. If priorities justify weighting, show the proposed weights and obtain human agreement before using a weighted total; otherwise use equal weights and label the result provisional.

### 5. Perform domain assessments
Evaluate data use for foundation-model training, fine-tuning, product improvement, human review, abuse monitoring, and telemetry. Examine opt-out scope, default settings, retention periods, deletion behavior, backup expiry, legal holds, residency, cross-border transfers, subprocessors, data ownership, output rights, confidentiality, and intellectual-property protections.

Evaluate authentication, SSO, MFA, lifecycle provisioning, least privilege, role separation, service accounts, encryption, key management, tenant isolation, secure development, vulnerability management, penetration testing, audit reports, log availability, incident notification, business continuity, disaster recovery, recovery objectives, and administrator visibility. Distinguish certification possession from control effectiveness and scope.

Evaluate onboarding, workflow change, training, acceptable-use rules, support escalation, accessibility, internal ownership, model-output review, accuracy limitations, bias or harmful-output risks, monitoring, incident handling, and shadow-use risk.

Evaluate licensing units, usage limits, implementation services, integration work, storage, support tiers, overages, taxes, price escalation, renewal notice, minimum commitments, suspension rights, indemnities, liability limits, insurance, service levels, termination assistance, export formats, deletion commitments, and switching costs. Separate known prices from estimates and identify estimate assumptions.

### 6. Build the risk and control package
Create a risk register covering privacy, security, compliance, legal, commercial, operational, integration, model behavior, vendor viability, and exit risk. Each risk must include cause, event, consequence, affected data or process, inherent severity and likelihood, evidence, existing controls, proposed treatment, residual risk, owner, due date, decision gate, and status.

Create a prioritized vendor evidence request. Questions must be answerable and tied to a specific gap, risk, score, or contract condition. Separate requests for documents, written confirmations, product demonstrations, technical validation, and contract amendments.

### 7. Form the recommendation
Choose exactly one recommendation:
- Approve
- Approve with conditions
- Pilot first
- Defer pending information
- Reject

An Approve recommendation requires supported critical controls, acceptable residual risk, adequate evidence coverage, and no unresolved mandatory gate. Approve with conditions requires named conditions, accountable owners, deadlines, and a consequence if a condition is not met. Pilot first requires a bounded hypothesis and safeguards; it is not production approval. Defer pending information must list the evidence needed to resume. Reject must identify the failed requirements and whether reconsideration is possible.

State which authorized functions must review or approve the decision. Never state that the vendor has been approved, verified, certified, contracted, configured, tested, deployed, monitored, or deleted unless supplied execution evidence proves that action occurred. Use proposed, pending, unverified, blocked, or not performed as appropriate.

### 8. Design a safeguarded pilot when appropriate
If the recommendation permits a pilot, define its business hypothesis, participants, duration, approved use cases, prohibited uses, data restrictions, synthetic-data preference, administrator configuration, access controls, logging, human review, user training, support route, incident procedure, success metrics, failure thresholds, review date, and accountable owner.

Include stop conditions for unauthorized data exposure, material control failure, unexpected training or retention behavior, security incident, harmful output, regulatory conflict, uncontrolled cost, or inability to export or delete pilot data. Provide a recovery and exit procedure covering disabled access, revoked integrations and tokens, data export, deletion request, backup-expiry confirmation, record preservation, user communication, and fallback workflow. Label all of these as proposed until execution evidence is supplied.

## Required deliverable

### A. Decision brief
Include the recommendation, confidence, evidence coverage, critical gates, top risks, expected business value, estimated commercial exposure, mandatory conditions, and required human approvers.

### B. Scope and data-flow map
Use a table with: Stage | Data or metadata | Sensitivity | Source | Recipient or subprocessor | Purpose | Region | Retention | Training or improvement use | Control | Evidence ID | Unknowns.

### C. Evidence ledger
Use a table with: Evidence ID | Artifact or source | Evidence class | Publisher or owner | Date | Product, plan, and region scope | Supported claim | Limitations | Status.

### D. Critical-gate assessment
Use a table with: Gate | Requirement | Expected observation | Actual supported observation | Evidence ID | Pass, Fail, or Unresolved | Consequence | Required resolution.

### E. Evaluation scorecard
Use a table with: Category | Score or NE or NA | Weight if agreed | Rationale | Evidence ID | Confidence | Requirement gap | Improvement condition. Show evidence coverage and any permitted aggregate calculation.

### F. Detailed assessments
Provide separate findings for data governance and privacy, security and assurance, compliance and contracting, operational and model-risk fit, integrations, commercial exposure, and portability. For each finding state the requirement, supported observation, evidence, gap, risk, and proposed control.

### G. Risk register
Use the fields defined in the risk and control package. Rank risks without hiding low-confidence or unresolved high-impact items.

### H. Vendor evidence requests and negotiation points
Prioritize each item as blocking, pre-contract, pre-pilot, pre-production, or monitor. Link every request to an evidence gap, risk identifier, score, or decision condition.

### I. Recommendation and approval path
State the selected recommendation, rationale, rejected alternatives, residual risks, conditions, owners, deadlines, escalation route, and required reviewers. Clearly distinguish advisory analysis from formal approval.

### J. Pilot, rollout, monitoring, and exit controls
When relevant, provide the proposed pilot plan, production-entry criteria, metrics, review cadence, incident triggers, stop conditions, fallback, and offboarding evidence requirements. If a pilot is unsafe or premature, state why instead of drafting an executable rollout.

### K. Verification and acceptance matrix
Use a table with: Check | Expected evidence or observation | Actual supplied evidence or observation | Evidence ID | Status | Owner | Resolution needed. At minimum reconcile data-use terms, retention and deletion, subprocessors and residency, identity controls, audit logging, incident terms, contractual requirements, price assumptions, export capability, pilot safeguards, and unresolved source conflicts.

### L. Final decision checklist
Confirm whether minimum inputs are complete, critical evidence is applicable and current, unsupported claims are labeled, scores trace to evidence, gate failures control the recommendation, commercial totals reconcile with stated assumptions, risk owners are assigned, conditions are measurable, required approvers are named, and proposed actions are not misrepresented as completed.

End with a short section titled Unresolved Items and Handoff that lists blocking unknowns, who must resolve them, the evidence required, and the next authorized decision point.

Variables to Replace

  • Business context
  • AI tool or vendor name
  • Vendor website or product summary
  • Intended use case
  • Departments or users
  • Data the tool will access
  • Data the tool will store or process
  • Integrations required
  • Compliance requirements
  • Security requirements
  • Budget or pricing information
  • Contract or procurement constraints
  • Existing alternatives
  • Risk tolerance
  • Definition of done

How to Use This Prompt

In ChatGPT, replace every bracketed variable with the organization’s details. Provide relevant source materials such as vendor security and privacy documentation, contract terms, data processing agreements, pricing, architecture and data-flow information, compliance requirements, audit reports, subprocessor lists, retention terms, and internal policies. Redact secrets and unnecessary personal data, then run the prompt and route the resulting analysis to authorized procurement, security, privacy, legal, compliance, and business reviewers.

Example Use Case

A company is considering an AI meeting assistant that records employee and customer calls. It supplies the vendor’s privacy terms, data processing agreement, security report, pricing, retention settings, integration design, and internal recording requirements to produce an evidence-qualified scorecard, identify unresolved training and deletion risks, negotiate contract conditions, and define a restricted pilot for human approval.

Published change

Major: Replace the legacy AI Vendor Evaluation and Procurement Risk Scorecard template with a domain-specific input, evidence, authority, safety, workflow, output, and verification contract.