The short answer: buy an operating system, not a demo
Choose an AI automation vendor by the evidence it can produce for your workflow: representative evaluation cases, explicit data and permission boundaries, human review where consequences are meaningful, observable failure and recovery paths, named operating owners, and a practical exit plan. A fluent demo is only evidence that the happy path can be presented.
Give every shortlisted partner the same task, sample cases, non-goals, systems, volume assumptions, critical failures, and acceptance rules. Then compare what each proposal proves, assumes, excludes, and leaves your team responsible for. The strongest answer may narrow the automation instead of promising everything.
Fit
A defined workflow and supported outcome.
Evidence
Tests that include difficult and unsafe cases.
Control
Owners, limits, recovery, and an exit path.
Define the workflow before comparing vendors
Write the current trigger, inputs, decisions, actions, exceptions, handoffs, and confirmed outcome. Name who does the work today and who will remain accountable after automation. Include representative records only after removing or protecting data that candidates do not need during selection.
Separate a deterministic step from a judgment step. Copying an approved field between systems, checking a required value, and routing by a known rule may not need generative AI. Classifying an ambiguous message or drafting a response may. The earlier AI agent versus workflow automation guide helps choose that boundary; the AI workflow readiness checklist helps decide whether the process is ready to automate at all.
Ask candidates to mark unknowns rather than hiding them inside a fixed promise. If access, data quality, exception rate, or a third-party API has not been checked, it belongs in discovery or a pilot—not in an unqualified estimate.
Score seven kinds of vendor evidence
| Area | Request | Warning sign |
|---|---|---|
| Workflow | Restated flow, users, exceptions, non-goals | Generic use-case deck |
| Evaluation | Versioned cases, rubric, results, critical failures | Anecdotal examples only |
| Data | Sources, retention, access, deletion, provider path | ‘Secure’ without a map |
| Controls | Permissions, approvals, limits, fallback, disablement | Model guardrails alone |
| Operations | Monitoring, incidents, changes, cost signals | Launch ends the plan |
| Ownership | Accounts, code, configuration, documents, responsibilities | Vendor-only access |
| Commercial | Assumptions, dependencies, recurring cost, pilot exit | Headline total only |
Score demonstrated, partially demonstrated, promised, unknown, or not applicable. Do not turn the matrix into false arithmetic: one unsafe permission boundary can outweigh several polished presentation categories.
Ask for evaluation evidence before architecture theatre
A useful evaluation set represents normal work, ambiguous cases, missing inputs, conflicting instructions, unsupported requests, duplicates, malicious or untrusted content, and the failures that would be costly even if rare. Define the expected outcome, acceptable variation, reviewer, and critical-failure rule for each class.
Ask which prompt, model, retrieval sources, tools, policies, and test-set version produced the result. Re-run the evaluation when one of those changes. NIST describes its voluntary AI Risk Management Framework as applying across design, development, use, and evaluation; vendor selection should therefore include the operating lifecycle, not only model choice.
Ask the vendor to show failed cases and what changed because of them. A credible team can explain where the system should defer, refuse, route to a person, or use a deterministic fallback. If all evidence is perfect, the test set is probably too gentle.
Trace every data source, permission, and consequential action
Draw where inputs originate, what is sent to each model or provider, what is stored, who can inspect it, and how retention or deletion works. Include logs, evaluation datasets, human review screens, vector stores, analytics, and support exports. A data-processing statement without the actual application path is incomplete.
Grant the narrowest tool permissions needed for the supported task. Reading a mailbox does not imply permission to send; drafting a refund does not imply permission to issue one. OWASP’s Excessive Agency guidance recommends minimizing extensions, permissions, and autonomy and using human approval for high-impact actions. Ask where those limits live in code and configuration—not only in a prompt.
Test both direct and indirect prompt injection when untrusted documents, websites, messages, or attachments can enter the context. OWASP documents how such content can alter intended behavior. No single filter removes the need for bounded tools, output validation, approval, monitoring, and recovery.
Use a bounded pilot to test the complete operating loop
Restrict the pilot by user group, task, data source, action, volume, language, or consequence. Preserve the current process as a fallback. Give reviewers a clear way to approve, correct, reject, and explain outcomes without training them to accept the model by default.
Define stop conditions before the first live case: unauthorized access, a critical action without approval, unacceptable failure on the task rubric, unavailable fallback, unexplained cost or latency, or an incident the team cannot diagnose. Define what permits expansion too. “Users liked it” is useful feedback, but it is not a release criterion by itself.
End with a decision: expand, narrow, hold, or stop. NIST’s Generative AI Profile organizes suggested actions around governing, mapping, measuring, and managing risk. A pilot should produce evidence for all four, proportionate to the workflow.
Write ownership and exit terms before launch
Name the business owner for supported behavior, the technical owner for releases, the incident owner, the person who can disable the system, and the approver for a broader permission or task. Record who maintains the evaluation set and who decides when production failures become regression cases.
Confirm who controls source repositories, deployment accounts, model and cloud accounts, secrets, domains, integrations, prompts, configuration, logs, dashboards, and documentation. Identify proprietary vendor components and the practical replacement path. “You own the output” is not the same as being able to operate or move the system.
Price model use, retrieval, storage, tool calls, monitoring, evaluation, reviewer time, support, retries, and minimum commitments separately where possible. Use expected ranges and assumptions from your workflow; avoid a universal automation ROI claim.
Compare every candidate in one selection brief
Record the workflow, non-goals, representative cases, critical failures, data and action boundary, required approval, fallback, pilot scope, acceptance and stop conditions, dependencies, recurring-cost assumptions, responsibilities, assets received, and exit path. Attach evidence; do not paste sensitive records into the brief.
If the vendor cannot respond to that common boundary, the proposals are not comparable. If the workflow is ready and the evidence request is clear, a focused AI automation engagement can start with the smallest controllable step rather than a platform-wide mandate.
AI automation vendor selection FAQ
What should I ask an AI automation vendor?
Ask the vendor to restate the workflow, show how success and critical failures will be evaluated, identify every data source and permitted action, explain human review and fallback behavior, describe monitoring and incident ownership, disclose dependencies and recurring costs, and define what your team receives at handover. Require evidence against representative cases rather than a generic demo.
How do I compare AI automation proposals fairly?
Give every vendor the same workflow, sample cases, non-goals, constraints, expected volume, systems, critical failures, and decision criteria. Compare assumptions and exclusions separately from the price. A proposal that narrows unsafe scope may be stronger than one that promises every requested feature.
Should an AI vendor offer a pilot first?
Usually when value or technical feasibility is still uncertain. The pilot should have a real task, representative test set, restricted users and permissions, a human checkpoint where consequences are meaningful, predefined stop conditions, measurable acceptance criteria, and a clear decision at the end: expand, narrow, hold, or stop.
What proof should an AI automation vendor provide?
Useful proof includes an evaluation plan and versioned results, examples of critical failures and mitigations, a data and tool boundary, permission design, audit and monitoring views, fallback behavior, incident and change processes, deployment and dependency details, and an explicit responsibility map. The exact depth should match the workflow's consequence and data sensitivity.
Who should own the AI automation after launch?
Your organization should retain a business owner for supported behavior and risk acceptance, access to its data and system accounts, and visibility into performance and incidents. The vendor or internal technical team may operate releases and monitoring, but responsibilities for approvals, credentials, changes, support, disablement, and exit should be written down.
