Free buyer checklist + CSV
AI vendor evaluation checklist
Compare AI development companies on evidence, not pitch quality. Ask the same 12 questions, score every answer and record the proof before you commit a production budget.
Download the free CSVScoring rule
Give each answer 0 for vague or unsupported, 1 for partial with evidence missing, or 2 for specific and independently verifiable. A score of 0 on data handling or ownership should pause the selection, regardless of the total.
Technical fit
Have you built this specific kind of AI system before?
Good evidence: A credible answer names a comparable system, the data and constraints, and what the team learned when it failed.
Red flag: A list of technologies with no delivered outcome attached.
Who will actually work on the project?
Good evidence: Meet the named engineers and understand each person's responsibility before signing.
Red flag: Technical questions are repeatedly deferred by sales or account staff.
What happens if the model provider changes price, removes a model or has an outage?
Good evidence: The architecture isolates model dependencies and the team has tested at least one practical alternative.
Red flag: The plan assumes today's model, pricing and uptime will remain unchanged.
Delivery proof
Can we speak with a reference client whose project resembles ours?
Good evidence: The vendor offers a candid conversation about a similar production system, including what went wrong.
Red flag: Only a written testimonial or a reference from an unrelated use case.
Which systems are still running in production today?
Good evidence: The answer clearly separates live systems from pilots and explains why some work stopped.
Red flag: Every prototype is presented as a production success.
What measurable business KPI did the work move?
Good evidence: You get a baseline, measured result, period and honest attribution caveats.
Red flag: The proof stops at capability or client-satisfaction claims.
Security and operations
Where does our data go and who can access it?
Good evidence: The vendor documents residency, retention, isolation, access controls and the compliance posture relevant to you.
Red flag: Your data is safe is the whole answer.
How will you monitor quality, failures and drift after launch?
Good evidence: The plan names task-level evals, alerts, owners, review cadence and incident response.
Red flag: Launch is treated as the finish line.
How will the system handle ten times the usage or a model change?
Good evidence: Load and unit-cost assumptions are explicit, and the model layer can change without a rewrite.
Red flag: There is no load-test, unit-cost or migration plan.
Commercial terms
What is included in the price and what is explicitly excluded?
Good evidence: Build, model, hosting, monitoring, support and change costs are separated in writing.
Red flag: One total hides the operating assumptions and future exclusions.
Who owns the code, data pipeline, prompts and model artifacts?
Good evidence: Transfer rights, repositories, documentation and the exit path are unambiguous.
Red flag: Ownership is vague or leaving requires hidden vendor dependencies.
Pilot design
What would a small paid pilot on our real workflow look like?
Good evidence: A time-boxed pilot uses real data, a named KPI and a clear scale, adjust or stop decision.
Red flag: The vendor pushes directly to a large commitment or offers only a polished demo.
Use the checklist before the proposal decides for you
Score vendors independently before the team compares notes. Require a link, document, reference call or working demonstration for every 2. Then run the strongest candidate through a small paid pilot on a real workflow with a measurable end state.
Read the complete reasoning in our guide to choosing an AI development company, review Mindela's delivery evidence, or get an independent AI architecture and vendor review.