Skip to content
Mindela

Free buyer checklist + CSV

AI vendor evaluation checklist

Compare AI development companies on evidence, not pitch quality. Ask the same 12 questions, score every answer and record the proof before you commit a production budget.

Download the free CSV

Scoring rule

Give each answer 0 for vague or unsupported, 1 for partial with evidence missing, or 2 for specific and independently verifiable. A score of 0 on data handling or ownership should pause the selection, regardless of the total.

Technical fit

01

Have you built this specific kind of AI system before?

Good evidence: A credible answer names a comparable system, the data and constraints, and what the team learned when it failed.

Red flag: A list of technologies with no delivered outcome attached.

02

Who will actually work on the project?

Good evidence: Meet the named engineers and understand each person's responsibility before signing.

Red flag: Technical questions are repeatedly deferred by sales or account staff.

03

What happens if the model provider changes price, removes a model or has an outage?

Good evidence: The architecture isolates model dependencies and the team has tested at least one practical alternative.

Red flag: The plan assumes today's model, pricing and uptime will remain unchanged.

Delivery proof

04

Can we speak with a reference client whose project resembles ours?

Good evidence: The vendor offers a candid conversation about a similar production system, including what went wrong.

Red flag: Only a written testimonial or a reference from an unrelated use case.

05

Which systems are still running in production today?

Good evidence: The answer clearly separates live systems from pilots and explains why some work stopped.

Red flag: Every prototype is presented as a production success.

06

What measurable business KPI did the work move?

Good evidence: You get a baseline, measured result, period and honest attribution caveats.

Red flag: The proof stops at capability or client-satisfaction claims.

Security and operations

07

Where does our data go and who can access it?

Good evidence: The vendor documents residency, retention, isolation, access controls and the compliance posture relevant to you.

Red flag: Your data is safe is the whole answer.

08

How will you monitor quality, failures and drift after launch?

Good evidence: The plan names task-level evals, alerts, owners, review cadence and incident response.

Red flag: Launch is treated as the finish line.

09

How will the system handle ten times the usage or a model change?

Good evidence: Load and unit-cost assumptions are explicit, and the model layer can change without a rewrite.

Red flag: There is no load-test, unit-cost or migration plan.

Commercial terms

10

What is included in the price and what is explicitly excluded?

Good evidence: Build, model, hosting, monitoring, support and change costs are separated in writing.

Red flag: One total hides the operating assumptions and future exclusions.

11

Who owns the code, data pipeline, prompts and model artifacts?

Good evidence: Transfer rights, repositories, documentation and the exit path are unambiguous.

Red flag: Ownership is vague or leaving requires hidden vendor dependencies.

Pilot design

12

What would a small paid pilot on our real workflow look like?

Good evidence: A time-boxed pilot uses real data, a named KPI and a clear scale, adjust or stop decision.

Red flag: The vendor pushes directly to a large commitment or offers only a polished demo.

Use the checklist before the proposal decides for you

Score vendors independently before the team compares notes. Require a link, document, reference call or working demonstration for every 2. Then run the strongest candidate through a small paid pilot on a real workflow with a measurable end state.

Read the complete reasoning in our guide to choosing an AI development company, review Mindela's delivery evidence, or get an independent AI architecture and vendor review.