Buyer's Guide · AI Agency Evaluation
How to Brief an AI Agency: What to Ask, Demand, and Walk Away From
More than 80% of enterprise AI projects fail to deliver intended business value (RAND Corporation). Most failures are not surprises — they are visible at the proposal stage, in the questions the agency never asks, the timelines quoted before your data is reviewed, and the success metrics that appear nowhere in the contract. This guide was written with no agency on our payroll. That matters more than it sounds.
|
July 7, 2026
|
Updated July 2026
|
15 min read
There is a structural problem with every AI agency evaluation guide on the internet: they are written by AI agencies. The red flags they list are the ones their competitors fail on. The criteria they emphasise are the ones they score well on. The evaluation frameworks they publish are designed to pass, not to genuinely filter. This is not malicious — it is simply the consequence of incentive. A vendor writing a buyer's guide is writing a document that, consciously or not, positions them favourably.
TechRadiant does not build AI systems. We verify the agencies that do. That structural independence — no engagement to win, no proposal to justify — produces a different kind of guide. The 12 red flags in this article will disqualify the majority of AI agencies currently operating in 2026. That is the point. The AI development market exploded from a handful of specialists in 2024 to thousands of firms claiming AI expertise. Most of them are repackaged chatbot development shops or basic LLM wrappers with a polished pitch deck. A structured evaluation process, applied rigorously, is what separates the fraction with real production capability from the majority without it.
80%
of enterprise AI projects fail to deliver intended business value
RAND Corporation
60%
of AI projects without AI-ready data will be abandoned through 2026
Gartner, 2026
3×
higher partner satisfaction from structured evaluation vs ad-hoc selection
SFAI Labs, March 2026
2.5×
higher project failure rate when evaluation is rushed under 3 weeks
SFAI Labs, March 2026
The brief — the document most buyers write wrong
Most organisations approach AI agencies with a technology requirement ("we want an AI chatbot") rather than a business problem ("we need to reduce tier-1 support first-response time from 4 hours to under 15 minutes without adding headcount"). The distinction determines everything that follows: the quality of proposals you receive, the agencies you attract, and whether the project has a measurable objective at its end state.
A brief that describes a technology choice instead of a business outcome attracts agencies optimised for selling, not for solving. The brief below forces a different response — and the quality of how an agency answers it tells you more than anything in their sales presentation.
1. Business problem
State the operational outcome you need — not the technology. What specifically is broken, slow, expensive, or error-prone in your current process?
Example: "Our claims-processing team manually reviews 1,400 documents monthly. Average review takes 3.2 hours per document. We want to reduce manual review time by 60% without increasing headcount."
2. Current state data
Current volume, team size, cycle time, error rate, cost per unit. Agencies that skip scoping your current state are skipping the work that determines feasibility.
Example: "14,000 monthly transactions, 12-person team, $340 per-document cost, 4.2% error rate at current scale."
3. Success metric
What number changes, by how much, in what timeframe, measured how. If you cannot define this, the project has no objective end state and no basis for measuring delivery.
Example: "Review time under 75 minutes per document, error rate below 1%, measured over 90 days post-launch."
4. Data situation
What data exists, where it lives, what format, how clean, who owns access.
Data is where most AI projects fail. Agencies need to assess it before scoping. If they don't ask, that is red flag #1.
Example: "5 years of historical claims data in AWS S3, mixed PDF and structured JSON, partially labelled, owned by the data team."
5. Technical environment
Cloud provider, existing software stack, integrations required, compliance requirements (HIPAA, SOC 2, GDPR, PCI DSS), and security certifications the agency must hold.
Example: "AWS infrastructure, Salesforce CRM integration required, SOC 2 Type II mandatory, data cannot leave our AWS region."
6. Budget + timeline
Not to anchor negotiation, but to filter mismatches. An agency whose minimum engagement is $500K should not spend two weeks in your RFP process if your budget is $80K.
Example: "$120K–$180K total budget. Decision by end of August. Aiming for initial deployment within 14 weeks of contract signing."
Send this brief, not a full RFP. You are testing three things in the initial response: responsiveness, question quality, and how the agency thinks about your problem. An agency that responds with a proposal template and no questions has not read it. An agency that asks a sharp clarifying question about your data structure or compliance requirements in the first response has demonstrated more real capability than any case study on their website.
The 3 questions — 60 seconds, 80% of the picture
Every agency evaluation should include a technical deep-dive with engineers present — not a sales presentation. But before you get there, three questions asked in the first meeting surface the majority of what you need to know about whether an agency has genuine production capability or not.
Question 1 of 3
"Show me one AI system running in production today — not a demo, not a case study — that is live, handling real volume, and that I can speak to a named client reference about."
✓ Acceptable answer
A specific system named, described with production metrics (volume, latency, accuracy, uptime), and a named or identifiable client willing to take a reference call. Bonus: they can show you a monitoring dashboard from a live deployment.
✕ Walk away if
Every example is a prototype, a proof-of-concept, or fully confidential with no verifiable reference. "Built a chatbot" tells you nothing. "Deployed a claims-processing agent handling 14,000 monthly transactions at 94% accuracy" tells you everything.
Question 2 of 3
"Tell me about an AI project that failed or significantly underperformed — what happened and what did you learn?"
✓ Acceptable answer
A specific failure named, a root cause identified (data quality, scope creep, integration complexity), a recovery or pivot described, and what changed in their process as a result. Agencies that have shipped real AI work have failures. This question rewards honesty and penalises polish.
✕ Walk away if
"We haven't had any failures" or vague deflection. Every agency that has shipped production AI has failures — data pipelines that broke, models that drifted, integrations that failed under real load. An agency claiming otherwise has either not shipped or is concealing. Both disqualify.
Question 3 of 3
"What is AI bad at for a use case like mine — specifically?"
✓ Acceptable answer
Specific failure modes relevant to your context: hallucination risk in high-stakes outputs, brittleness when edge cases fall outside training distribution, cost unpredictability at scale, latency constraints in real-time workflows. A good agency tells you what AI cannot do, not just what it can.
✕ Walk away if
"Nothing — AI can do anything" or vague hedging without your specific context named. An inability to identify AI's failure modes for your use case is a fundamental expertise failure, not a confidence signal. You are talking to a salesperson, not an engineer.
The 12 red flags — visible at the proposal stage
The majority of AI project failures are not surprises. They are predictable from the proposal, the first meeting, or the questions an agency does not ask. These 12 red flags are the ones that appear most consistently across the 2026 agency evaluation research — and unlike most lists, they are documented with the specific question that surfaces each one.
1
They quote a timeline before reviewing your data
Walk away immediately
Any firm timeline — "12 weeks to production," "first deployment in 8 weeks" — quoted before the agency has seen your dataset, your API constraints, or your production environment is sales math, not engineering math. Gartner predicts 60% of AI projects unsupported by AI-ready data will be abandoned through 2026. Data readiness is the variable that determines feasibility; an agency that scopes without assessing it is either guessing or planning to cut corners when the data turns out to be what it always is: messier than expected.
Question that surfaces this
"What do you need to see from us before you can give a timeline?" A serious agency names specific data artifacts, access requirements, and infrastructure questions. A red-flag agency gives you a timeline anyway.
2
Every answer is yes
Walk away immediately
AI is powerful but specific. A trustworthy agency will tell you what AI cannot do for your use case — not just what it can. If every question in a scoping conversation gets a confident "yes, AI can do that," you are talking to a salesperson, not an engineer. Real AI teams have opinions and push back. "The retrieval approach you're describing will hit latency issues at your query volume; here's what we'd change" is a green flag. "Sounds great, we can build that" to every item is not. Agreement without qualification surfaces problems later — on your budget, not theirs.
Question that surfaces this
Ask them to name one thing about your brief they would push back on or scope differently. A red-flag agency cannot name one.
3
Case studies are 100% confidential with nothing verifiable
High risk — probe further
Some anonymisation of client case studies is normal and legally expected. But an agency with no named client willing to take a reference call, no live demo accessible, no public production deployment of any kind, and no GitHub repository showing prior work has no verifiable track record at all. The absence of anything verifiable is not evidence of discretion — it is evidence of absence. An agency that has shipped production AI work will have at least one client willing to speak about it and at least one system you can see in operation.
Question that surfaces this
"Can you connect me with one client reference for a 20-minute call?" A red-flag agency hedges, offers a written testimonial instead, or needs to "check with the client" and never follows up.
4
The proposal spends more time on technology than on your problem
Significant concern
A proposal that leads with technology — the models they use, the platforms they have built on, the architecture diagrams of their standard stack — before dedicating substantial space to your specific current state, your operational context, and the measurable gap between where you are and where you need to be, is a proposal written for their pipeline, not for your problem. BotsCrew's May 2026 analysis of failed AI proposals found this pattern in the majority of cases that went on to fail in delivery. The best proposals dedicate most of their space to the problem, and deploy technology choices as the answer to a problem already clearly articulated.
Question that surfaces this
Count the words in the proposal dedicated to your specific situation versus the agency's standard capabilities. If it is more than 50/50 toward their capabilities, something is wrong.
5
Success metrics are vague or absent from the contract
Walk away immediately
"Improved efficiency," "reduced manual work," "better customer experience" — these are not success metrics. A success metric states what number changes, by how much, in what timeframe, measured by what method. Any AI agency should be willing — and required — to define success criteria before the build starts. If the contract does not contain a specific, measurable definition of done, you are paying for effort, not results. An agency resistant to contractual success criteria is an agency that does not expect to hit them.
Question that surfaces this
"What does success look like in the contract — specifically?" If the answer is not a number, a timeframe, and a measurement method, it is not a success metric.
6
Data is treated as an assumption
Walk away immediately
"We'll sort out the data in discovery" is the most expensive phrase in AI development. Data readiness — volume, quality, labelling, access, format, and regulatory constraints — is the primary variable in whether an AI use case is feasible at all and on what timeline. An agency that treats data as something to figure out after contract signing has either never shipped a complex AI system or is intentionally deferring the conversation that would surface a scoping problem. Gartner's prediction that 60% of AI projects without AI-ready data will be abandoned is not a warning about future risk — it describes projects already in progress in 2026.
Question that surfaces this
"What do you need to assess about our data before you can confirm this use case is feasible?" A red-flag agency gives you a generic answer. A serious agency names specific data artifacts, formats, volumes, and access requirements.
The TechRadiant alternative
Skip the evaluation risk entirely — start with outcome-verified agencies
TechRadiant evaluates AI development agencies on documented production deployments, real client outcomes, and technical delivery depth — before you ever speak to them. The agencies in our verified index have already passed the kind of scrutiny this guide describes. Share your brief and get matched in 48 hours.
7
Hourly billing with no outcome milestones
Significant concern
Hourly billing in AI development means the vendor earns more when things go wrong. Scope creep, model iterations, debugging cycles, integration failures — all of these generate hours. You want a commercial structure where the vendor has skin in the outcome. Milestone-based billing tied to defined deliverables, pod-based retainers tied to production metrics, or fixed-bid with clearly defined done criteria all align incentives better than open-ended hourly. Pure hourly is the commercial structure that maximises agency revenue on a struggling project — which is specifically the scenario where incentive alignment matters most.
Question that surfaces this
"What does your billing structure look like, and how is it tied to delivery milestones?" A red-flag agency defaults to hourly. A serious agency proposes milestone-based billing or is open to negotiating it.
8
Senior engineers disappear after the pitch
High risk — verify staffing
The people who sold you the project are frequently not the people who build it. A pitch led by senior engineers or technical founders who then hand off to a junior delivery team is one of the most consistent patterns in failed AI engagements. Ask, specifically, who will be assigned to your project — by name — and what percentage of their time will be dedicated to it. Agencies that cannot answer this before the contract is signed either have not decided yet or plan to staff your project with whoever is available when it starts.
Question that surfaces this
"Who specifically will be building this — by name and seniority — and what is their current availability?" A red-flag agency gives a general answer. A serious agency names people and confirms their current allocation.
9
No answer for post-launch
Significant concern
AI systems are not static. Data changes, which affects model performance. Foundation models update, which can break integrations. APIs evolve, which requires maintenance. Users find edge cases that were not in the training distribution. Production monitoring surfaces patterns that require ongoing adjustment. An agency that sells you the build with no articulated answer for what happens after week one of go-live is selling you a prototype, not a production system. Ask for a specific post-launch support model, an SLA, and a cost structure before the contract is signed.
Question that surfaces this
"What does the first 90 days post-launch look like — monitoring, maintenance, and cost?" A red-flag agency redirects to a separate conversation. A serious agency has a documented answer.
10
IP ownership is ambiguous or defaults to the agency
Walk away immediately
Who owns the model? Who owns the training data derivatives? Who owns the codebase? Who owns the prompt library, the evaluation methodology, and the architecture documentation if you part ways? These questions must have written answers before development begins. Some agencies retain rights to reuse approaches, architectures, or frameworks across clients — which may be acceptable depending on your competitive sensitivity — but ambiguity about IP ownership discovered after contract signing is always resolved in the agency's favour, never yours. If IP is not explicitly addressed in the proposal, assume the default is unfavourable.
Question that surfaces this
"Who owns the model, the codebase, and the approach documentation when the engagement ends?" A red-flag agency defers this to legal. A serious agency answers clearly, in writing, before the contract is final.
11
Compliance is treated as an afterthought in regulated industries
Walk away immediately if regulated
For healthcare (HIPAA), financial services (SOC 2, PCI DSS), or any EU operation (GDPR, EU AI Act effective August 2026), compliance must be proactively addressed in the proposal — not raised when you ask about it. The EU AI Act classifies AI systems used in medical diagnosis, credit scoring, and employment screening as high-risk, imposing conformity assessment, transparency, and monitoring requirements before deployment. An agency that treats these as something to figure out during development is either not working in your industry regularly or is understating the scope of what compliance-grade AI delivery actually requires.
Question that surfaces this
"Walk me through your compliance architecture for a [HIPAA/SOC 2/GDPR] environment." A red-flag agency gives a generic answer. A serious agency names specific controls, certifications they hold, and how data handling is structured.
12
No documented production failures in their history
Significant concern
An agency that presents only success stories has either not shipped anything risky or is concealing failures from you. Every team that has shipped real AI systems in production — with real data, real users, and real performance requirements — has encountered failures: data pipelines that broke, models that drifted, integrations that failed under load, edge cases that were not in the training set. A team that can name a specific failure, describe what they learned, and demonstrate how their process changed as a result is more trustworthy than one with a spotless public record — because the spotless record means the hard work has not been done yet.
Question that surfaces this
"Tell me about something that went wrong on an AI project and what you changed as a result." A red-flag agency cannot name a specific failure. A serious agency names one immediately and explains the recovery.
The evaluation scorecard — how to compare agencies objectively
Weight the evaluation criteria before you receive proposals — not after. Once you have proposals in hand, anchoring bias from the most impressive deck will distort your assessment unless you have pre-committed to a weighted framework. The weighting below is drawn from SFAI Labs' 2026 agency evaluation research, which found technical expertise (30%) and relevant portfolio (25%) to be the strongest predictors of successful AI delivery.
| Evaluation criterion |
Weight |
What good looks like |
What to score 1 (walk away) |
| Technical expertise in your required stack |
30% |
Demonstrated proficiency in RAG, agentic systems, LLM fine-tuning, and MLOps — not claimed, verified through technical deep-dive with engineers |
Cannot walk through architecture decisions from a past deployment |
| Relevant production portfolio |
25% |
At least one verifiable production deployment in your industry or use case category with a named reference willing to take a call |
Every case study is anonymous, confidential, or demo-only |
| Data and scoping process |
20% |
Requires data assessment before confirming timeline; asks specific questions about data format, volume, quality, and access |
Quotes timeline before reviewing any data |
| Commercial structure and outcome alignment |
15% |
Milestone-based billing with defined success criteria in the contract; open to performance-linked terms |
Hourly only with no outcome milestones |
| Compliance and security depth |
10% |
Holds relevant certifications (SOC 2, HIPAA BAA, ISO 27001); proactively addresses compliance architecture for your industry |
Defers compliance to a later conversation |
"The best AI development company is rarely the one with the most impressive demo. A low score on production experience is a prototype-to-production risk that compounds across every other evaluation category."
Azumo — AI Development Company Evaluation Checklist, June 2026
The 7 contract terms you must insist on
Non-negotiable contract terms for any AI development engagement
Defined success criteria in the contract body — not the proposal. What the deliverable looks like, what performance threshold it must hit, and how it is verified. If it is not in the contract, it does not exist as a commitment.
Milestone-based billing tied to delivery and performance. Not hourly. Milestone billing aligns vendor incentives with your outcomes — not with hours spent on a struggling project.
Full IP ownership explicitly transferred to you — model, training data derivatives, codebase, prompt library, evaluation methodology, and architecture documentation. Define what happens to each element if the engagement ends early.
Data handling, access controls, and security terms in writing. Which employees can access your data, how it is stored and processed, what certifications the agency holds, and how data is deleted at engagement end.
Post-launch support obligations with a defined SLA and cost structure. What the maintenance model is, what the response time is for production incidents, and what the monthly operating cost projection covers for the first 12 months.
Knowledge transfer obligations. Specifically: what documentation they commit to producing, how many hours of internal team training are included, and what the overlap period looks like between the agency and your team before handoff.
Exit terms. What you receive if you terminate early, whether you can operate the system independently after the engagement, and what the handoff process looks like. A system you cannot operate without the agency is a dependency, not a delivery.
For teams ready to move from evaluation to selection, TechRadiant's verified AI agency index is the starting point that eliminates most of the evaluation work described in this guide — because the agencies in our index have already been assessed on documented production outcomes, compliance depth, and technical delivery track record, rather than self-reported capability. See our verified AI development companies for the full index, or use our project matching service to get matched based on your specific brief, industry, and compliance requirements.
Frequently asked questions
What should a brief to an AI development agency include?
Six components: (1) The business problem in plain language — not the technology you want, the operational outcome you need. (2) Current state data — volume, team size, cycle time, error rate. (3) The success metric — what number changes, by how much, in what timeframe, measured how. (4) Your data situation — what exists, where it lives, format, quality, access ownership. (5) Your technical environment — cloud provider, stack, integrations, compliance requirements. (6) Budget range and decision timeline — to filter mismatches, not to anchor negotiation. The response quality to this brief is the first filter. An agency that responds with no clarifying questions has not read it carefully enough to scope it.
What are the biggest red flags when evaluating an AI agency?
The 12 non-negotiable red flags: timeline quoted before data is reviewed; every answer is yes; 100% anonymised case studies with nothing verifiable; proposal focuses on technology rather than your problem; success metrics are vague or absent from the contract; data treated as an assumption; hourly billing with no outcome milestones; senior engineers disappear after the pitch; no answer for post-launch maintenance; ambiguous IP ownership; compliance treated as an afterthought in regulated industries; no documented production failures in their history. Any one of these is a significant concern. Two or more is a walk-away signal.
What three questions should you ask any AI agency in the first meeting?
Question 1: "Show me one AI system running in production today — not a demo — that I can speak to a named client reference about." Walk away if every example is a prototype or confidential with no verifiable reference. Question 2: "Tell me about an AI project that failed or underperformed — what happened?" Walk away if they claim no failures exist. Question 3: "What is AI bad at for a use case like mine, specifically?" Walk away if the answer is "nothing" or generic. These three questions surface 80% of what you need to know in under 60 seconds of honest conversation.
How long should the AI agency evaluation process take?
4–6 weeks total: 2 weeks for research and shortlisting (3–5 agencies); 2 weeks for proposals and evaluation; 1–2 weeks for final negotiation and contracting. Rushing under 3 weeks correlates with 2.5× higher project failure rates. Extending beyond 8 weeks typically signals internal stakeholder misalignment, not evaluation complexity (SFAI Labs, March 2026). The evaluation must include a technical deep-dive with engineers present — not just a sales presentation — and reference calls with at least 2–3 clients per finalist.
What contract terms should you insist on with an AI agency?
Seven non-negotiables: defined success criteria in the contract body (not just the proposal); milestone-based billing tied to delivery and performance; full IP ownership explicitly transferred including model, codebase, and documentation; data handling and security terms in writing; post-launch support obligations with a defined SLA and 12-month operating cost projection; knowledge transfer obligations specifying documentation and training; and exit terms defining what you receive if you terminate early and whether you can operate the system independently. Any contract missing these terms should not be signed.
How do you verify an AI agency's claimed expertise before hiring them?
Four reliable methods: (1) Request a named reference call — not a written testimonial — and ask specifically about production performance, timeline, and what happened when something went wrong. (2) Ask to see monitoring dashboards or architecture documentation from a live deployment — agencies with real production experience can show a running system. (3) Hold an engineers-only technical session with no salespeople present and ask specific architecture questions. (4) Use a third-party verified source — independent marketplace platforms that evaluate agencies on documented delivery outcomes, not self-reported capability or paid placement, eliminate the fundamental information asymmetry in agency evaluation. TechRadiant's AI agency index is built on this model.