Most companies that end up with a bad development agency did not choose carelessly. They chose wrong because the information available during evaluation was almost entirely controlled by the agency being evaluated. The agency knew your company, your competitive context, and your budget range before your first call. You knew their website, their case studies, and their sales presentation — all created specifically to make them look good. That asymmetry is the real agency selection problem, and a list of red flags does not solve it unless each flag comes with a way to verify it that the agency cannot stage.
This guide has one structural difference from every other red flag / green flag article: every signal includes a verification method. Looking for the flag is the easy part. Confirming it independently — before the contract is signed — is what the evaluation actually requires.
The red flags — each with a verification method
These are listed in approximate order of how often they predict project failure, not in order of how obvious they are during evaluation. The most dangerous red flags are the ones that do not look like red flags — they look like enthusiasm, confidence, or flexibility. Those are treated separately in the section that follows.
Pricing more than 10% below all competing bids
Any vendor that substantially low-balls their price is either trying to buy the business or misunderstands the scope — and both outcomes end the same way. Deep discounts almost always recover via change orders, scope disputes, or silent quality compromise. The cheaper the initial quote, the more you pay to get to a working product. This pattern is documented explicitly in CIO.com's practitioner guidance: "Any vendor that low-balls their price is either trying to buy the business or doesn't understand the scope" — Mark Ruckman, Sanda Partners. Budget benchmarks for orientation: simple MVP $10K–$50K; basic mobile app $40K–$120K; mid-size application $80K–$250K; enterprise solution $250K+.
Agreeing to every RFP term without questions or exceptions
The provider that says yes to everything usually doesn't know or doesn't care what they are doing — Esteban Herrera, HfS Research (CIO.com). A competent agency will push back on unrealistic timelines, flag scope ambiguities, identify data risks they need resolved before committing, and name at least one term they want to negotiate. Blanket agreement signals either that the team did not read the brief carefully or that they plan to surface problems after the contract is signed when your leverage is lowest.
Bait-and-switch staffing
Senior principals present in the pitch, junior staff assigned to delivery. This is one of the most consistent patterns in failed outsourcing engagements and one of the hardest to detect in advance because it happens after the contract is signed. The pitch team are sales professionals whose job is to win the deal — the delivery team are whoever is available when the project starts. The gap between the two can be significant in technical seniority, domain knowledge, and English proficiency.
Case studies with no verifiable outcome and no available reference
Every agency has case studies. The distinction that matters is between case studies that describe what was built and case studies that document what changed as a result — with a specific metric, a named client, and a reference available. "Built a healthcare platform for a leading provider" tells you nothing. "Reduced manual claims processing time by 62% for a regional health system, with the CTO available for a reference call" tells you everything. Some anonymisation is legitimate; total anonymisation with no verifiable reference is not.
A timeline quoted before data and integrations are assessed
Any firm delivery timeline quoted before the agency has seen your data quality, your integration constraints, and your production environment is sales math. Data readiness — volume, quality, format, labelling, access, and regulatory constraints — is the variable that most frequently determines whether a timeline is feasible at all. Gartner predicts 60% of AI projects without AI-ready data will be abandoned through 2026; the equivalent pattern applies to any project with significant data migration or integration requirements.
No defined post-launch plan in the proposal
A software product is not finished at launch. It requires monitoring, bug fixes, performance tuning, dependency updates, security patching, and feature iteration. An agency proposal that ends at "delivery" with no articulated post-launch structure — monitoring, SLA, support model, cost — is a proposal for a handoff, not a partnership. What happens after week one of go-live reveals more about the agency's actual service model than anything in the sales presentation.
The flags that look like green flags — but predict failure
The most dangerous signals in agency evaluation are the ones that appear positive during the pitch but reverse into failure modes once the contract is signed. These are the patterns that catch experienced buyers — not just first-time outsourcers — because they exploit signals that are genuinely good in an honest partner and genuinely manipulative in a bad one.
- "We're very flexible" — without being specific about what flexibility means contractually. Flexibility is a green flag when it means milestone-based engagement with clear exit terms. It is a red flag when it means no defined scope, no success metrics, and an open-ended time-and-materials contract where scope creep has no ceiling.
- Extremely fast response to your initial brief. A proposal returned within 24 hours of receiving a complex brief has not been read carefully. A strong team asks clarifying questions before scoping. Speed in the sales process predicts speed over substance in the delivery process.
- "We've worked with companies just like yours." Name-dropping similar clients without specifics about what was built, what the outcome was, and whether there is a reference available is a sales tactic, not evidence. Ask for the specific project, the specific outcome, and the specific contact. If these cannot be provided, the similarity claim is not backed by evidence.
- A proposal that exactly matches your stated budget. An agency that quotes precisely at your stated budget maximum has optimised for winning the bid, not for delivering the work. A legitimate proposal is scoped from what the work requires, not from what you said you had available. The exact-budget match is a strong signal of price anchoring rather than independent scoping.
- An impressive portfolio of logos without case study depth. A logo wall of recognisable brands is a sales asset, not evidence of delivery capability. What matters is what was built for those brands, at what scale, with what outcome, and whether a reference is available. Logos without project-specific depth are decorative.
"Price is usually the weakest signal. Focus on how they actually work: discovery process, communication quality, QA structure, and delivery track record."
The green flags — each with a verification method
They push back on your brief
A team that challenges your assumptions, identifies one thing they would scope differently, and names a specific technical risk in your requirements before you have even discussed it has demonstrated independent judgment. This is the single most reliable green flag in an agency evaluation — not because pushback is inherently good, but because it means the team read your brief, thought about it, and prioritised honest engagement over sales compliance. As TalkThinkDo's 2026 evaluation framework notes: "A strong team with adequate technology outperforms a weak team with the latest tools."
Case studies document business outcomes, not just deliverables
The difference between "built a claims-processing platform" and "reduced manual review time by 62% across 14,000 monthly transactions" is the difference between an agency that tracks whether their work produced value and one that tracks whether they shipped. Two case studies with specific, measurable outcomes say more about delivery quality than twenty polished screenshots. The metric, the timeframe, and the scale all matter — and if any of these are missing, the case study is describing activity, not impact.
Engineers answer technical questions directly in the first meeting
The best predictor of what a working relationship will feel like is what the first technical meeting feels like. If the engineers on the call answer specific architecture questions directly — naming trade-offs, flagging constraints, showing opinions — the team is technically confident and not hiding behind account management. If every technical question gets escalated to a follow-up with "our team," the person in the room is not the person who will build the project.
The proposal spends more space on your problem than on their capability
A proposal written for your project rather than from a template will spend the majority of its pages articulating your specific current state, your specific desired outcome, and the gap between them — before deploying technology and methodology as the answer to a problem already clearly understood. If the word count ratio skews more than 50/50 toward the agency's generic capabilities versus your specific situation, you are reading a template, not a proposal. A template means nobody senior read your brief.
They propose a paid discovery phase before the full build
An agency that recommends a bounded discovery engagement before full commitment — producing a detailed scope, architecture recommendation, and realistic timeline — is demonstrating two things simultaneously: confidence that the work product of a discovery will justify full engagement, and commitment to scoping based on evidence rather than assumptions. This is the single best evaluation mechanism available, because it tests the actual working relationship with real stakes before you commit the full budget. Any agency that resists a paid discovery for a complex project is protecting themselves, not you.
They name what AI cannot do for your use case
In 2026, every agency claims AI capability. The agencies that are genuinely AI-mature can tell you what AI is bad at for your specific use case — hallucination risk in high-stakes outputs, latency constraints in real-time workflows, brittleness when edge cases fall outside training distributions, cost unpredictability at scale. Agencies that can describe AI failure modes specifically are agencies that have shipped AI into production and seen those failures. Agencies that describe only AI's capabilities have either not shipped it or are concealing the failures.
Start with agencies already verified on delivery outcomes
TechRadiant evaluates agencies on documented production results — verified case studies with specific outcomes, confirmed client references, and assessed technical depth. The agencies in our index have already passed the scrutiny this guide describes. Share your brief and get matched in 48 hours.
AI maturity in 2026 — a baseline requirement, not a differentiator
Two years ago, AI tooling in a development agency was a differentiator worth asking about. In 2026, it is the baseline. GitHub's Octoverse report found AI-assisted coding adoption crossed 97% among developers in 2025. An agency not actively using AI tools in their workflow is not being more careful or more human — they are delivering more slowly, at higher cost, than every AI-augmented competitor. The right question is no longer "do you use AI?" It is "how specifically, and can you prove it?"
- Names specific tools with specific use cases — code generation vs test generation vs code review vs documentation have different tools and different workflows
- Provides commit-level or PR-level attribution records showing where AI was used in prior projects — not claims, documented records
- Can state measurable outcomes from AI adoption — delivery speed improvement, defect density reduction, review cycle compression
- Has a clear IP position on AI-generated code — you own it, including AI-generated output, confirmed in the contract
- Can name one thing AI tools got wrong on a recent project — and what the recovery looked like
- "We use AI across our entire workflow" — without naming which tools for which specific functions
- Cannot show a project where AI tooling changed delivery speed or quality measurably
- Ambiguous on IP ownership of AI-generated code — who owns code produced by Copilot, Cursor, or equivalent tools when used on your project
- No governance documentation covering how AI tool usage is logged and disclosed to clients
- Uses "AI-powered" in their marketing copy for services that do not involve AI in the development workflow itself
The weighted evaluation scorecard — apply before proposals arrive
Set criteria weights before receiving any proposals — not after. Once proposals are in hand, anchoring bias from the most impressive deck distorts judgment in ways that are difficult to override consciously. The weights below draw from CodeGeeks Solutions' 2026 evaluation framework and SFAI Labs' 2026 partner assessment research, which identified technical expertise and verified portfolio as the criteria most predictive of successful delivery.
| Evaluation criterion | Weight | Score 5 — green flag | Score 1 — walk away |
|---|---|---|---|
| Technical expertise verified via live production case study | 30% | Named production deployment, specific metrics, reference call available within the week | All case studies confidential, no named reference, demo-only examples |
| Domain and industry experience | 20% | Multiple delivered projects in your specific industry or closest adjacent, with compliance depth if regulated | Generic "we work across all industries" positioning with no sector-specific evidence |
| Delivery process and AI maturity | 20% | Documented Agile/DevOps process with CI/CD, named AI tools with specific use cases, measurable AI adoption outcomes | Vague methodology description, no AI tool specifics, no QA structure articulated |
| Communication and team stability | 15% | Named team members with verified availability, engineers answer questions directly, structured async documentation standards | Bait-and-switch signals (senior in pitch, junior in delivery), account manager as only point of contact |
| Commercial structure and outcome alignment | 10% | Milestone-based billing with defined success metrics in the contract, open to paid discovery, clear post-launch structure | Hourly only with no outcome milestones, pricing more than 10% below all others, all-RFP-terms agreement |
| Independent verification quality | 5% | Verified reviews on G2, Clutch, or equivalent with specific project outcomes; reference call completed and positive | Only self-published testimonials, no third-party reviews, reference call declined or avoided |
Score each agency 1–5 on each criterion, multiply by the weight column, and sum the totals. Any criterion scored 1 — particularly technical expertise, compliance, or IP ownership — should function as a gating question rather than a rounding error. A high total score that includes a 1 on production experience conceals prototype-to-production risk that compounds across every other dimension of the engagement.
Build a parallel red-flag log: record every concern from every agency conversation, no matter how minor. After three vendor meetings, the pattern across the log typically reveals which agency poses the lowest residual risk more reliably than any single evaluation conversation. For the equivalent framework applied specifically to AI agency selection — including the 12 red flags specific to AI development and the 3-question screening method — see our complete AI agency briefing and evaluation guide.