Most companies that end up with a bad development agency did not choose carelessly. They chose wrong because the information available during evaluation was almost entirely controlled by the agency being evaluated. The agency knew your company, your competitive context, and your budget range before your first call. You knew their website, their case studies, and their sales presentation — all created specifically to make them look good. That asymmetry is the real agency selection problem, and a list of red flags does not solve it unless each flag comes with a way to verify it that the agency cannot stage.

This guide has one structural difference from every other red flag / green flag article: every signal includes a verification method. Looking for the flag is the easy part. Confirming it independently — before the contract is signed — is what the evaluation actually requires.

66%
of technology projects end in partial or total failure
Standish Group CHAOS Report
64%
of IT leaders outsource development despite knowing the failure rates — because the talent alternative is worse
Aksa Karsa / Medium, 2025
97%
of developers now use AI-assisted coding tools — making AI maturity a baseline, not a differentiator
GitHub Octoverse, 2025
80%
of executives plan to maintain or increase outsourcing investment (Deloitte, 2024) — but are shifting toward verified, long-term partners
Deloitte Global Outsourcing Survey, 2024

The red flags — each with a verification method

These are listed in approximate order of how often they predict project failure, not in order of how obvious they are during evaluation. The most dangerous red flags are the ones that do not look like red flags — they look like enthusiasm, confidence, or flexibility. Those are treated separately in the section that follows.

Pricing more than 10% below all competing bids

Any vendor that substantially low-balls their price is either trying to buy the business or misunderstands the scope — and both outcomes end the same way. Deep discounts almost always recover via change orders, scope disputes, or silent quality compromise. The cheaper the initial quote, the more you pay to get to a working product. This pattern is documented explicitly in CIO.com's practitioner guidance: "Any vendor that low-balls their price is either trying to buy the business or doesn't understand the scope" — Mark Ruckman, Sanda Partners. Budget benchmarks for orientation: simple MVP $10K–$50K; basic mobile app $40K–$120K; mid-size application $80K–$250K; enterprise solution $250K+.

✕ How to verify
If one quote is significantly lower than three others on identical scope, ask them to itemise what is included and what is not. The scope exclusions — QA, DevOps, documentation, deployment, post-launch support — will explain the price gap. A quote with no itemisation is a bid to be avoided.

Agreeing to every RFP term without questions or exceptions

The provider that says yes to everything usually doesn't know or doesn't care what they are doing — Esteban Herrera, HfS Research (CIO.com). A competent agency will push back on unrealistic timelines, flag scope ambiguities, identify data risks they need resolved before committing, and name at least one term they want to negotiate. Blanket agreement signals either that the team did not read the brief carefully or that they plan to surface problems after the contract is signed when your leverage is lowest.

✕ How to verify
Send a brief with at least one deliberately ambitious element — a timeline that is tight, a scope that has a known ambiguity, or a requirement that has an obvious technical risk. Count how many pushbacks you receive. Zero pushbacks is the red flag. One or two specific, reasoned pushbacks is a green flag.

Bait-and-switch staffing

Senior principals present in the pitch, junior staff assigned to delivery. This is one of the most consistent patterns in failed outsourcing engagements and one of the hardest to detect in advance because it happens after the contract is signed. The pitch team are sales professionals whose job is to win the deal — the delivery team are whoever is available when the project starts. The gap between the two can be significant in technical seniority, domain knowledge, and English proficiency.

✕ How to verify
Ask, specifically and in writing before the contract is signed: "Who will be assigned to this project by name, role, and seniority, and what is their current availability?" Then ask to speak with the technical lead — not the account manager — in a 30-minute engineering conversation before committing. An agency that cannot arrange this is signalling that the delivery team does not yet exist.

Case studies with no verifiable outcome and no available reference

Every agency has case studies. The distinction that matters is between case studies that describe what was built and case studies that document what changed as a result — with a specific metric, a named client, and a reference available. "Built a healthcare platform for a leading provider" tells you nothing. "Reduced manual claims processing time by 62% for a regional health system, with the CTO available for a reference call" tells you everything. Some anonymisation is legitimate; total anonymisation with no verifiable reference is not.

✕ How to verify
Ask for one named client reference and schedule a 20-minute call. Ask the reference specifically: was the project delivered on time and budget? What happened when something went wrong? Would you work with them again on something more complex? The answers to those three questions reveal more than any case study document.

A timeline quoted before data and integrations are assessed

Any firm delivery timeline quoted before the agency has seen your data quality, your integration constraints, and your production environment is sales math. Data readiness — volume, quality, format, labelling, access, and regulatory constraints — is the variable that most frequently determines whether a timeline is feasible at all. Gartner predicts 60% of AI projects without AI-ready data will be abandoned through 2026; the equivalent pattern applies to any project with significant data migration or integration requirements.

✕ How to verify
Ask: "What do you need to assess about our data and integrations before you can confirm this timeline?" A serious agency names specific data artifacts, integration dependencies, and infrastructure questions. An agency that gives you a timeline anyway has not actually scoped your project.

No defined post-launch plan in the proposal

A software product is not finished at launch. It requires monitoring, bug fixes, performance tuning, dependency updates, security patching, and feature iteration. An agency proposal that ends at "delivery" with no articulated post-launch structure — monitoring, SLA, support model, cost — is a proposal for a handoff, not a partnership. What happens after week one of go-live reveals more about the agency's actual service model than anything in the sales presentation.

✕ How to verify
Ask: "What does the first 90 days post-launch look like — monitoring, maintenance, and cost?" A red-flag agency redirects to a separate proposal. A serious agency has a documented answer with specific SLAs and a cost structure ready to discuss before you sign.

The flags that look like green flags — but predict failure

The most dangerous signals in agency evaluation are the ones that appear positive during the pitch but reverse into failure modes once the contract is signed. These are the patterns that catch experienced buyers — not just first-time outsourcers — because they exploit signals that are genuinely good in an honest partner and genuinely manipulative in a bad one.

Green-flag appearances that are actually red flags
  • "We're very flexible" — without being specific about what flexibility means contractually. Flexibility is a green flag when it means milestone-based engagement with clear exit terms. It is a red flag when it means no defined scope, no success metrics, and an open-ended time-and-materials contract where scope creep has no ceiling.
  • Extremely fast response to your initial brief. A proposal returned within 24 hours of receiving a complex brief has not been read carefully. A strong team asks clarifying questions before scoping. Speed in the sales process predicts speed over substance in the delivery process.
  • "We've worked with companies just like yours." Name-dropping similar clients without specifics about what was built, what the outcome was, and whether there is a reference available is a sales tactic, not evidence. Ask for the specific project, the specific outcome, and the specific contact. If these cannot be provided, the similarity claim is not backed by evidence.
  • A proposal that exactly matches your stated budget. An agency that quotes precisely at your stated budget maximum has optimised for winning the bid, not for delivering the work. A legitimate proposal is scoped from what the work requires, not from what you said you had available. The exact-budget match is a strong signal of price anchoring rather than independent scoping.
  • An impressive portfolio of logos without case study depth. A logo wall of recognisable brands is a sales asset, not evidence of delivery capability. What matters is what was built for those brands, at what scale, with what outcome, and whether a reference is available. Logos without project-specific depth are decorative.

"Price is usually the weakest signal. Focus on how they actually work: discovery process, communication quality, QA structure, and delivery track record."

Xmethod — How to Choose a Software Development Company, May 2026

The green flags — each with a verification method

They push back on your brief

A team that challenges your assumptions, identifies one thing they would scope differently, and names a specific technical risk in your requirements before you have even discussed it has demonstrated independent judgment. This is the single most reliable green flag in an agency evaluation — not because pushback is inherently good, but because it means the team read your brief, thought about it, and prioritised honest engagement over sales compliance. As TalkThinkDo's 2026 evaluation framework notes: "A strong team with adequate technology outperforms a weak team with the latest tools."

✓ How to verify
Count the number of specific, reasoned pushbacks in their initial response to your brief. One or two is a green flag. Zero means they did not read it. Five or more without any concession means they are difficulty-signalling rather than scoping. The right number is small and specific.

Case studies document business outcomes, not just deliverables

The difference between "built a claims-processing platform" and "reduced manual review time by 62% across 14,000 monthly transactions" is the difference between an agency that tracks whether their work produced value and one that tracks whether they shipped. Two case studies with specific, measurable outcomes say more about delivery quality than twenty polished screenshots. The metric, the timeframe, and the scale all matter — and if any of these are missing, the case study is describing activity, not impact.

✓ How to verify
For each case study they reference, ask: "What specifically changed after you delivered this — what number moved, by how much, over what timeframe?" An agency that can answer this without hesitation has been measuring outcomes. One that deflects to the technology used or the features built has not.

Engineers answer technical questions directly in the first meeting

The best predictor of what a working relationship will feel like is what the first technical meeting feels like. If the engineers on the call answer specific architecture questions directly — naming trade-offs, flagging constraints, showing opinions — the team is technically confident and not hiding behind account management. If every technical question gets escalated to a follow-up with "our team," the person in the room is not the person who will build the project.

✓ How to verify
Ask a specific technical question relevant to your project in the first meeting — your data migration approach, your preferred architecture pattern, your integration constraint. A green-flag team answers it directly and shows judgment. A red-flag team defers to a technical document or a follow-up.

The proposal spends more space on your problem than on their capability

A proposal written for your project rather than from a template will spend the majority of its pages articulating your specific current state, your specific desired outcome, and the gap between them — before deploying technology and methodology as the answer to a problem already clearly understood. If the word count ratio skews more than 50/50 toward the agency's generic capabilities versus your specific situation, you are reading a template, not a proposal. A template means nobody senior read your brief.

✓ How to verify
Redact your company name from the proposal and ask: could this proposal have been written for a different client in a different industry? If yes, it is a template. A genuinely custom proposal contains at least three references to specifics from your brief that could not have been inferred from the generic RFP.

They propose a paid discovery phase before the full build

An agency that recommends a bounded discovery engagement before full commitment — producing a detailed scope, architecture recommendation, and realistic timeline — is demonstrating two things simultaneously: confidence that the work product of a discovery will justify full engagement, and commitment to scoping based on evidence rather than assumptions. This is the single best evaluation mechanism available, because it tests the actual working relationship with real stakes before you commit the full budget. Any agency that resists a paid discovery for a complex project is protecting themselves, not you.

✓ How to verify
If they do not propose a discovery phase, ask whether they would be willing to structure one. An agency with genuine confidence in their delivery capability will say yes readily. One that cannot define what the discovery deliverable looks like — specific scope document, architecture rationale, timeline, and cost estimate — does not have a real discovery process.

They name what AI cannot do for your use case

In 2026, every agency claims AI capability. The agencies that are genuinely AI-mature can tell you what AI is bad at for your specific use case — hallucination risk in high-stakes outputs, latency constraints in real-time workflows, brittleness when edge cases fall outside training distributions, cost unpredictability at scale. Agencies that can describe AI failure modes specifically are agencies that have shipped AI into production and seen those failures. Agencies that describe only AI's capabilities have either not shipped it or are concealing the failures.

✓ How to verify
Ask: "What is AI bad at for a use case like mine?" A green-flag agency gives a specific answer relevant to your context within 30 seconds. A red-flag agency gives a vague general answer or says nothing AI cannot do. The quality of the failure-mode answer is one of the most reliable indicators of genuine AI expertise.
Skip the evaluation risk

Start with agencies already verified on delivery outcomes

TechRadiant evaluates agencies on documented production results — verified case studies with specific outcomes, confirmed client references, and assessed technical depth. The agencies in our index have already passed the scrutiny this guide describes. Share your brief and get matched in 48 hours.

AI maturity in 2026 — a baseline requirement, not a differentiator

Two years ago, AI tooling in a development agency was a differentiator worth asking about. In 2026, it is the baseline. GitHub's Octoverse report found AI-assisted coding adoption crossed 97% among developers in 2025. An agency not actively using AI tools in their workflow is not being more careful or more human — they are delivering more slowly, at higher cost, than every AI-augmented competitor. The right question is no longer "do you use AI?" It is "how specifically, and can you prove it?"

AI maturity — what green looks like vs what red looks like
Green — genuine AI maturity
  • Names specific tools with specific use cases — code generation vs test generation vs code review vs documentation have different tools and different workflows
  • Provides commit-level or PR-level attribution records showing where AI was used in prior projects — not claims, documented records
  • Can state measurable outcomes from AI adoption — delivery speed improvement, defect density reduction, review cycle compression
  • Has a clear IP position on AI-generated code — you own it, including AI-generated output, confirmed in the contract
  • Can name one thing AI tools got wrong on a recent project — and what the recovery looked like
Red — AI washing
  • "We use AI across our entire workflow" — without naming which tools for which specific functions
  • Cannot show a project where AI tooling changed delivery speed or quality measurably
  • Ambiguous on IP ownership of AI-generated code — who owns code produced by Copilot, Cursor, or equivalent tools when used on your project
  • No governance documentation covering how AI tool usage is logged and disclosed to clients
  • Uses "AI-powered" in their marketing copy for services that do not involve AI in the development workflow itself

The weighted evaluation scorecard — apply before proposals arrive

Set criteria weights before receiving any proposals — not after. Once proposals are in hand, anchoring bias from the most impressive deck distorts judgment in ways that are difficult to override consciously. The weights below draw from CodeGeeks Solutions' 2026 evaluation framework and SFAI Labs' 2026 partner assessment research, which identified technical expertise and verified portfolio as the criteria most predictive of successful delivery.

Evaluation criterion Weight Score 5 — green flag Score 1 — walk away
Technical expertise verified via live production case study 30% Named production deployment, specific metrics, reference call available within the week All case studies confidential, no named reference, demo-only examples
Domain and industry experience 20% Multiple delivered projects in your specific industry or closest adjacent, with compliance depth if regulated Generic "we work across all industries" positioning with no sector-specific evidence
Delivery process and AI maturity 20% Documented Agile/DevOps process with CI/CD, named AI tools with specific use cases, measurable AI adoption outcomes Vague methodology description, no AI tool specifics, no QA structure articulated
Communication and team stability 15% Named team members with verified availability, engineers answer questions directly, structured async documentation standards Bait-and-switch signals (senior in pitch, junior in delivery), account manager as only point of contact
Commercial structure and outcome alignment 10% Milestone-based billing with defined success metrics in the contract, open to paid discovery, clear post-launch structure Hourly only with no outcome milestones, pricing more than 10% below all others, all-RFP-terms agreement
Independent verification quality 5% Verified reviews on G2, Clutch, or equivalent with specific project outcomes; reference call completed and positive Only self-published testimonials, no third-party reviews, reference call declined or avoided

Score each agency 1–5 on each criterion, multiply by the weight column, and sum the totals. Any criterion scored 1 — particularly technical expertise, compliance, or IP ownership — should function as a gating question rather than a rounding error. A high total score that includes a 1 on production experience conceals prototype-to-production risk that compounds across every other dimension of the engagement.

Build a parallel red-flag log: record every concern from every agency conversation, no matter how minor. After three vendor meetings, the pattern across the log typically reveals which agency poses the lowest residual risk more reliably than any single evaluation conversation. For the equivalent framework applied specifically to AI agency selection — including the 12 red flags specific to AI development and the 3-question screening method — see our complete AI agency briefing and evaluation guide.