Concept Testing
Putting a concept in front of real target users to gather structured evidence on whether they would actually use or buy it — before committing to build.
Everyone says your idea sounds great. The only question that matters is what they do when you ask them to commit, and whether you decided in advance what “yes” looks like.
What it is
Stated interest is almost always high. Only revealed commitment tells the truth.
Concept testing is the structured evaluation of a proposed solution with real members of the target audience, to learn whether it resonates, whether it solves a real problem, and whether people would actually use or buy it — before significant resources are committed to building it. Its purpose is to replace internal opinion and debate with external evidence from the people who actually matter.
Its entire value rests on one hard distinction: stated preference versus revealed preference. Ask people whether they would be interested in a concept and they will, overwhelmingly, say yes — warmly and meaninglessly — because agreeing is free, agreeable, and costs nothing. Ask them to actually do something (sign up, provide payment, commit, use it) and the warm agreement collapses into the far smaller number who genuinely want it. That gap between what people say and what they do is where the truth of a concept lives, and surfacing it honestly is what concept testing exists to do.
The discipline that makes the test honest is defining success in advance. Before testing, you set the criterion: the specific level of real commitment the concept must reach to proceed. Without that pre-set threshold, warm verbal interest can always be read as validation after the fact, and the test becomes theater that confirms whatever the team already wanted. With it, the concept faces a verdict it cannot spin.
The gap and the line
What they say, what they do, and the line you drew first.
Click each element to explore what it measures and why it matters. Toggle the scenario to see what it looks like when revealed commitment clears the threshold and when it does not.
When to deploy it
A validation tool — not a substitute for prototyping or for a launch.
Use it when
- →You have a concept concrete enough for real people to react to honestly, and you are before the point of committing to full-scale development or launch.
- →You need to resolve genuine uncertainty about desirability — whether people actually want this — not just whether it functions.
- →You need to choose between competing concepts based on real external response rather than internal preference or executive conviction.
- →You want to replace internal debate with evidence from real target users who have something on the line.
Do not lean on it when
- ×The concept is too vague to react to honestly. Make it tangible first — prototype it — then test. Reactions to vague descriptions are worthless.
- ×The decision has already been made and the test would be theater to justify it. A test whose result cannot change the decision is not a test; it is a confirmation exercise.
- ×You are unwilling to define and honor a success threshold in advance. Without it, the test has no power: any warm outcome can be spun as success.
The honest limit: concept testing measures desirability at a moment, with the fidelity and audience you put in front of it. Tested with the wrong audience, at misleading fidelity, or with leading questions that fish for approval, it produces confident but false signal. Its rigor comes entirely from real target users, an honest commitment ask, and a pre-set bar — remove any of those and it becomes reassurance, not evidence.
How it works
Six moves, from pre-set criterion to honest verdict.
Define what you need to learn, and set the success criteria in advance.
State the specific question and the threshold for a go decision — the level of real commitment the concept must reach. Deciding this beforehand is what makes the result interpretable rather than a vibe to be spun. Commit it to paper and share it with the stakeholders who will act on it before the test runs. A threshold that lives only in someone's head, decided after the results are seen, is no threshold at all.
Recruit real target users.
Test with actual members of the target audience, not colleagues, friends, or convenient stand-ins. The reactions of the wrong audience are worse than no data, because they carry false confidence. Screen recruits against the actual target profile; getting the right people is often the hardest logistical part and the most important.
Present the concept at appropriate fidelity.
Make it concrete enough for honest reactions, without overbuilding. The point is a real reaction, not a finished product. Match the fidelity to what you need to learn: rough for early desirability questions, higher when the reaction depends on details. Overbuilding wastes the cost advantage of testing before commitment; underbuilding produces reactions to something too vague to judge.
Ask for commitment, not opinion.
Do not ask 'would you be interested?', which harvests warm, meaningless agreement. Ask people to actually do something that reveals preference: sign up, provide payment details, place a pre-order, try to use it. A landing-page test with a real sign-up, a fake-door test with a click-through, or a prototype session asking the participant to attempt a real task all produce revealed preference that a survey cannot.
Read revealed demand, not politeness.
Watch for genuine signals: do they lean in, ask when they can have it, try to use it, commit or pay? Treat mild verbal approval as noise, not signal. The gap between warm words and actual behavior is the truth the method exists to surface. Probe the gap — ask why someone said yes but hesitated to commit — the explanation is often the finding that drives the next iteration.
Synthesize across sessions and decide honestly.
Find the pattern across several tests, compare the revealed result to the pre-set threshold, and make the clear call the criterion demands: proceed, refine, or stop — even when the verdict is unwelcome. The hardest part of a concept test is honoring a negative result against executive pressure or team attachment to the concept. The threshold is there precisely for that moment.
Best practices
What good looks like — and what to avoid.
When it goes well
- ✓Success criteria are defined before testing, so the result drives a real decision rather than confirming what the team already wanted.
- ✓Real target users — not internal stand-ins — provide the reactions. The commitment ask is something real: sign-up, payment, attempt to use.
- ✓The team reads behavior and genuine demand signals rather than polite approval. Mild verbal interest is noted but not weighted as evidence.
- ✓The team honors the verdict, including an unwelcome one, and stops or repositions when the evidence says to.
- ✓The test is designed for the right question at the right fidelity — not overbuilt to the point of defeating the cost advantage of testing first.
The mistakes, and how to avoid them
Testing with the wrong audience.
Colleagues, friends, and convenient stand-ins give warm, useless reactions. They want you to succeed, they are not your customer, and they cannot reveal what a real target user would do when their own money or time is on the line. Recruit carefully; screen against the actual target profile.
Asking leading questions that fish for validation.
"You'd love this, right?" or "This would solve your problem, wouldn't it?" harvest the answer you want. Ask for commitment and observe behavior instead. If someone says yes but hesitates to commit, the hesitation is the finding.
Defining success after seeing the results.
The cardinal sin. With no pre-set threshold, any warm outcome can be spun as a win. The discipline of setting the criterion before the test — and sharing it publicly — is what makes the result a verdict rather than a Rorschach test the team reads the way it wants.
Measuring stated interest and calling it validation.
A high "would you be interested?" number feels like proof and means almost nothing. It is the baseline. Warm verbal interest is what you get before any test. Only revealed commitment — what people do when they have to actually do something — counts as signal.
Running a test whose result cannot change the decision.
If the launch is already locked and the test is being run to produce the slide that says "validated," it is theater. Only test when the answer can genuinely move the outcome. Otherwise you are spending time and money building false confidence.
Logistics
Designing the test from commitment ask to honest verdict.
A structured concept test runs from a few days to a couple of weeks, depending on how long it takes to recruit real target users and run enough sessions to see a pattern. The most important logistical act — setting the success threshold in writing before anything else — takes minutes and determines whether the result is a verdict or a vibe.
Write the success threshold down before you start
Commit the go/no-go criterion to paper before designing the test, and share it with the stakeholders who will act on it. A threshold that lives only in someone's head after the fact is no threshold at all. Pre-committing it publicly is what protects the test from being rationalized when the warm stated interest comes in higher than the revealed commitment.
Design a real commitment ask
The heart of a good concept test is the mechanism that converts stated interest into revealed preference: a sign-up with payment details, a pre-order, a deposit, a fake-door click-through to a 'coming soon' page, an actual attempt to use the concept. This ask is where the truth is measured. Design it deliberately for the concept and the question you need to answer.
Recruit carefully and screen for the real target
Getting genuine target users is often the hardest logistical part and the most important. Screen recruits against the actual target profile — not just demographics, but the behavioral and situational criteria that define who the product is actually for. A test run on the wrong people produces confident, false results.
Choose fidelity to fit the question
Match the concept fidelity to what you need to learn. A rough description or simple landing page suffices for early desirability questions; higher fidelity is needed when the reaction depends on details of the experience. Overbuilding wastes the cost advantage of testing before commitment; underbuilding produces reactions to something too vague to generate honest signal.
Run enough sessions to see a pattern
A handful of sessions (five to eight users is common for qualitative signal; larger samples for a quantitative commitment rate) lets you see past individual quirks. Throughout, structure the test around behavior and commitment rather than verbal approval. The participant's desire to be encouraging is the single biggest threat to honest signal — design around it by asking for commitment, not opinion.
AI and this method
AI can simulate a customer who says yes. It can never simulate a customer who actually commits.
Toggle between modes to see where AI genuinely helps a concept test — and the one thing it fundamentally cannot do, which happens to be the whole point.
In-depth example
The same concept. Two approaches. One finds the truth; one cannot.
A consumer goods company has a promising concept for a premium subscription product and confident executives under pressure to launch fast. Before committing, the team wants to know whether real demand exists. Toggle between the traditional test and a hypothetical AI-run version to see why only one of them can surface the gap.
What the team did
PRE-SET THRESHOLD
Defined before testing: 40% of real target customers had to actually sign up and provide payment details for the concept to proceed.
THE COMMITMENT ASK
Instead of asking “would you be interested?”, they asked real target customers to sign up and provide payment details for a pilot — something real on the line.
RECRUITING
Real target customers, not colleagues or friends. Screened against the actual target profile to ensure the reactions came from the right people.
What the test found
A large majority said the concept sounded appealing — warm, encouraging, and meaningless.
Far fewer actually signed up and provided payment. The gap between saying yes and doing it was stark.
Threshold: 40% · Verdict: × FAILS
28% did not clear the pre-set 40% threshold — and because the bar was set before the test, there was no way to spin the warm verbal interest into a green light.
What changed
Probing the gap revealed the concept solved a problem people recognised but did not feel acutely enough to pay a premium for. The team repositioned to a narrower segment that felt the pain far more intensely.
They re-tested against the same 40% threshold. The narrower segment returned a 52% commitment rate — above the bar. They launched to that segment.
The gaps prevented a costly launch into weak demand and pointed to the segment where real demand existed.
Why the threshold mattered
The 76% stated interest, seen without the pre-set threshold, would have looked like strong validation. It was not. The threshold turned the result from a vibe to be interpreted into a verdict that required action. Without it, the warm verbal agreement could have been read as a green light for an expensive launch into weak demand.
Frameworks
Where concept testing shows up.
Concept testing maps to the validation moment in nearly every framework — the point where a concept meets real external evidence before the team commits to build.
Related methods
What to combine with concept testing.
Sources & further reading
The work behind this method.
Testing Business Ideas
David Bland and Alexander Osterwalder (2019)
The definitive guide to structured concept and assumption testing. Bland and Osterwalder provide a library of experiment types — from landing-page tests to fake-door studies to pre-order campaigns — each mapped to the type of assumption being tested and the risk profile of the concept. The emphasis on pre-set success criteria and revealed-preference measurement is the core discipline the book teaches and that makes the whole method trustworthy.
The Mom Test
Rob Fitzpatrick (2013)
On getting honest signal rather than polite encouragement — essential for reading past stated interest. Fitzpatrick's central insight is that people will say anything to avoid being unkind, which means almost all concept feedback is warm, misleading, and useless until you ask for commitment or watch actual behavior. The book teaches how to design questions and interactions that surface the truth people are too polite to volunteer.
Sprint
Jake Knapp, John Zeratsky, and Braden Kowitz (2016)
For the five-user test format used at the end of a Design Sprint — the most widely practiced form of rapid concept testing. Knapp and team are explicit about what Friday is: a test of genuine user response to a prototype, not a show-and-tell. Their discipline around interviewing for honest reactions rather than fishing for approval reflects the same stated-vs-revealed insight that makes concept testing work as a method.