The Sean Ellis test asks users how they would feel if they could no longer use your product. If 40% say "very disappointed", the usual reading is that you have product-market fit.
The threshold comes from Ellis benchmarking roughly 100 startups and noticing that the ones growing well cleared 40% and the ones stalling did not. He never published the underlying data, and researchers at MeasuringU found no peer-reviewed validation of the item, concluding that "there is little compelling evidence to support its promotion for use in practice."
That is a fair criticism, and it is not the practical problem. The practical problem is arithmetic, and it applies even if 40% is exactly the right number.
If you have a result in hand, the product-market fit survey calculator returns its interval for any sample size, along with the survey questions and a template. The rest of this article is where those numbers come from.
What a survey result actually means
Survey results carry a margin of error that depends on how many people answered. Running the numbers for a measured result of 40%, at 95% confidence:
| Responses | True value is somewhere in | Margin |
|---|---|---|
| 10 | 17% – 69% | ±26% |
| 20 | 22% – 61% | ±20% |
| 30 | 25% – 58% | ±17% |
| 50 | 28% – 54% | ±13% |
| 100 | 31% – 50% | ±9% |
| 200 | 33% – 47% | ±7% |
| 400 | 35% – 45% | ±5% |
| 1,000 | 37% – 43% | ±3% |
At 20 responses, a normal number for an early product, measuring 40% means the real figure is somewhere between 22% and 61%. That range contains both "nowhere near fit" and "comfortably past it".
Below 100 responses the test cannot even separate a 40% product from a 30% one. The intervals overlap.
Eight out of twenty and twelve out of twenty are the same result
Turn it around. With 20 responses, here is every outcome whose confidence interval still contains 40%:
| Result | Reads as | True value could be |
|---|---|---|
| 4 of 20 | 20% | 8% – 42% |
| 6 of 20 | 30% | 15% – 52% |
| 8 of 20 | 40% | 22% – 61% |
| 10 of 20 | 50% | 30% – 70% |
| 12 of 20 | 60% | 39% – 78% |
Every row is statistically consistent with a true value of 40%. A founder who gets 6 out of 20 and concludes they lack fit, and a founder who gets 12 out of 20 and concludes they have it, have run the same experiment and learned the same amount, which is nothing.
What it takes to actually clear the bar
To be confident the true value is above 40%, the whole confidence interval has to sit above 40%. That depends on both the sample size and how far above the line you land:
| If you measure | You need |
|---|---|
| 45% | 369 responses |
| 50% | 93 responses |
| 55% | 41 responses |
| 60% | 24 responses |
| 70% | 11 responses |
A result close to the threshold is the expensive one. Squeaking over at 45% requires 369 responses to defend. Landing at 60% is defensible with 24.
This is the useful shape of the rule, and it inverts the usual advice. If your early sample lands anywhere near 40%, you do not have an answer, and collecting a few more responses will not give you one.
What to do instead at small n
The survey is not useless below 100 responses. It is useless as a threshold below 100 responses.
Read the free text, not the percentage. The follow-up questions about what you would use instead and what the main benefit is produce usable signal at n = 15. The percentage does not.
Treat a low score as informative and a borderline score as noise. 3 of 20 has an interval of 5% – 36%, which genuinely rules out 40%. 8 of 20 rules out nothing.
Use behaviour where you can. Retention, repeat purchase and referral are measured on every user rather than on the handful who answer a survey, so the same population gives you a far tighter estimate. This is the same reason the Mom Test prefers what people have already done over what they say they would do.
Ask the question the survey is standing in for. The PMF survey is a proxy for whether people would miss the product enough to pay for it. Willingness to pay and existing spend on alternatives answer that directly, without a sample-size problem, and they are available before the product exists.
Method
The statistics. Wilson score intervals at 95% confidence, which behave better than the normal approximation at small samples and near the extremes. Every figure above is reproducible from the sample size and the proportion alone; no proprietary data is involved. The "responses needed" column is the smallest n at which the lower bound of the interval exceeds 0.40, holding the measured proportion fixed. Requiring the count of respondents to be a whole number pushes the 45% row from 369 to 380, since 45% of 369 is not an integer.
What this does not argue. Nothing here says 40% is the wrong threshold. The benchmark may well be right. The claim is narrower: at the sample sizes founders actually use, a measured result cannot distinguish clearing the threshold from missing it.
One assumption worth stating. These intervals assume responses are a random sample of your users. Survey responses are not random, since the engaged answer more often, which pushes real-world results upward and makes small samples worse rather than better than the table suggests.
Sources
MeasuringU — The Product-Market Fit Item A research review of the Sean Ellis item. The authors trace the 40% threshold to its originator's experience rather than to published data, report finding no peer-reviewed papers describing research with the item, and advise against giving the threshold undue weight in business decisions. Their reasoning includes the same precision problem covered above: a top-box score takes considerable effort to estimate precisely.
The intervals themselves. Wilson score intervals, standard in the statistics literature for binomial proportions. Every number in this article can be reproduced from the sample size and the proportion with any statistics package, and we would rather you checked than took our word for it.
Cite this
A product-market fit survey with 20 responses has a margin of error of ±20%: a measured 40% means the true value lies between 22% and 61%. Confirming a 45% result requires 369 responses. — Scoutr, The 40% Rule and Sample Size