Fake door test examples: six shapes, and three famous ones that don't count
The patterns fake door tests actually take in a live product, what each one can establish, and why Dropbox and Zappos — cited in every example roundup — were not fake door tests at all.
Search for fake door test examples and you get the same three case studies: Dropbox’s video, Zappos’ shoes, Buffer’s landing page. Two of those were not fake door tests, and the confusion matters — because each technique answers a different question, and picking the wrong one gives you a confident answer to something you weren’t asking.
So: the shapes a fake door test actually takes inside a live product, then the famous examples and what they really were. If you want the method itself first, start with how fake door testing works.
Six shapes a fake door test takes
1. The nav item
A new entry in the main navigation — Reports, Automations, Integrations — that leads to an honest “not built yet” panel.
Establishes: whether people go looking for this category of thing at all. Doesn’t establish: what they expected to find behind it. Two users clicking Reports may want completely different things. The trap: navigation is the highest-traffic surface you have, so it inflates early numbers through sheer novelty. Run it past the first week.
2. The plan probe
A fourth tier on the pricing page — usually above your current top plan — with a price and a feature list, leading to a “join the waitlist” form.
Establishes: willingness to consider a price point. This is the only shape that gets anywhere near pricing. Doesn’t establish: willingness to actually pay. A waitlist signup is a fraction of the commitment a card entry is. The trap: it is visible to competitors, to customers on lower tiers who now feel under-served, and to your sales team who will start quoting it. Of all six shapes this is the one to think hardest about before shipping.
3. The integration tile
A logo grid of integrations where some tiles are real and some aren’t, each unbuilt one leading to a “request this integration” flow.
Establishes: relative demand — which is the useful part. You’re not asking do people want integrations, you’re asking which three of these twelve. Doesn’t establish: depth. Someone wanting a Slack integration might want a notification or a full two-way sync, and the tile can’t tell them apart. The trap: partner logos imply a relationship you may not have. Use the partner’s name in text, not their mark, unless you have permission.
4. The export button
A format option that doesn’t exist yet — Export to PDF, Download as CSV, Send to Sheets — sitting alongside the ones that do.
Establishes: demand for a specific, unambiguous output. This is the cleanest shape, because there’s almost no interpretation gap between the label and what the user expects. Doesn’t establish: frequency. One-off export needs and daily-workflow export needs look identical at the click. The trap: it’s often cheap enough to just build, which makes it the shape most likely to fail the “is this test worth more than the feature?” question.
5. The empty state offer
An empty state that offers a capability you haven’t built — “No templates yet. Browse the template library” — where the library doesn’t exist.
Establishes: demand at the exact moment of need, which is the highest-quality signal of the six. These users are stuck right now. Doesn’t establish: anything about your engaged users, who never see empty states. The trap: empty states are disproportionately seen by new users, so you’re measuring a population skewed towards people who haven’t yet decided to stay.
6. The upgrade interstitial
A locked feature behind an upgrade prompt, where the feature doesn’t exist at either tier.
Establishes: almost nothing you should trust, and it is the shape I’d argue against most often. You’re not just showing a door that isn’t there — you’re asking for money to open it. The trap: this is the one that crosses from experiment into something users would reasonably call a dark pattern. If you run it, run it without the payment step.
The famous examples that weren’t fake door tests
Dropbox’s video
Drew Houston posted a demo video of Dropbox working before the sync engine was finished, and used signups to the beta waitlist as the demand signal.
That’s a demand test, not a fake door. There was no door inside a product — there was no product. Nobody was shown a button that lied to them; they were shown a description of something that didn’t exist yet and asked whether they wanted it. Useful, and a genuinely different instrument.
Zappos’ shoes
Nick Swinmurn photographed shoes in local shops, listed them online, and when an order came in, went and bought the pair at retail to ship it.
That’s a Wizard of Oz test — the front end is real, the back end is a human doing the work by hand. The crucial difference from a fake door: the customer got the shoes. Nothing was faked from the buyer’s point of view. That’s why it could measure willingness to pay, which no fake door can.
Buffer’s two-page landing test
Joel Gascoigne put up a page describing Buffer with a “plans and pricing” link. Clicking it led to a second page saying, in effect, you caught us before we’re ready — with a place to leave an email.
This one is genuinely close, and it’s the origin of a lot of the technique’s popularity. But it’s a smoke test: a landing page shown to traffic he had to go and earn, not a door inside a product shown to existing users. It measured demand and his ability to describe the thing compellingly, bundled together.
Three techniques, three different questions:
| Technique | Question it answers |
|---|---|
| Fake door | Do my existing users want this enough to go looking for it? |
| Smoke test | Can I attract anyone to this idea? |
| Wizard of Oz | Will people pay for this, and what does delivering it teach me? |
Write the result down like this
Whatever shape you run, the artefact at the end is four lines, and writing it before the test is what stops the result being reinterpreted afterwards:
- Decision rule: the threshold and the consequence, written before launch.
- The number: click-through rate among users who saw the door, plus the denominator.
- The segments: who clicked, broken down by the two or three dimensions that would change your mind.
- The verdict and the date: what you’re doing about it, and when you’d revisit.
Teams that keep this record accumulate something more valuable than any individual result — an internal benchmark. Your third fake door test is interpretable in a way your first never is, because you finally have a comparable number from your own product rather than a figure from someone else’s blog post.
If discovery in your team keeps producing research nobody acts on, the problem is usually not the technique — it’s that nobody agreed what would change as a result. That’s a fixable habit, and it’s most of what I work on in coaching.