North star metric: the five ways it goes quietly wrong

4 min read

A bad north star doesn't fail loudly — the number goes up and the business doesn't. The test that kills most candidate metrics, and what to do when the honest answer is that you need two.


A north star metric is one number that captures the value your product delivers, with a small set of inputs beneath it that teams can actually move. The idea is sound and the practice is where it comes apart, because a badly chosen north star doesn’t announce itself. It just goes up, quarter after quarter, while the business doesn’t.

The test that kills most candidates

Before adopting any metric, ask this out loud:

If this number doubled and nothing else changed, would the business genuinely be twice as healthy?

Most candidates die here, and the deaths are informative.

Weekly active users — doubled, with no change in retention or willingness to pay, is a much more expensive business. Fails.

Sign-ups — doubled, with the same activation rate, is a bigger leaky bucket. Fails, and it’s the most commonly adopted north star in early-stage companies.

Messages sent — doubled might mean people are communicating more, or that something broke and they’re re-sending. Ambiguous, which is its own kind of failure.

Weekly active teams with at least three collaborators — doubled, that’s harder to fake and much closer to delivered value. Survives.

The pattern: metrics that survive tend to be qualified. Not “users” but “users who did the valuable thing, at the frequency that indicates it stuck.”

The five failure modes

1. It’s chosen for legibility, not truth

Weekly actives is a north star that survives contact with a board deck. That’s precisely why it gets picked, and the legibility is doing the work rather than the accuracy. If your north star was chosen partly because it’s easy to explain to investors, you have a reporting metric, not a north star.

2. It measures activity, not value

The cleanest tell: could a user hit this metric while having a bad time? Session count goes up when people can’t find things. Time-in-app goes up when the interface is confusing. Any metric that rises with friction is measuring effort, not value.

3. It’s not movable by the team that owns it

A north star three layers of causation away from anything a team does produces learned helplessness. If the honest answer to “what did we do that moved it” is “nothing we can identify,” you need better input metrics beneath it — and it’s the inputs, not the star, that teams should be steering by week to week.

4. It hides the segment that matters

An average across a base with two very different populations tells you about neither. If your enterprise accounts and your self-serve users behave nothing alike, one north star will be moved by whichever group is larger, and the smaller one may be where all the revenue is.

5. It doesn’t survive its own success

Some metrics stop being meaningful once you optimise them. Anything that can be gamed by the team will eventually be gamed by the team — not dishonestly, but because that’s what optimisation pressure does. Ask: what’s the cheapest way someone could move this without creating value? If the answer is easy, you’ve found your future problem.

The input layer is where the work is

The north star is a communication device. The inputs are the operating tool, and most of the value comes from getting those right.

For a collaboration product with “weekly active teams with 3+ collaborators” at the top, a reasonable input set might be: teams reaching three members, teams with a shared artefact in week one, week-four retention of activated teams. Each one is ownable by a team, each moves the star, and each fails visibly.

Three to five inputs is the range that works. Beyond that, the causal story stops being legible and you’ve built a dashboard rather than a strategy.

When you honestly need two

Purists say one number. In practice there are two situations where insisting on one causes damage:

Marketplaces and two-sided products. Supply and demand health are genuinely different questions and one number will hide a collapse on the thin side.

A serious quality or trust dimension. A growth star with no counter-metric invites optimising growth at the cost of the thing that makes growth durable. Pair it with a guardrail you commit to not degrading.

Two is defensible. Five is a scorecard, and a scorecard is what you have when nobody was willing to decide.

Reviewing it

Set a date to revisit — annually is reasonable — and actually keep it. North stars go stale as the business changes, and the failure is silent: the metric keeps working as a metric long after it stopped describing value. If your north star has survived a change in business model unchanged, that’s not stability, that’s a metric nobody is examining.


Choosing this well is one of those decisions that’s cheap to make and expensive to have made badly, because a whole organisation orients around it for years. If you’re picking one now, the useful thing is someone outside your company running the doubling test on your candidates — that’s twenty minutes of work and it’s a fair sample of what coaching actually looks like.