North Star Metric Selection for Product and Growth Teams
Choosing one metric above departmental KPIs aligns teams around shared customer value.

Most fast-scaling companies don't have a data shortage. They have a data surplus, and that surplus lets every team define success on its own terms. Marketing chases acquisition. Sales chases revenue. Product chases feature releases. Support chases resolution time. Dashboards multiply, reports pile up, and somehow the company gets less clear on what it's actually trying to do, not more.
The visible symptom appears in the roadmap meeting. Instead of a conversation about what customers need, it turns into a negotiation over whose metric wins the sprint. A North Star Metric fixes this by sitting above departmental KPIs instead of next to them, giving every team the same question to ask of every proposal: does this move the number that represents real customer value?
What a North Star Metric is and is not
A North Star Metric is the single number that best captures the core value a product delivers to customers. When that number climbs, customers are genuinely getting more out of the product, and the business's long-term health follows close behind. It sits in a specific spot: close enough to the top-line business outcome to matter to leadership, but close enough to daily work that a team can actually move it week over week.
Three categories of bad candidates appear consistently across practitioner guidance. Vanity metrics like total registered users, downloads, or page views climb steadily without telling anyone whether a single customer found lasting value. Pure lagging financials, revenue chief among them, are the output of dozens of upstream decisions rather than a lever any one team can pull directly. And sign-up counts reward the moment someone creates an account while saying nothing about whether they ever came back to use the thing.
Raw Studio draws a clean line between a North Star and a KPI: a North Star predicts where the business is headed, while a KPI reports what already happened inside one function. They're complementary tools doing different jobs, not competing versions of the same idea.
The clearest way to spot a real North Star is to look at what it measures: value delivered, not activity generated. Airbnb tracks nights booked, not signups. Spotify tracks time spent listening, not app installs. Slack tracks messages sent within an organization, a team-level signal of real engagement rather than a count of individual logins. Netflix tracks hours watched per subscriber, a deliberate shift the company made once it recognized that subscription counts said nothing about whether people were actually watching or sticking around.
Four criteria that separate a real North Star candidate from a number that just felt right in a meeting
Picking a North Star is a structured screen, where every candidate has to clear four tests before it earns a pilot.
The first test is customer value: does moving this number require customers to actually get more value, or can the number be pushed up artificially? Statspresso puts this first in its own screening process. If a metric can rise while customers are quietly getting less out of the product, it's the wrong metric no matter how good it looks on a dashboard. Daily active users is the textbook cautionary tale here. DAU can be inflated through notification spam and dark-pattern re-engagement tricks. Usage climbs, resentment builds underneath it, and users eventually leave once they've had enough. The metric went up the whole time the business was hollowing out.
The second test is whether the metric behaves as a leading indicator, moving before revenue does so teams can act on it instead of explaining it after the fact. A well-chosen North Star can flag a retention problem months before it occurs in the churn line of a revenue report. A metric that only tells the story in hindsight isn't a steering tool, it's an autopsy.
The third test is measurability: can the team track it continuously with data already on hand, or data it can reasonably start collecting soon? Raw Studio is specific about the bar here. If hitting that bar requires six months of new instrumentation, the metric isn't a near-term North Star candidate.
The fourth test is team influence: can at least three different teams see their own work reflected in the number? A metric only three people in the whole company can move is a personal KPI for those three people, not something the organization can rally around. Statspresso treats this as a hard filter rather than a nice-to-have: fail more than one of the four questions, and the right move is to keep looking rather than force the fit.
It's fair to push back here and ask whether any single metric can really satisfy all four conditions at once, since choosing a North Star Metric is a structured screen that every candidate must pass before it earns a pilot. That's a reasonable doubt. They're a filter for ruling out bad candidates fast. A number that clears all four is worth piloting, not worth adopting on the spot.
How to generate and validate candidates before committing
Clearing the four criteria tells a team a candidate is plausible. It doesn't tell them the candidate is right, and that's where validation through cohort comparison comes in. Skipping that step means choosing a North Star on gut feeling, which is exactly the trap the whole framework exists to prevent.
The process starts with language, not data. Before anyone looks at a spreadsheet, the team writes one sentence: "Our product helps [customer] achieve [outcome]." The point of the exercise is to name the outcome customers actually want, rather than defaulting to whatever activity happens to be easiest to log. Raw Studio treats this sentence as the foundation the rest of the selection process builds on.
From that sentence, the team generates three to five candidate metrics that could represent the outcome described, and runs each one back through the four criteria from the previous section. Candidates that survive that screen move to the harder test: correlation checks. The team pulls cohorts of customers who hit high versus low levels of each candidate metric, then compares how those cohorts perform on retention and revenue over the following weeks and months. A candidate that doesn't actually predict retention or revenue, no matter how appealing it sounded in the room, gets dropped here.
The last step is a pilot that runs before any launch. The strongest surviving candidate gets reported alongside the existing scorecard for one full quarter before the old metrics get retired. Raw Studio frames stability as a design feature of the whole exercise: swapping the North Star every quarter because growth slowed down usually points to an execution problem, not a measurement problem, and a team that keeps changing the metric is often dodging the harder question of why the number isn't moving. Run well, with the right agenda going in, this entire selection process fits into a single focused working session.
A metric chosen this carefully is still just a number on a dashboard until teams know what to do about it day to day. That's where input metrics come in.
Structuring input metrics for team action
A North Star Metric without input metrics underneath it is a number nobody actually knows how to move. The input tree is what connects a team's daily work to that single top-line goal, giving each group a direct line of sight between what they ship and what the company is steering toward.
The framework works in three layers. The North Star itself sits at the top as the output. Below it sit input metrics, which function as the levers teams can actually pull. Below that sits the work itself: the features, experiments, and research that move those levers. Raw Studio points to Amplitude's North Star Playbook on this point, framing the team's real job as pulling the right levers in the right order rather than staring at the top-line number and hoping it moves.
Practitioner guidance organizes input metrics around four consistent dimensions. Breadth measures how many customers engage with the core value. Depth measures how much value each engaged customer gets out of a given interaction. Frequency measures how often customers come back to get that value again. Efficiency measures how much friction sits between a customer's intent and the outcome they're after. A SaaS product might track weekly active accounts for breadth, features used per session for depth, and login frequency for frequency.
Most frameworks cap this list at three to five input metrics, enough to cover the main drivers of the North Star without spreading ownership so thin that nobody feels responsible for any single number. Each input metric should belong to a named team: breadth to growth, depth to product, frequency to engagement or CRM, so that accountability is built into the structure rather than assumed and then quietly forgotten. Palo Alto Networks narrowed its own candidate inputs down to a small, clearly owned set rather than letting the list sprawl, which is the whole point of keeping the input tree lean: a short list with clear owners actually gets used, while a long list with fuzzy ownership just becomes more dashboard noise.
Data Infrastructure for the North Star Framework
None of this holds together if the CEO and the product team are computing the same North Star from two different queries and getting two different answers. Metric integrity is fundamentally an infrastructure problem, not a definition problem.
The failure mode is familiar to anyone who has sat through a metrics review meeting that devolved into "that's not what my spreadsheet shows." Two teams produce two numbers for the same metric, and the alignment the North Star was supposed to create collapses into a source of conflict instead. Raw Studio is direct about what prevents this: the data pipelines, the product analytics stack, and the reporting tools all have to support accurate, timely tracking of the North Star and its input metrics, and this requirement is a direct driver of how companies choose their BI tools in the first place.
Thoughtbot's experience is a useful, concrete illustration of what happens when infrastructure lags behind ambition. The team computed its North Star metric with a SQL query run against the read-only follower of a production database on Heroku Postgres, using Dataclips to count distinct active repositories per week as a stand-in for "teams". That's a workable stopgap, but it also shows how easily North Star computation can end up riding on infrastructure that was never built for the job.
The industry has been moving to close that gap at the platform level. Databricks announced in May 2025 that Unity Catalog metric views had entered Public Preview, giving organizations a centralized way to define governed, reusable core business metrics once and use them consistently across dashboards, AI agents, and alerts. That kind of governed semantic layer is a direct answer to the "two teams, two numbers" problem, because it removes the room for two different queries to produce two different truths.
A permissions question is easy to overlook here, yet broken identity stitching or ungoverned data access produces measurement noise that misleads the teams acting on the metric. If identity stitching between a user and their account is broken, or if data access is ungoverned, the resulting measurement noise doesn't just create inconvenience. Permissions are not only a security concern here but a metric integrity concern, and getting them wrong quietly undermines everything the selection process in the earlier sections was built to produce.
Self-serve data access and the North Star framework
Even a perfectly governed, perfectly defined North Star Metric is nominal rather than operational if the teams meant to own its input metrics have to file a ticket every time they want to check the number. Adoption of the whole framework depends on broad, team-level access to the data behind it.
That gap between having a dashboard and actually steering by it is wider than most companies admit. After years of self-serve BI promises from the analytics industry, most organizations still see only a minority of employees actively using the tools that were supposed to democratize access. A beautifully designed input-metric tree does nothing for a marketing lead who owns a breadth metric if checking that metric still means waiting on someone else, since the leading-indicator test asks whether a metric moves before revenue does so teams can act rather than explain after the fact.
Real self-serve access changes what's possible day to day. A marketing lead who owns a campaign-specific input metric can watch it move in real time instead of waiting for someone else to pull the number. A product team responsible for a depth or frequency metric can slice its own cohorts without submitting a formal data request and waiting for a turnaround. Mercor's experience working with Hex is a useful data point here: people across the business, including a finance leader, who had never written a line of SQL or Python were still able to work directly with the data they needed.
That's the condition the whole framework rests on. A North Star Metric, screened against four clear criteria, broken into a handful of owned input metrics, and backed by governed, trustworthy infrastructure, still depends on the people closest to the work being able to see the numbers themselves. Without that access, the framework reverts to what it was built to replace: a number on a slide that only the data team really understands.


