How to build a customer health score that works
Most customer health scores are ignored within a year of implementation. Not because the math is wrong, but because a number that moves without explaining why produces no action — so people stop looking at it.
Who this is for: Customer success leaders and RevOps teams building or fixing a health model.
In short
Key takeaways
A health score is a hypothesis about what predicts churn. If you have never tested it against accounts that actually left, you do not know whether it works.
Product usage is over-weighted in most models because it is the easiest data to get, not because it predicts best.
Explainability determines adoption. If a CSM cannot see why a score moved, the rational response is to ignore it.
Start with three signals. Complex models are harder to debug and rarely more accurate.
Trajectory beats level: an account at 65 and falling is more urgent than one at 55 and stable.
What a health score is actually for
A health score exists to answer one operational question: which accounts should someone look at this week. That is all. It is a prioritization device for a team whose attention is the scarce resource, and every design decision should follow from that purpose.
This framing rules a lot out. A score designed to be reported to executives will optimize for looking stable. A score designed to prioritize attention will optimize for surfacing change early, accepting some false positives. You cannot have both from one number, and trying to produces something that serves neither.
Choosing signals: what predicts versus what is available
There is a systematic bias in how health models get built. Teams start from available data — product events, ticket counts, CRM fields — rather than from predictive value, because those sources have clean APIs. The result is a model heavily weighted toward product usage that fails to catch the most common B2B churn causes.
Signals worth including, roughly in order of predictive value: stakeholder stability (is your champion still in role), engagement breadth (how many contacts are actively engaged), engagement recency and reciprocity (are they replying, and how fast), unresolved commitments on either side, value realization (can they name an outcome), support sentiment (not just volume), and product usage depth and breadth.
Note where product usage sits. It matters, and for high-frequency products it matters a lot. But an account can use your product daily and churn when the budget owner — who never logs in — decides the spend is not justified. Usage describes behaviour; it does not describe the decision.
Weighting, and why one global model fails
Login frequency may be decisive for a daily-use collaboration tool and nearly meaningless for a quarterly reporting product. Applying one weighting across both produces noise in each. Weight by segment, and accept that this means maintaining several models rather than one.
Keep weights summing to 100% for interpretability, and keep the count low. A four-signal model where every weight is defensible beats a fifteen-signal model where nobody can explain why support sentiment is 7%. Complexity is not accuracy; it is usually just harder to debug.
Validating the score against reality
This is the step almost everyone skips, and it is the one that separates a working score from a decoration. Take every account that churned in the last twelve months. Look up what each scored 90 days before they cancelled. Plot the distribution.
If churned accounts were predominantly green at 90 days, your model has no predictive value and needs different signals, not different weights. If they were mostly yellow or red, your model works and you can tune the thresholds. If the result is scattered, one segment is probably being scored with the wrong weighting.
Run this annually. Products change, segments shift, and a model calibrated against 2024 churn can be actively misleading by 2026. A score nobody revalidates decays silently, which is worse than having no score at all, because it carries false authority.
Explainability is a product requirement, not a nice-to-have
When a score drops from 82 to 64, the CSM needs to know whether to call the champion, escalate to support, or do nothing. A composite number cannot tell them. So the score must decompose: which signal moved, by how much, and what specific evidence caused it.
The strongest version of this is a score where every movement links to the underlying artifact — the email that went unanswered, the meeting where a concern was raised, the ticket that escalated. That converts the score from a claim into evidence, and evidence is what people act on. It is also what makes the score defensible in a renewal forecast review, where "the model says 64" does not survive scrutiny.
Operating procedure
How to do it, in order
Define the decision the score must improve
Usually: which accounts get attention this week. Write it down, because it constrains every later choice.
Pick three to five signals, weighted toward relationship not just usage
Stakeholder stability, engagement breadth, value realization, support sentiment, usage. Resist adding more until these are validated.
Normalize each signal to 0–100
Define explicitly what 0 and 100 mean for each input, per segment. Ambiguity here produces scores nobody trusts.
Assign weights summing to 100% per segment
Build one model per meaningfully distinct segment rather than a single global weighting.
Backtest against accounts that actually churned
Score last year's churned accounts as of 90 days pre-cancellation. If they were green, change signals — not weights.
Make every score decomposable to its evidence
A CSM must be able to click through from the score to the specific email, meeting, or ticket that moved it.
Track trajectory alongside level, and revalidate annually
Surface direction of travel, not just current position. Re-run the backtest every year as the product and segments change.
Failure modes
Common mistakes
Never testing the model against churn
The single most common failure. A score that was never backtested is an untested hypothesis presented as a fact.
Over-weighting product usage
Available is not the same as predictive. Usage-heavy models miss champion departure, budget shifts, and unrealized value — the most common B2B churn causes.
One global weighting across all segments
What predicts churn for a daily-use tool differs from a quarterly reporting product. A single model produces noise in both.
A score with no explanation
If the CSM cannot see which signal moved and why, ignoring the score is the rational response. Explainability drives adoption more than accuracy does.
Refreshing monthly
The entire value of a health score is lead time. A score updated monthly is a historical record, and batching destroys the thing you built it for.
Showing raw scores to customers
Scores encode internal judgments and incomplete data. A customer who sees themselves marked at risk will redirect the conversation unproductively. Share the findings and the plan instead.
FAQ
Questions, answered
What should go into a customer health score?+
Stakeholder stability, engagement breadth and recency, value realization, support sentiment, and product usage depth. Weight by segment. The most common design error is over-weighting product usage because it is the easiest data to obtain, which leaves the model blind to champion departure and budget-owner decisions.
Why is my customer health score not predicting churn?+
Usually because relationship signals are missing or under-weighted relative to product usage. Validate by scoring last year's churned accounts as of 90 days before they left — if they were green, the problem is which signals you chose, not how you weighted them.
How many signals should a health score use?+
Three to five is a good working range. Complex models are harder to debug, harder to explain to the team expected to act on them, and rarely more accurate. Add signals only after validating that the existing ones predict.
How often should health scores update?+
Continuously, or at minimum daily. Lead time is the entire purpose of a health score; refreshing monthly turns it into a historical record.
Should health scores be shared with customers?+
Generally not as a raw number. Scores contain internal judgments and incomplete information, and a customer seeing an "at risk" label tends to redirect the conversation toward the label rather than the underlying issue. Share the specific findings and the plan.
Keep reading
Related resources
Guides
Customer churn
How to measure, diagnose, and reduce customer churn in B2B SaaS — cohort analysis, leading indicators, the signals that actually predict, and a 90-day operating plan.
Customer success metrics
Which customer success metrics to track and which to drop — retention, expansion, health, and efficiency metrics, what each one is actually for, and how to avoid a dashboard nobody uses.
Customer retention
How to build a customer retention program that works — the retention metrics that matter, onboarding's outsized role, multi-threading, and a renewal process that starts early enough.
Calculators
Your next account move is already in the signals
Knowing what to do is half of it.
Aartha surfaces which accounts need the play — with the cited evidence behind why.