Customer Health
Why "the health score went red" is not an answer
Every customer success platform ships a health score. Gainsight, ChurnZero, Totango, Planhat, Vitally, Custify — all of them let you weight a set of inputs, roll them into a number, and colour it red, amber, or green. It is the single most standard feature in the category.
It is also the feature CS leaders quietly stop trusting first.
The pattern is consistent enough to be predictable. The score goes live. For a quarter, people look at it. Then an account goes red for reasons nobody can explain, a CSM overrides it, a green account churns, and within two quarters the score has become something that exists in the QBR deck but does not change what anyone does on a Tuesday.
The problem is not the weighting. Teams re-tune weights constantly and it does not fix this. The problem is that a score is a conclusion without an argument.
What a health score actually computes
Almost every implementation is some version of the same thing: take a handful of measurable inputs, normalise them, multiply by weights, sum.
- Product usage — logins, DAU/MAU, feature adoption
- Support signals — ticket volume, CSAT, time to resolution
- Engagement — email opens, meeting attendance, QBR cadence
- Lifecycle and commercial — time to onboard, NPS, days to renewal
These are all real signals and none of them are wrong. But notice what they have in common: they are all things that were already structured. They are counts, timestamps, and enum fields — the data that was easy to get because someone had already put it in a box.
The reasons B2B accounts actually churn are mostly not in boxes:
- The champion who sponsored the purchase moved to a different team in April.
- The original business case was a cost-reduction target that the customer has quietly stopped tracking.
- On the last QBR someone from finance asked what the month-to-month option looked like.
- Your team committed to an integration in Q2 and it slipped twice.
None of that moves a usage metric. Several of them move usage in the wrong direction for the model — an account doing a careful evaluation of alternatives often has more logins that month, not fewer.
Why "explainability" features do not close the gap
Most platforms have noticed this and shipped something in response: a score breakdown, a contribution chart, an AI-generated summary of the account, a sentiment indicator. These help, but they explain the wrong thing.
A score breakdown tells you which input moved the number. It says: usage contributed −12, support contributed −5. That is arithmetic transparency. What the CSM needs before they walk into a renewal call is causal, sourced explanation: not "usage is down 12 points" but "the two power users who drove 60% of usage both left in May, and here is the meeting where their replacement said they were re-evaluating scope."
The distinction matters because of what happens next. Arithmetic transparency produces a follow-up question — why is usage down? — which sends the CSM to go read call notes and Slack threads. Sourced explanation produces an action. If your explainability feature reliably generates more manual investigation, it has moved the work rather than removed it.
AI summaries have a related failure. A generated paragraph about an account is only as good as what it read, and if it read unstructured notes with no notion of what is current, it will confidently blend a stale fact from March with a live one from July. A summary that is right most of the time is worse than no summary, because you cannot tell which times.
Three properties that make risk defensible
If the goal is a risk signal a CSM will actually act on — and, increasingly, one an AI agent can act on without supervision — it needs three things a weighted score structurally cannot provide.
1. Provenance. Every contributing fact points to its source: the specific call, the specific email, the timestamp. Not "sentiment: negative" but "on the 12 June call, the VP of Ops said the rollout had not delivered the headcount savings they modelled — 00:14:32." A claim that can be verified in one click gets acted on. A claim that cannot gets overridden.
2. Bi-temporality. The system must know when a fact became true and when it stopped being true. This is the single most common failure in AI-generated account context. "The customer is happy with onboarding" and "the customer is frustrated with support" are not contradictory if the first is from February and the second is from July — but a retrieval system with no time model will surface both with equal confidence and produce an assistant that argues with itself.
3. Reconciliation. When new evidence contradicts an existing fact, the old one has to be invalidated, not merely outranked. Otherwise the graph accumulates every version of the truth forever, and the score is computed over a pile that includes things everyone knows are no longer the case.
Together these turn "this account is at risk" into "this account is at risk because three specific things happened, each of which you can go read." That is the difference between a number people override and an argument people act on.
What this changes operationally
The practical effect shows up in four places:
- Renewal prep stops being archaeology. The question "what changed on this account in the last quarter" has an answer with citations, rather than requiring someone to re-listen to calls.
- Escalation gets specific. "Red" escalates to a meeting. "The economic buyer changed in May and has never been briefed" escalates to a briefing.
- Handoff survives turnover. When a CSM leaves, a weighted score transfers nothing. A cited fact graph transfers the actual relationship context.
- Agents become safe to point at customers. This is the part that is about to matter most. Drafting an outreach email is only safe if the system knows what is currently true. Ungrounded agents on top of ungrounded context is how a customer gets an email referencing a problem they solved four months ago.
How to tell which kind you have
You do not need to audit the model. Pick an account your platform currently shows as at risk and ask for the answer to one question:
"Why?"
Grade the response:
- "Usage is down and CSAT dropped." That is the score restating itself. Arithmetic, not explanation.
- "Here is an AI summary of the account." Better — but ask when each claim became true and where it came from. If it cannot say, you cannot rely on it in front of a customer.
- "These four things happened, on these dates, and here are the sources." That is a risk signal worth building a renewal motion on.
Then run the same test on an account showing green that you personally suspect is not. That second test is the more revealing one, because false negatives are what actually cost renewals, and a KPI-weighted score is at its blindest precisely where the risk is relational rather than behavioural.
The underlying point
Health scores were a reasonable design for a world where the only machine-readable customer data was product telemetry and CRM fields. In that world, compressing structured inputs into a number was the best available summary.
That constraint is gone. The conversations where the truth actually lives — calls, emails, QBRs, support threads — are now readable by machines and can be turned into structured, cited, time-aware facts. Once that is possible, compressing everything into a single number stops being a clever summary and starts being lossy for no reason.
The useful output is not a better score. It is the evidence underneath one.
Aartha builds a cited, time-aware Customer Memory Graph from meetings, email, tickets, and CRM activity — so renewal risk comes with the evidence, not just the colour.
Turn customer signals into intelligence.
See how Aartha builds durable customer memory from the tools your team already uses.
Book a demo