If you're an HR leader who has bought a digital mental health benefit in the last few years, this rhythm will be familiar. You buy the product. Three months in, the vendor sends you a deck. The deck has aggregate PHQ-9 and GAD-7 numbers. You ask whether the deltas are meaningful. The vendor says they are. You schedule the next QBR.
There's a problem buried in this rhythm. The metrics on the deck are lagging indicators. By the time they shift, the conditions that produced the shift have been in place for weeks or months. If something is going wrong with your benefit, your standard reporting will not tell you until well after it started going wrong.
This isn't a complaint about the PHQ-9 or GAD-7: they're good instruments. They're validated, they're widely used, and they have decades of clinical evidence behind them. The problem isn't the tool. It's the cadence.
What lagging indicators actually tell you
A lagging indicator confirms what has already happened. It's useful for understanding outcomes, but it's not useful for changing them in time.
In mental health, the standard lagging indicators are symptom severity (PHQ-9 for depression, GAD-7 for anxiety) and user satisfaction (NPS and PMF). They're assessed with surveys, and they describe the past, not the future.
This works for an annual benefits review. It doesn't work for the question, "is this working right now?"
What leading indicators look like
A leading indicator tells you, before the lagging indicator would, whether the underlying conditions for a good outcome are in place. In sales, pipeline coverage is a leading indicator for revenue. In retention, product engagement is a leading indicator for renewal. In mental health, leading indicators are session-level signals that predict what we expect a survey to later confirm.
Here's what we've found in our work:
Behavioral activation — whether a session produces movement away from behavioral inertia and toward engagement in valued or pleasant activities — is a leading indicator for PHQ-9 improvement. Activation is the mechanism through which evidence-based depression treatment actually works. If depressed users are engaging in sessions that produce activation, we tend to see improvement in their PHQ-9 scores down the road.
Willingness — whether a session produced movement toward reduced avoidance and increased willingness to face discomfort — is the parallel construct for anxiety. Avoidance is both the maintenance mechanism for anxiety and the primary target of evidence-based anxiety treatment. When anxious users engage in sessions that produce willingness, we tend to see improvement in their GAD-7 scores.
Resonance—whether a session landed with users emotionally—is a leading indicator for NPS. When resonance starts trending up across users, the survey numbers follow within weeks.
Why the category defaulted to lagging indicators
The category defaulted to lagging indicators for two reasons. The first is that they were what the industry knew how to measure: validated instruments collected via survey, comfortable for clinicians, easy to put on a slide.
The second is harder to admit. Leading indicators show how the product's really performing, warts and all. They make it harder to hide where the product isn't producing change. If a vendor isn't tracking their product's leading indicators, or worse - if they don't know what their leading indicators are, they will have no way to catch that things aren't working until the damage is already done.
Lagging-only reporting is conventional, not dishonest — it's what the industry knew how to build. The question is whether conventional is still good enough
What to ask your vendor
If you want to know whether your vendor would catch it if the benefit weren't working, here are the questions that will tell you.
What do you measure session by session, before the survey cycle? Vendors should be able to answer this with specifics, not by saying "we look at engagement."
What's your leading indicator for symptom improvement? If they can't name one and explain how it's related to clinical improvement, they're flying on retention data and hope.
How often do you review session-level data, and how do you act on it? Daily and weekly cadence is reasonable. Monthly is the floor.
Can I see the leading-indicator dashboard for my population? If they can't show you, they probably don't have one.
The shift the category is going to make
The industry is going to move toward session-level measurement. It has to. The lagging-indicator era worked when the assumption was that mental health benefits would be evaluated annually and renewed on broad utilization numbers. That assumption is breaking. Employers are getting sharper, brokers are getting sharper, and the products that produce session-level evidence are going to be the ones that survive.
This isn't about replacing PHQ-9 and GAD-7. It's about adding the layer in front of them — the daily, session-level read that tells you whether the conditions for improvement are present before you wait six weeks to find out.
That's the difference between watching the rear-view mirror and driving.