Mental Health
September 23, 2026|Last updated September 9, 2026

Your PHQ-9 is improving.

Are your clients?

Written by Audrey Smith

At a glance

  • The PHQ-9 scores how often nine symptoms occurred in the past two weeks. Its functioning question is unscored, so the impact on a client's daily life never reaches the total.
  • Scores and client experience usually move together. When they split, the gap is clinical information worth exploring in session, and researchers have documented it in both directions.
  • Measurement-based care helps when scores are collected, shared with the client, and acted on together. Fewer than 20% of behavioral health clinicians use it routinely, mostly for structural reasons like time and payer paperwork.
  • Three questions close most of the gap: what would better look like in your life, does this trend match what you're noticing, and what isn't the scale asking about.
Your PHQ-9 is improving. Are your clients?

Your 3:00 client’s PHQ-9 has dropped from 16 to 9 since intake. A drop that size counts as “response,” the research term for meaningful improvement on the scale. Then, on her way out, she mentions she still hasn’t called her sister back and still dreads Sunday nights. Both things are true at once, because the PHQ-9 answers a narrower question than the one you’re treating: it asks how often nine specific symptoms showed up over the past two weeks. Whether someone feels like themselves again is a bigger question, and research from the past few years keeps finding gaps between the two.

You don’t have to choose between trusting the scale and trusting your client. The useful move is knowing exactly what the score covers, where it runs out, and what to ask next.

What the PHQ-9 actually measures

The Patient Health Questionnaire-9 asks a client to rate how often each of the nine depression symptoms bothered them over the past two weeks, from “not at all” to “nearly every day.” The answers add up to a score from 0 to 27. It was built and validated as a screening and severity tool in primary care, and a 2021 systematic review supports it in that role.

Two design details matter for how you read it. First, the scale measures frequency, so a symptom that shows up less often but hits just as hard can pull the score down without the client feeling much relief. Second, the questionnaire’s tenth item, which asks how difficult these symptoms have made daily life, sits outside the score entirely. A client’s work, relationships, and sense of self can stay stuck while the total falls.

That’s the design working as intended. A 2023 analysis by the nonprofit research organization Sapien Labs found that fewer than half of people scoring in the “severe” range (20 or higher) rated their symptoms’ life impact at the most severe level. Score severity and life disruption are related, and they’re far from the same thing.

When the score and the client disagree

The mismatch runs in both directions. A score can improve while the client reports feeling no different, often because the items driving the drop (sleep, appetite, energy) matter less to that client than the ones holding still (low mood, little interest in things they love). And a client can report real change the score never registers. Research on clinically meaningful change exists precisely because statistically detectable movement on a symptom scale and improvement a client can feel are two different thresholds.

What clients count as recovery is also broader than nine symptoms. A 2024 qualitative meta-analysis in The Lancet Psychiatry found clients name outcomes like deeper self-understanding, a stronger sense of agency, and more social connection. A 2022 review of what matters to people with persistent depression reached similar conclusions. None of those appear on the PHQ-9.

You’ll often sense this gap before any instrument shows it. The hard part is pinning it down consistently, client after client. In the 2026 Practice Success Report: The five domains of a thriving therapy practice, 42.9% of marriage and family therapists we surveyed named ‘defining improvement consistently’ as an obstacle to tracking outcomes, compared with 29.9% of counselors.

Why the score carries so much weight

The PHQ-9 is free, takes about two minutes, and produces a number that travels cleanly between clinicians, practices, and payers. That’s why standardized depression measures keep showing up in payer quality programs and value-based contracts. When a number moves that easily through the system, it starts standing in for outcomes it was never designed to capture. That pressure comes from how the system pays for care, and it lands on your desk either way.

To be fair to the measure, the evidence for using it well is solid. A 2021 multilevel meta-analysis found that tracking progress and feeding it back into treatment modestly improves outcomes, reduces deterioration, and cuts dropout, with the clearest gains for clients whose treatment is going off track. The American Psychological Association’s practice guidelines on measurement-based care, first released in 2022, describe it as three moves: collect client-reported outcomes routinely, share the results with the client, and act on them in the context of your clinical judgment and the client’s own account. The score informs the work. It doesn’t run it.

Most practices aren’t doing this, and the reasons are structural. Implementation research focused on solo and small-group practices estimates fewer than 20% of behavioral health clinicians use measurement-based care routinely, and only about 5% follow an evidence-based schedule. The barriers it names are time, workflow, and competing demands, the same forces behind the documentation backlog in mental and behavioral health.

Three questions to ask alongside the score

You don’t need new instruments or a research protocol to close the gap between the number and the person. You need three questions, asked at the right moments.

1. “What would better look like in your week?”

Ask at intake, and again when treatment goals shift. You’re listening for two or three concrete markers the client chooses: answering the phone when a friend calls, cooking dinner instead of skipping it, going back to choir. Write them into the treatment plan next to the baseline score, so both live in the client’s progress notes and treatment plan from day one. This is the client-defined half of outcome tracking that symptom scales can’t supply.

2. “Here’s your trend. Does it match what you’re noticing?”

Sharing the score is a core part of measurement-based care as the APA guidelines define it, and it’s the step that turns a form into a conversation.

If the score is down but the client feels flat, look at the item level together: which symptoms moved, and are those the ones the client cares about? 

If the score is flat but the client reports change, check their markers from question one, because their gains may live there.

Either way, you have a session topic, a documentation note, and a reason to adjust the plan with the client in the room.

3. “What is this scale not asking you about?”

Functioning, relationships, work, and meaning. If you want a structured supplement, the World Health Organization’s WHODAS 2.0 is a free measure of day-to-day functioning. If you’d rather keep it conversational, ask directly and document what you hear. The APA guidelines describe measurement-based care as covering symptoms and functioning; a symptom scale alone covers half of that.

Where to start

Pick one client whose score has improved over the past two months. In your next session, show them the trend and ask whether it matches what they’ve noticed. Write their answer, plus their own two or three markers of better, into the treatment plan next to the number. If you run your practice on TheraNest® by Ensora Health, the note and the treatment plan already live in the same client record, so keeping the score and the client’s account side by side takes minutes. If you’re deciding which numbers deserve attention across the whole practice, start with the KPIs worth tracking in a solo practice, then give every clinical score a client-defined marker to sit beside.

Frequently asked questions
Is the PHQ-9 a screening tool or an outcome measure?
Arrow Icon
It was built and validated as a screening and severity measure in primary care. It’s widely used for outcome monitoring because it’s brief, free, and consistent, and it performs reasonably there as long as it isn’t the only measure of success. Pair it with functioning measures or client-defined goals for a fuller picture.
What should I do when the PHQ-9 improves but my client says they don’t feel better?
Arrow Icon
Take both seriously. Review the item-level changes together, revisit the markers your client named at intake, and document the discrepancy along with your clinical reasoning. Researchers expect this mismatch, which is why clinically-meaningful-change thresholds exist. It’s also a strong prompt to revisit the treatment plan with the client.
Does measurement-based care actually improve outcomes?
Arrow Icon
Yes, modestly, when it’s done as designed. The 2021 meta-analysis and a 2023 research review of routine outcome monitoring both find better outcomes and fewer dropouts, with the largest effect for clients going off track. The benefit comes from discussing the feedback with the client, so administering a measure without talking about it forfeits most of the value.
What can I use alongside the PHQ-9?
Arrow Icon
WHODAS 2.0 for functioning, the GAD-7 when anxiety is part of the picture, and client-defined goal tracking from question one above. None of these require new tools, just a decision with each client about what improvement means for them.