July 19, 2026
12 min read
By Albert Wong, PhD · Clinical Psychologist
The short answer
Give the PHQ-9 (depression, 9 items, 0–27) and GAD-7 (anxiety, 7 items, 0–21) at intake, then every two to four weeks. A drop of 5 or more on the PHQ-9 is a commonly used threshold for clinically meaningful change; smaller wiggles are usually noise. Any positive response on PHQ-9 item 9 — the suicidality item — gets a direct conversation the same day, every time. Share the trend line with the client, put the scores in the note as objective data, and where plans cover it, bill CPT 96127 per administered instrument. They're screening and monitoring tools, not diagnostic tests — but used this way, they'll catch a quiet decline months before the room does.
You know your client is better. You can feel it in the room — the shoulders have come down, the jokes have come back, the silences are comfortable instead of heavy. And your feel for the room is real clinical data. Here's the problem: "I can feel it" doesn't survive an insurance review. It doesn't catch the quiet slide in the client who's gotten very good at performing okay for fifty minutes. And it can't show the client the distance they've traveled the way a line on a chart can — a line that starts at 19 and ends at 6 and belongs to them.
That's the whole case for measurement-based care, and it's worth being precise about what it is and isn't. It is not a replacement for clinical judgment, any more than a depth sounder replaces the person at the helm. Sailors call navigating by feel and experience dead reckoning, and good sailors are genuinely good at it — but they still take soundings, because the water lies. Measurement-based care just means routinely administering a brief symptom measure, looking at the result, and letting it inform — not dictate — what you do next. A robust literature backs this up: across settings and populations, routinely measuring outcomes and feeding the results back to clinician and client is consistently associated with better outcomes, and especially with earlier detection of the clients who aren't responding — the ones dead reckoning misses, because they're the ones who've learned to look fine.
You could build a measurement practice from dozens of instruments. Almost everyone builds it from two, and for good reasons: they're brief, they're validated for both screening and monitoring over time, they're free and in the public domain (no licensing fees, no permission letters), and every payer, collaborating physician, and utilization reviewer in America already knows how to read them.
The PHQ-9 is nine items covering depression symptoms over the past two weeks, each scored 0–3 ("not at all" to "nearly every day"), for a total of 0–27. The GAD-7 is seven items on anxiety symptoms over the same window, same 0–3 scale, total 0–21. A client can do both in under five minutes. The severity bands:
| Severity | PHQ-9 (0–27) | GAD-7 (0–21) |
|---|---|---|
| Minimal | 0–4 | 0–4 |
| Mild | 5–9 | 5–9 |
| Moderate | 10–14 | 10–14 |
| Moderately severe | 15–19 | — |
| Severe | 20–27 | 15–21 |
On the GAD-7, a score of 10 or above is the commonly used screen-positive threshold — the point at which anxiety warrants a closer clinical look. And now the caveat that belongs in bold in your mind: these are screening and monitoring tools, not diagnostic instruments. A PHQ-9 of 16 does not diagnose major depressive disorder; it tells you the client is endorsing significant symptoms and the clinical interview — your differential, your history-taking, your judgment — does the diagnosing. Grief elevates these scores. Poor sleep from a newborn elevates them. Thyroid problems elevate them. The score is a sounding, not the chart.
One item deserves its own paragraph. Item 9 of the PHQ-9 asks about thoughts of being better off dead or of hurting oneself. Any positive response — a 1, a 2, a 3, anything above zero — needs a direct follow-up conversation, that day, every time. This is standard clinical practice, not an exotic protocol: you ask about the thoughts, assess intent, plan, and means, document your assessment and your reasoning, and safety plan where indicated. Most positive item-9 responses turn out to be passive ideation without intent, and the conversation takes five minutes. But the item exists so that the question always gets asked, including of the client who would never have raised it themselves — and a documented follow-up is also exactly what a reviewer will look for if they ever pull the chart. If you use digital administration between sessions, know how you'll be alerted to a positive item 9 and what your response window is. A screener you don't look at until next Tuesday is not monitoring.
The rhythm that works in private practice is simple: baseline at intake, then every two to four weeks. Monthly is fine for stable, longer-term work; every two weeks fits an active treatment phase or a new medication trial you're tracking alongside a prescriber. What about every session? You can — some clinicians do, and for brief, protocol-driven work it has its place — but for most practices, weekly administration mostly teaches you to chase noise. These are two-week symptom windows filled out by a human being who slept badly, had a fight with their sister, or is feeling optimistic because it's finally sunny. Scores wobble. If you measure every week, you will be tempted to steer with every wobble, and that's how you end up zigzagging across the harbor responding to chop instead of current.
So what counts as a real change? A commonly used heuristic: a PHQ-9 change of 5 points or more is treated as clinically meaningful — reliable movement rather than measurement noise. A client who goes from 18 to 16 has, most usefully, stayed the same. A client who goes from 18 to 11 has genuinely moved. The same logic applies in the wrong direction, and that's where measurement earns its keep: three administrations over eight weeks that read 14, 15, 16 are telling you something the pleasant, cooperative fifty minutes in between are not.
When the trend says non-response, the score hasn't failed you — it's handed you the most useful hard conversation in therapy: "We've been at this eight weeks and your numbers aren't moving. I don't think that's you failing; I think it's the plan needing to change. Let's talk about what we adjust." From there the options are the classic ones — change the approach or the dose of therapy, seek consultation, add a medication evaluation, or refer to a better-fitting level of care. None of that is comfortable. All of it beats discovering at month nine what the instruments were showing at week eight.
A score that goes into a drawer is clerical work. Three uses turn it clinical:
Here's the part almost nobody tells solo practitioners: the screener you're already giving is often billable. CPT 96127 — brief emotional/behavioral assessment with a standardized instrument — is covered by many commercial plans, billed per instrument administered, scored, and documented, alongside your regular session code. Give a PHQ-9 and a GAD-7 at the same visit and many plans allow a unit for each. Reimbursement is modest and coverage varies enough that you should verify with your payers rather than budget on it — but it's real money for work you were doing anyway, and it compounds across a caseload and a year. The details live in our CPT codes cheat sheet.
It isn't skepticism. Most therapists are persuaded by the evidence within a paragraph. What kills it is the clipboard: printing the form, remembering to hand it over, watching the client fill it out on your session time, hand-scoring nine items while they watch, transcribing the number into the note, and then — the step that always dies first — keeping some spreadsheet somewhere so you can see the trend. That's five clerical steps per client per administration, and by February the whole practice has quietly sunk. Not because you stopped believing in it. Because it was a second job.
The fix is to make every one of those steps automatic: the measure goes to the client's portal a day or two before session, they fill it out on their phone in the waiting room or at home, scoring is instant, the result lands in the chart, and the trend line draws itself. This is one of the places your EHR either carries the practice or becomes the reason it doesn't exist — in Practice Harbor, the PHQ-9 and GAD-7 are built in: sent from the portal, scored automatically, and trended on the client's chart, so the only part left for you is the part that was always yours — looking at the line and deciding what it means.
That's the standard worth holding any system to, ours or anyone's: measurement-based care should be a byproduct of your normal workflow, not a discipline you maintain. You didn't get licensed to hand-score forms. You got licensed to read instruments and steer.
PHQ-9 and GAD-7 sent from the portal, scored automatically, and trended on the chart — measurement-based care as a byproduct of your week, not a second job. Free for pre-licensed clinicians, $19/mo licensed.
A practical cadence for private practice is a baseline at intake, then readministration every two to four weeks — every two weeks during an active treatment phase or medication change, monthly for stable longer-term work. Weekly administration is possible but tends to produce noise, since the PHQ-9 asks about a two-week symptom window; small week-to-week wobbles usually reflect ordinary life variation rather than clinical change.
A change of 5 or more points on the PHQ-9 is a commonly used threshold for clinically meaningful change. A shift from 18 to 16 is best read as staying the same, while a drop from 18 to 11 reflects genuine improvement. The same threshold applies in reverse: a sustained 5-point rise, or a steady upward trend across several administrations, is an early signal of non-response worth acting on.
Often, yes. CPT 96127 (brief emotional/behavioral assessment with a standardized instrument) is covered by many commercial plans and is billed per instrument administered, scored, and documented — so a PHQ-9 and a GAD-7 given at the same visit can be two units where the plan allows it. Reimbursement is modest and coverage varies by payer, so verify with each plan before counting on it.
No. The PHQ-9 and GAD-7 are validated screening and monitoring tools, not diagnostic instruments. An elevated score indicates significant self-reported symptoms and warrants a closer clinical look, but diagnosis comes from the clinical interview — grief, sleep deprivation, and medical conditions can all elevate scores. Their strength is tracking symptom severity over time, where they are free, well-validated, and universally understood by payers and physicians.