
When PHQ-9 Scores Mislead: Rating Scales and Clinical Judgment
A PHQ-9 score can mislead as easily as it clarifies. The panel unpacks the gap between rating-scale remission and how patients actually experience recovery, using real cases where appearances and scores diverged sharply.
In "When PHQ-9 Scores Mislead: Rating Scales and Clinical Judgment," the panel examines how rating scales capture, and sometimes miss, what remission really looks like.
Dr. Citrome opens by defining response as a fifty percent reduction on a rating scale and remission as falling below a set threshold, then asks whether that distinction means anything to a patient. Dr. Harding admits her views are mixed: clinicians running clinical trials and clinicians who are mostly patient-facing often see rating scales very differently. Her practice, which accepts all insurance, still needs objective measures to justify continued coverage to payers, so scales retain real practical value.
She traces today's common scales, including the PHQ-9 and clinician-rated tools like the Hamilton, back to a broader screening packet developed in the late 1990s for primary care, later repurposed to track symptom severity over time. The risk, she explains, is that clinicians fixate on the number and forget the person behind it. A PHQ-9 total of three can look like near-remission, yet if question nine, which asks about thoughts of death or self-harm, scores high, that patient is not truly in remission at all.
Dr. Citrome adds that question ten, which measures how much these symptoms affect daily functioning, is the most overlooked and arguably most important item on the scale. He notes that patients often score zero on the suicide item out of fear of being hospitalized, so he normalizes morbid thoughts directly to invite honesty. He then shares two contrasting cases from independent medical exams: a well-groomed, articulate woman who scored twenty on the PHQ-9 despite appearing composed, because she had learned to hide her depression from her children and coworkers, and a woman who cried for an entire hour yet scored only seven, explaining that the visit itself, not her baseline state, had upset her. Both cases show how scale scores and clinical impression can mislead alone, reinforcing that patient-reported outcomes matter most read with clinical judgment.
In "Raising the Bar in MDD: What True Remission Looks Like," Erin Crown, MHS, PA-C, CAQ-Psychiatry explains why a return to life before depression, not a scale score, is the recovery standard she holds herself to.
Related to this article








