October is the month the mental health calendar gets crowded. Mental Illness Awareness Week runs October 4 through 10 this year. The Thursday of that week, October 8, is National Depression Screening Day, when clinics, campuses, and workplaces set out a short questionnaire and invite people to fill it in. Two days later is World Mental Health Day, and this year’s theme, chosen by the World Federation for Mental Health, is “Lived Experiences Heard: Real Voices, Real Change.”

I want to write about the questionnaire, because I think it sits exactly at the point where those two ideas meet. The form most of those tables will be handing out is nine questions long. I use it in my own practice. I built it into the app I make. And I have watched it do something quietly remarkable and something quietly misleading, often on the same afternoon. This is what I would tell you about it before you pick up the pen.

Where the nine questions came from

The Patient Health Questionnaire-9, the PHQ-9, was developed in the late 1990s by Robert Spitzer, Kurt Kroenke, and Janet Williams, and validated in a study published in 2001 across several thousand patients in primary care and obstetric clinics. Its design is almost austere. Each of the nine questions corresponds to one of the nine symptoms that define a major depressive episode in the diagnostic manual: low mood, loss of interest, sleep, energy, appetite, self-worth, concentration, slowing or restlessness, and thoughts of death. You are asked how often each has bothered you over the past two weeks, on a four-point scale from “not at all” to “nearly every day.” The answers are scored zero to three and added up. The total runs from zero to twenty-seven.

The score bands are the ones you may have seen on a results sheet: five to nine is described as mild, ten to fourteen as moderate, fifteen to nineteen as moderately severe, twenty and above as severe. In the validation study, a score of ten or more identified major depression with a sensitivity and a specificity of about 88 percent each. That is a good instrument. It is also free to use, which is a large part of why it is everywhere: on clipboards in waiting rooms, in electronic health records, in the app on your phone.

Notice what the design gives up. There is no question about why. There is no question about what happened two weeks ago, or thirty years ago. There is no question about whether you have ever had the opposite experience, a stretch of days when you needed almost no sleep and felt magnificent. There is no question about grief, or about the medication you started in August, or about your thyroid. The form is a count of symptoms over a fortnight. That is all it ever claimed to be, and it is the source of both its usefulness and its limits.

What the number is, and what it is not

Here is the sentence I say before every screening and again after it, and the sentence the app shows on both sides of every assessment: this is a screening tool, not a diagnosis.

People nod at that and then, understandably, read the number as a diagnosis anyway. A fourteen feels like a verdict. So let me be precise about what it is.

A screening score is a snapshot of two weeks, taken by the person being photographed, with the camera they had at the time. It is meant to answer one question: is this worth a closer look? A ten or above says yes. It does not say what the closer look will find.

The closer look is the diagnosis, and a diagnosis is a clinical judgment, not a sum. It takes in how long this has been going on and whether it is new. It asks about the events surrounding the change, and about losses, because grief and depression overlap on almost every item of this form and are not the same thing. It asks about sleep deprivation, medical illness, alcohol, and medications that can produce every one of the nine symptoms. And it asks, crucially, about the other pole. A person whose depression is part of bipolar disorder can score exactly the same as a person whose depression is not, and the right treatment for the two can be very different. A screening for depression alone cannot see that. A clinician who asks about elevated periods can.

There is also a plainer statistical point. When a test is given to a lot of people, most of whom do not have the condition, a fair share of the positive screens turn out, on closer look, to be something else. That is not a flaw in the PHQ-9. It is what screening means. A positive screen is an invitation to a conversation, not the conclusion of one.

And the reverse is true. Some people score low and are in real trouble. People who minimize, people who have been depressed so long the questions describe their baseline, people who answer the way they think they are supposed to. The number is what you told it. If you are reading this and thinking that your score has never captured how bad it actually is, that is worth saying out loud to someone. The form cannot hear you. A person can.

The ninth question

One of the nine questions is not like the others.

It asks whether, over the past two weeks, you have had thoughts that you would be better off dead, or of hurting yourself in some way. It is scored like the rest, zero to three, and it counts toward the total like the rest. But every competent use of this form treats it differently, and you should know why.

The total score can be low and the ninth answer can be anything but zero. When that happens, the total does not matter. Any answer other than “not at all” on that question means a follow-up conversation, that day, with a human being. Not because the form has diagnosed anything, but because large health-system studies have found that how people answer this one item is associated with what happens to them afterward, and because it is the only item on the form where the cost of being wrong is not symmetrical.

I built the app to behave exactly this way. If someone answers the ninth question with anything other than “not at all,” the app does not wait for the total. It stops, checks in, and puts crisis resources on the screen, every time, regardless of the score. It notifies no one. It keeps a record that a safety check happened and never a record of what was said. I would want you to demand that of any app that asks you this question, and to walk away from any that does not.

If that is you, right now: in the United States, call or text 988. You do not have to finish the questionnaire first.

What a count of symptoms cannot see

I am a psychoanalytic therapist trained in depth psychology, and my tradition has a long, sometimes prickly relationship with symptom checklists. I want to state the objection fairly, because I think it is right, and then say why I use the form anyway.

Two people score fourteen. The first is a woman whose father died in June, who has not slept properly since, and who is doing, slowly and painfully, the work of mourning. The second is a man who has carried, since childhood, a voice that tells him he is worthless, and which has grown louder since a promotion he did not think he deserved. Same number. Same band. On the form they are indistinguishable. In the room they could not be more different, and the work with each of them will look nothing alike.

The PHQ-9 counts the weather. It does not describe the climate, and it has no way of asking about the internal story a person is living inside: what the low mood is about, whom it is addressed to, what it protects them from, why now. Those questions are where treatment actually happens. The best evidence we have suggests that therapies which go after that story, and not only the symptom count, produce changes that keep growing after the sessions stop, which is not something a fortnightly score would ever have predicted.

So the objection stands. A screening is a surface. And this is where I think this year’s World Mental Health Day theme is quietly exactly right. Lived experiences heard. The number is not the experience. But in my practice the number is very often how the experience first gets into the room. A person who could not have said “I have been thinking about death” can circle a two. Once it is circled, it can be asked about. The form does not replace the voice. Used well, it is what gets the voice a hearing.

Why take it anyway: the value of the second time

If the single score is so limited, why do I hand the form out at all? Because of the second score. And the fifth.

There is a well-established finding in psychotherapy research that clinicians, including very experienced ones, are poor at noticing when a client is getting worse. In one widely cited study, therapists were asked to predict which of their clients would deteriorate. They identified almost none of the ones who did. A brief measure, given routinely, caught most of them. When that information was fed back to the therapist, the clients who had been drifting off track did measurably better. The approach has a name, measurement-based care, and the striking thing about it is how rarely it is used: fewer than one in five mental health practitioners routinely measure symptoms over time.

The reason the repeated measure works is the same reason the single measure disappoints. Memory for mood is reconstructed from the current mood. If you are low today, last month will look low in retrospect; if you feel better, you will underestimate how bad it got. A score written down on a particular date is immune to that. It is a message from the person you were two weeks ago, and it says something you cannot otherwise know: whether the direction is up or down. A change of about five points is commonly treated as meaningful. Two scores a month apart can tell you that. No amount of introspection can.

This is the honest job of a screening questionnaire, and it is a good one: not to name what you have, but to give the observing part of your mind a fixed point to measure from, so that when the low mood insists that nothing is changing, there is a number in your own handwriting that disagrees.

What to do with your number

If you take a screening this month, here is what I would do with the result.

Under five. You are probably not in a depressive episode right now, which is not the same as being fine. If something brought you to the table, that something is still worth a conversation.

Five to nine. Mild by the bands, but mild over many months is not mild. Take it again in a few weeks. If it holds or climbs, mention it to a doctor or therapist. This is often the range where changing sleep, movement, and isolation actually moves the number.

Ten and above. This is the range the form was built to flag. Bring the number to a professional. Bring the date you took it. If you can, bring more than one score, and bring anything you have tracked about sleep, because the first questions you will be asked are the ones the form skipped: how long, what happened, and whether it has ever gone the other way.

Any answer but zero on the ninth question. Talk to someone today. 988 is there for exactly this, and so is the app’s safety plan if you have made one.

Whatever the number. Say it out loud to a person. That is the part the form cannot do for you, and it is the part that changes things.

Real voices, real change

I have a private theory about why the screening tables on that Thursday in October do more good than their nine questions can account for. It is not that the form is wise. It is that filling it in is the first time many people have been asked, in plain words, whether they have felt hopeless, whether they have thought about death, whether the last two weeks have been bearable. The questions are ordinary. Being asked them is not.

That is what a good screening does, and it is all it does. It does not diagnose you, it does not know your story, and it cannot tell your grief from someone else’s illness. It gives you a number to hold and a date to hold it against, it flags the one answer that cannot wait, and it turns something private into something that can be said. What happens next is not the form’s job. It is a conversation, with someone who can listen for the climate underneath the weather.

This World Mental Health Day, take the nine questions if they are offered. Then do the harder thing, which is to tell someone the answer.


The Observing Ego is a mood tracking app, not a substitute for professional care, a crisis service, or a diagnostic tool. The PHQ-9, GAD-7, and every other assessment in it are labeled screening tools, not diagnoses, before and after each use. Its safety features, including crisis resources, the safety plan, and suicide-risk screening, are free permanently and require no account, and the app notifies no one. If you are in crisis in the United States, call or text 988, or chat at 988lifeline.org. Outside the US, local crisis lines are listed in the app’s safety section.