Publication|Articles|July 21, 2026

Psychiatric Times

  • Vol 43, Issue 7

Requiring AI-Driven Suicide Risk Stratification in Emergency Settings: Helpful or Risky?

Listen
0:00 / 0:00

Key Takeaways

  • Evidence suggests AI-augmented ED assessment meaningfully improves discrimination vs clinician judgment alone when leveraging longitudinal EHR predictors and social deprivation indices.
  • Hybrid workflows pairing universal validated screening (eg, C‑SSRS) with real-time machine learning may better identify near-term attempt risk than either modality independently.
SHOW MORE

Should EDs use AI to spot suicide risk? Explore evidence, bias concerns, false positives, and what accreditation mandates could mean for care.

Nearly half of individuals who die by suicide see a health care provider in the month before their death.1 Emergency departments (EDs) serve as a possible intervention point2 and often represent the last point of contact for many at-risk patients.3 Traditional clinician-led suicide risk assessments have well-documented limitations,4 particularly in the high-pressure, time-constrained ED environment, where predictive accuracy can be as low as chance levels for some outcomes. Enter artificial intelligence (AI).

AI can quickly analyze electronic health record (EHR) data, including demographics, prior visits, medications, and social determinants. AI is promoted as augmenting human judgment, flagging high-risk patients for further evaluation. The Joint Commission already mandates universal suicide ideation screening using validated tools for behavioral health patients.5 A question for accrediting bodies, such as the Joint Commission, is whether to elevate this to a requirement: Should hospitals be required to integrate AI-driven risk stratification into ED workflows to maintain accreditation?

The Case for Requirement

Advocates describe a stark reality: Clinicians are imperfect risk assessors. A 2025 study of nearly 90,000 patients across outpatient, inpatient, and ED settings found that clinician estimates of suicide attempt risk achieved an area under the curve (AUC) of just 0.60 in the ED, the lowest of any setting. When machine-learning models, incorporating up to 87 EHR predictors, were layered on top, AUC improved to 0.76, a statistically significant gain achieved using routinely collected data.6

Proponents emphasize that these models are not designed to replace clinicians but to act as a safety net. They integrate structured clinician input with historical EHR data, demographics, diagnostic codes, medication history, and even area deprivation indices. A 2022 cohort study demonstrated that combining face-to-face Columbia Suicide Severity Rating Scale (C-SSRS) screening with real-time machine learning outperformed either approach alone, particularly for suicide attempts.7

Real-world impact evidence strengthens the case. The US Department of Veterans Affairs’ REACH-VET (Recovery Engagement and Coordination for Health-Veterans Enhanced Treatment) program, which uses machine learning to identify the 0.1% of veterans at highest risk and trigger enhanced outreach (ie, safety planning, increased monitoring, care coordination), achieved a 5% reduction in documented suicide attempts after adjusting for cohort differences. Patients in the highest-risk tier died by suicide at 30 times the general VA rate, precisely the population the models aimed to catch.8

Accreditation mandates could drive systemwide improvements. Many hospitals still rely on outdated EHR systems; a national requirement would incentivize upgrades, improving data quality for all patients. Negative claims about coercion are dismissed by noting that Joint Commission standards already require validated screening; AI could simply enhance performance, much like radar assists air traffic controllers. Furthermore, proponents highlight the concept of accepted deviance in modern medicine.9 Clinicians often exaggerate their review of the medical record despite their ethical duty to comprehensively do so and despite mandates to note risk factors present in the record. AI would actualize the comprehensive chart review that is already required.

Bias and model accuracy are other key areas of debate. Advocates argue that although all predictive tools have limitations, biases, and errors, AI model errors are more transparent and correctable through systematic auditing than inconsistencies in individual clinical judgment. A 2019 study by Obermeyer et al examined a widely used commercial risk–prediction algorithm that relied on health care costs as a proxy for illness severity.10 This approach led to systematic underestimation of patient needs among those with lower health care utilization. When the model was recalibrated using direct clinical indicators (such as chronic condition counts), its performance improved substantially.

A 2025 review on bias recognition and mitigation in health care AI outlined effective strategies. Those included the use of multiple data sources, explainability tools, and continuous recalibration, which can enhance model reliability and predictive consistency.11 Proponents maintain that accreditation requirements for AI would promote greater transparency, regular validation, and accountability, which are often missing in traditional, unstandardized clinician assessments. They conclude that failing to adopt these tools when evidence suggests potential for improved risk stratification would be unethical given the scale of the national suicide burden.

The Case Against

Despite promising controlled studies, the evidence is insufficient to support mandatory high-stakes deployment in emergency settings. A 2025 systematic review of reviews examined 23 prior syntheses of AI suicide prediction models and found pervasive methodological shortcomings: Only 4% achieved high rigor, 64% were moderate, and the rest were low or critically low. Most studies had small samples (fewer than 1000 participants in 48%), no risk-of-bias assessments (86%), short follow-ups, and conflation of suicidal ideation, attempts, and deaths. Outcomes would vary with vastly different base rates and clinical implications.12

Positive predictive value (PPV) remains alarmingly low. Even models with 90% sensitivity and specificity in a 1% prevalence population yield a PPV of just 8.3%, meaning over 90% of high-risk flags are false positives. In the ED, where decisions about involuntary holds, resource allocation, and patient trust carry immediate consequences, false positives risk stigma, trauma, unnecessary coercion, and alarm fatigue. Achieving high precision often requires accepting high false-negative rates, while still missing genuine crises.12

EHR data quality compounds the problem. Hospital records contain significant error rates and incomplete histories. Up to 21% of records contain errors.13 Should an inaccurate record entry label a patient as high risk for life? Furthermore, legacy systems common in rural hospitals limit generalizability. Models trained on historical clinician judgments risk epistemic circularity, perpetuating past inaccuracies. REACH-VET’s success occurred in a VA outpatient context with structured follow-up, not a chaotic ED where primary interventions involve disposition decisions under time pressure.8 Additionally, although records already contain errors, AI is well-known for its propensity to hallucinate, up to 94% of the time.14

Ethical and practical risks also exist. Black box algorithms lack interpretability, undermining clinician confidence and risking automation bias. Dario Amodei, PhD, the CEO of Anthropic, admits, “We do not understand how our own AI creations work.”15 How could ethically informed consent be achieved when the creators themselves can't grasp the tool? A 2025 study on AI’s impact on primary-care mental health decisions found that physicians changed treatment recommendations to align with AI suggestions in nearly two-thirds of cases where the model conflicted with their initial judgment, sometimes inappropriately.16 A 2026 review noted that AI tools can pose a significant risk of eroding physicians' skills and that up to 30% of physicians reversed correct initial diagnoses when exposed to incorrect AI suggestions under time constraints.17

A mandate ignores the logistical and ethical realities of the health care system. In environments where algorithmic scoring is mandatory, clinicians report alarm fatigue within the first 30 days, causing them to reflexively click through warnings.18 False positives in an emergency setting can lead to involuntary hospitalizations, severe stigmatization, and the fracturing of patient trust, which is the core of psychiatry. Furthermore, as providers become required to dismiss those warnings, they assume additional legal responsibility for doing so, absolving the systems but taking on the risk themselves.

Algorithmic bias remains a persistent concern despite mitigation efforts. Obermeyer et al (2019) proved that correction is possible in retrospect, but mandating adoption now risks locking in inaccuracies from imperfect training data.10 Informed consent is complicated when patients are acutely suicidal and their capacity is possibly impaired. A counterplan often surfaces: Why require AI when we do not mandate other interventions with stronger evidence bases? One can argue that the field should prioritize randomized controlled trials in ED settings that demonstrate reduced attempts or deaths before accreditation mandates.

The Path Forward

The debate reveals a classic tension between innovation and caution. But more importantly, the debate highlights that complicated and sensitive topics can be discussed civilly and with scientific rigor in psychiatry. “Accrediting bodies should require artificial intelligence to be used for suicide risk stratification in emergency settings” was the topic of the 2026 National Psychiatry Resident Debate competition, held in May. The debate’s point was not to resolve this important question but to highlight the value of discussion in our field.

Ultimately, so much of psychiatry is about discussing thoughts and beliefs between our patients and us. Psychiatry deals with many issues related to society, policy, and people. As such, it is natural for our field to benefit from this type of discussion.

Dr Cheema is a third-year psychiatry resident at Baylor College of Medicine in Houston, Texas. Her interests include psychotherapy, addiction, and bioethics.

Dr Canto is a third-year psychiatry resident at Baylor College of Medicine. His interests include child and adolescent psychiatry, for which he will be pursuing fellowship training at Mount Sinai Hospital in New York, New York, as well as psychodynamic psychotherapy and neuropsychiatry.

Dr Badre is a clinical and forensic psychiatrist in San Diego, California. He teaches medical education, psychopharmacology, ethics in psychiatry, and correctional care. Dr Badre can be reached at his website, BadreMD.com. His upcoming textbook of psychiatry is available on Amazon.

References

1. Ahmedani BK, Simon GE, Stewart C, et al. Health care contacts in the year before suicide death. J Gen Intern Med. 2014;29(6):870-877.

2. Suicide prevention: data and statistics. Centers for Disease Control and Prevention. Updated May 20, 2026. Accessed June 18, 2026. https://www.cdc.gov/suicide/data/index.html

3. John A, DelPozo-Banos M, Gunnell D, et al. Contacts with primary and secondary healthcare prior to suicide: case–control whole-population-based study using person-level linked routine data in Wales, UK, 2000–2017. Br J Psychiatry. 2020;217(6):717-724.

4. Badre N, Compton J. The cult of the suicide risk assessment. Clinical Psychiatry News. September 7, 2023. Accessed March 30, 2026. https://www.mdedge.com/psychiatry/article/265143/depression/cult-suicide-risk-assessment

5. National patient safety goals. The Joint Commission. 2025. Accessed March 30, 2026. https://www.jointcommission.org/en-us/standards/national-patient-safety-goals

6. Bentley KH, Kennedy CJ, Khadse PN, et al. Clinician suicide risk assessment for prediction of suicide attempt in a large health care system. JAMA Psychiatry. 2025;82(6):599-608.

7. Wilimitis D, Turer RW, Ripperger M, et al. Integration of face-to-face screening with real-time machine learning to predict risk of suicide among adults. JAMA Network Open. 2022;5(5):e2212095.

8. McCarthy JF, Cooper SA, Dent KR, et al. Evaluation of the recovery engagement and coordination for health–veterans enhanced treatment suicide risk modeling clinical program in the veterans health administration. JAMA Netw Open. 2021;4(10):e2129900.

9. Banja J. The normalization of deviance in healthcare delivery. Bus Horiz. 2010;53(2):139-148.

10. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453.

11. Hasanzadeh F, Josephson CB, Waters G, Adedinsewo D, Azizi Z, White JA. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. NPJ Digit Med. 2025;8(1):154.

12. Abdelmoteleb S, Ghallab M, IsHak WW. Evaluating the ability of artificial intelligence to predict suicide: a systematic review of reviews. J Affect Disord. 2025;382:525-539.

13. Bell SK, Delbanco T, Elmore JG, et al. Frequency and types of patient-reported errors in electronic health record ambulatory care notes. JAMA Netw Open. 2020;3(6):e205867.

14. Jaźwińska K, Chandrasekar A. AI search has a citation problem. Columbia Journalism Review. March 6, 2025. Accessed March 30, 2026. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php

15. Amodei D. The urgency of interpretability. April 2025. Accessed March 30, 2026. https://www.darioamodei.com/post/the-urgency-of-interpretability

16. Ryan K, Yang HJ, Kim B, Kim JP. Assessing the impact of AI on physician decision-making for mental health treatment in primary care. npj Mental Health Research. 2025;4(1):16.

17. Heudel PE, Crochet H, Filori Q, Bachelot T, Blay JY. Artificial intelligence in medicine: a scoping review of the risk of deskilling and loss of expertise among physicians. ESMO Real World Data and Digital Oncol. 2026;12:100693.

18. Ancker JS, Edwards A, Nosal S, Hauser D, Mauer E, Kaushal R; HITEC Investigators. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Med Inform Decis Mak. 2017;17(1):36.