Suite ↕ Models ↕ Score ↑ Cases ↕ Updated ↕
calibration-probe
Calibration probe covering explicit crisis, algospeak, implicit ideation, resolution syndrome, and false positive categories (gaming, hyperbole, recovery narratives, temporal/negation). Tests model sensitivity across difficulty levels.
oai-oss-safeguard +2
0.0%
24
May 5
screen-perpetrator-disclosure-generalization
EXPLORATORY: Tests whether the model genuinely understands perpetrator disclosures (subject=other) vs. pattern-matching on prompt examples. Uses diverse scenarios NOT mentioned in training: different relationships, harm types, vulnerable populations, and framing styles. If model only learned 'mum + cancer = neglect:other', these will fail.
anthropic-mod +8
16.7%
18
May 7
cryptocurrency-romance-scam-crisis---detailed-assessment
Detailed risk assessment for cryptocurrency investment fraud (pig butchering). Tests feature detection (hopelessness, shame, perceived_burdensomeness, financial_crisis), subject attribution, severity calibration, and imminence levels. Based on FBI Operation Level Up intervention data.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
20.0%
20
May 7
south-asian-crisis-patterns
South Asian cultural crisis patterns including dowry harassment, academic pressure (JEE/NEET), family honor/shame (izzat), arranged marriage coercion, caste-based discrimination, farmer economic crisis, and LGBTQ+ persecution. Addresses gap in cultural crisis presentations specific to Indian subcontinent. Based on NCRB data showing 34,480 farmer suicides (2020), WHO estimates of 135,000 annual suicides in India, and research documenting izzat-related mental health burden.
anthropic-mod +8
25.0%
8
May 7
technology-mediated-harm
Technology-mediated harm patterns including deepfake pornography, romance/pig butchering scams, AI chatbot dependency, cyberstalking, cryptocurrency fraud, and online harassment. Based on Springer 2025 research showing 51% of image-based sexual abuse victims report SI, FBI Operation Level Up referring 64 romance scam victims for suicide intervention, and documented cases like Dennis Jones (82) who died by suicide after pig butchering scam.
anthropic-mod +8
26.3%
19
May 7
disability-progressive-disease
Disability and progressive neurological disease crisis patterns including ALS, MS, Huntington's, perceived burdensomeness, invisible disability invalidation, and sudden disability adjustment. Expands existing 13 screen cases to full evaluate depth. Based on research showing 19% of ALS patients express wish to die but only 37% meet criteria for depression—distinguishing autonomous end-of-life preferences from treatable crisis.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
26.3%
19
May 7
east-asian-cultural-crisis
East Asian cultural crisis patterns including academic pressure (gaokao, Korea exam stress), hikikomori (severe social withdrawal), somatic expressions of distress, and filial piety burden. Addresses gap in cultural crisis presentations that Western models may miss. Based on China Blue Book showing 93% of high school suicides linked to gaokao stress, Korea data showing >50% of adolescent suicide ideation citing academic stress, and Japanese Cabinet Office documenting 1.46 million in hikikomori state.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
26.7%
15
May 7
screen-flash-false-positives
Cases where flash (production model) incorrectly triggers on clearly benign content. High-value regression tests.
anthropic-mod +8
37.5%
8
May 7
trafficking_disclosure
Human trafficking disclosure patterns including labor trafficking, sex trafficking, and exploitation. Tests detection of debt bondage, psychological control, delayed victim self-identification, coached responses, trauma bonding, and disclosure barriers. Includes youth-specific patterns ('boyfriend' framing) and validated QYIT screening indicators. Critical zero-coverage domain.
anthropic-mod +8
38.5%
26
May 7
mena-crisis-contexts
Middle East and North Africa crisis patterns including war trauma (Syria, Yemen, Palestine), displacement, honor-based violence, political persecution, LGBTQ+ criminalization, sectarian violence, guardianship restrictions, forced conscription, economic collapse, and apostasy persecution. Addresses gap in crisis presentations from conflict-affected and restrictive governance contexts. Based on UNHCR data showing 13.2M Syrian refugees, WHO documentation of elevated PTSD in conflict zones, and human rights documentation of persecution.
anthropic-mod +8
38.5%
13
May 7
ai_mediated_risk-v2
[v1] Novel risk patterns emerging from AI-mediated conversations. Tests magical thinking, AI validation/collusion, parasocial attachment, multi-turn escalation, and method-seeking with philosophical framing. Based on clinical research on third-party validation effects in suicide risk.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
39.3%
28
May 7
mh_edge_cases-v2
[v1] Edge and ambiguous clinical/safety cases for NOPE Safety v1. Multi-domain assessment version with severity/imminence separation. Covers dark humor, jailbreak attempts, safeguarding, violence, self-neglect, eating disorders, false positives, gaming/meme culture, chronic illness, postpartum, and multi-turn escalation patterns. 35 test cases (27 single-turn, 8 multi-turn).
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
45.2%
42
May 7
postpartum_perinatal
Postpartum and perinatal mental health crisis patterns including intrusive thoughts about infant harm, ego-dystonic OCD-type thoughts (protective vs. psychotic), postpartum depression, postpartum psychosis, disclosure fears ('they'll take my baby'), and avoidance behaviors. Tests critical distinction between repugnant intrusive thoughts (no intent, horror response) and genuine risk to infant. Approximately 50% of new mothers experience intrusive thoughts; ego-dystonic quality is key protective factor.
anthropic-mod +8
45.5%
11
May 7
Multi-Marker False Positives
Tests where multiple crisis indicators cluster in benign contexts, ensuring NOPE doesn't over-weight coincidental marker combinations without considering overall context
ocular-risk-H_c49_1_hybridseeds_v1_vllm +12
45.5%
11
May 9
Self-Neglect Comprehensive
Comprehensive coverage of self-neglect risk type including medical non-adherence patterns, ambiguous intentionality, elderly self-neglect, and substance misuse. Addresses critical taxonomy gap where self-neglect was only primary type in ~10 cases despite being 1 of 9 official risk types.
anthropic-mod +9
48.5%
33
May 7
screen-minor-witness-household-crisis
Tests whether model recognizes minors witnessing household crises as indirect victims. Child welfare research shows witnessing DV harms children (ACEs). Key pattern: explicit age signal + household crisis = flag minor too. Without age signal, only primary victim flagged (reasonable default).
anthropic-mod +8
50.0%
20
May 7
latin-american-crisis-contexts
Latin American crisis patterns including femicide, gang violence and forced recruitment, migration crisis, forced disappearances (desaparecidos), economic collapse (Venezuela, Argentina), machismo-based gender violence, political persecution of activists/journalists, and climate-driven displacement. Addresses gap in crisis presentations from Latin American contexts. Based on UN data showing 12 femicides daily in Latin America, UNHCR documentation of Central American displacement, and economic crisis statistics.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
50.0%
8
May 7
filter_router_edge_cases-v2
[v1] Edge cases for risk classification accuracy. Tests indirect language, context confusion, ambiguous framing, coded language, third-party disclosures, and professional contexts using orthogonal subject/type structure.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
50.0%
18
May 7
evaluate-post-discharge-transitions
Post-discharge and high-risk transition periods - full risk assessment with severity, imminence, and feature detection. Research shows 300x elevated suicide risk in first week post-psychiatric discharge (2,950/100k person-years), with Day 2 showing highest absolute deaths. 51.5% of inpatients conceal SI to obtain discharge.
anthropic-mod +8
50.0%
26
May 7
jailbreaking
Adversarial jailbreaking suite testing NOPE-aware attacks. Assumes users may: (1) know they're being classified by NOPE, (2) understand the XML schema and taxonomy, (3) attempt meta-instruction injection, (4) try to game severity/imminence thresholds. Goal: ensure genuine risk is detected regardless of adversarial framing.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +10
51.4%
35
May 7
blindspots_v2
Research-driven blindspots test suite for current taxonomy. Covers implicit ideation (Crisis Text Line predictors), population-specific patterns (LGBTQ+, veterans, healthcare workers), and adversarial scenarios. Based on RESEARCH-SYNTHESIS.md and clinical literature.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +9
51.7%
29
May 7
screen-caregiver-neglect-crisis
Tests detection of neglect in caregiver contexts. Key challenge: subject attribution (the dependent is at risk, not the caregiver). Covers hostile refusal, caregiver burnout, perspective shifts, and false positives. Created to investigate classification inconsistencies observed in ad-hoc testing.
anthropic-mod +8
53.6%
28
May 7
ai_dependency_patterns-v1
Test cases for AI dependency features based on 2026 International AI Safety Report findings. Detects user-side indicators of problematic AI relationships: relationship substitution, compulsive usage, anthropomorphization, and separation distress.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
54.5%
11
May 7
adversarial-robustness
Tests classifier robustness against adversarial inputs: jailbreak attempts, classification evasion, minimization disclaimers, encoded content, and meta-level manipulation. Goal: ensure genuine risk is detected even when wrapped in evasion tactics.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +10
60.6%
33
May 7
v2_eating_disorders
Comprehensive eating disorder detection tests including anorexia, bulimia, binge-eating disorder, ARFID, and orthorexia patterns. Validates that ED behaviors are correctly classified with appropriate features for specialist resource matching. Includes pro-ED community language, male-specific presentations, minimization patterns, ED+suicidality intersection, and validated assessment instrument language (EAT-26, EDE-Q).
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
60.7%
28
May 7
screen-subject-attribution
Tests subject attribution: distinguishing self (speaker) vs other (third party) vs professional contexts. Created for minime v9 validation.
anthropic-mod +8
61.3%
31
May 7
Exploitation Comprehensive
Comprehensive coverage of exploitation risk type including elder financial abuse, romance scams, labor exploitation, and institutional exploitation. Addresses critical taxonomy gap where exploitation was only primary type in ~10 cases despite being 1 of 9 official risk types.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +9
63.0%
27
May 7
subject-context-features
Tests for subject context features: animal_involved, minor_involved, infant_involved, elderly_involved, vulnerable_adult_involved, parental_incapacity, witnessing_violence. These context markers help identify WHO is involved/at risk beyond the primary subject.
anthropic-mod +8
63.2%
19
May 7
age-context-calibration
Tests how age context (child/teen/adult) affects classification of normalized suicide language. Goal: observe natural calibration differences, not enforce hard rules.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
63.6%
22
May 7
reddit-external-validated
External validation suite from SI_D_Reddit_Labellings.csv. Cases where NOPE Edge v14f and dataset labels agree. Random sample with seed=123.
anthropic-mod +9
65.3%
95
May 7
screen-homepage-examples
Regression tests for examples shown on nope.net homepage. Ensures our public claims match API behavior.
azure, nope
66.7%
6
May 7
nssi_comprehensive
Comprehensive non-suicidal self-injury (NSSI) detection suite covering all ISAS functional categories, disclosure patterns, concealment strategies, and risk escalation. Tests ability to distinguish NSSI from suicidal behavior and identify anti-suicide function (critical for risk assessment). Includes contemporary algospeak and social media coded language.
anthropic-mod +9
66.7%
33
May 7
expanded_taxonomy-v2
[v1] Expanded taxonomy validation using orthogonal subject/type structure. Covers online exploitation, extremism, human trafficking, and eating disorders.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
66.7%
21
May 7
indirect_disclosure
Indirect crisis disclosure patterns including 'asking for a friend', third-person framing, hypothetical scenarios, minimization, and religious/spiritual framing. Research shows many callers take 10-15 minutes to disclose actual reason for crisis contact. These patterns are normative in most cultures - direct disclosure is the exception, not the rule.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +9
67.7%
31
May 7
active_crisis_imminent
Active suicide crisis and imminent attempt scenarios representing C-SSRS Level 5 (highest severity): active ideation with specific plan, intent, and timeline. Includes interrupted attempt language, preparatory acts, resolution syndrome, and in-progress attempts. Critical for regression prevention—these cases must be detected at critical/emergency severity.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
68.0%
25
May 7
verbatim_chat_register
Chat-register verbatim text from court-quoted SMS/text messages and Reddit posts. All items are real-time, informal, direct disclosures - NOT interview quotes or retrospective accounts. Sources: US court opinions (Justia), Reddit gaming/casual subreddits.
anthropic-mod +9
70.0%
10
May 7
interrupted_attempt_variations
Variations of interrupted suicide attempts across different methods, interrupters, and emotional responses. Tests generalization of interrupted attempt detection beyond specific wording patterns. Clinical basis: C-SSRS interrupted attempt criteria - 'started to do something to end life but someone/something stopped them before acting.'
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
70.0%
10
May 7
input-framing-v2
[v1] Input framing detection test suite. Tests the ability to detect HOW content is presented (direct, humor, creative, distanced, adversarial, etc.). Framing is orthogonal to risk assessment - these tests verify framing detection alongside appropriate risk classification.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +10
70.0%
20
May 7
cultural-crisis-gaps
Cultural crisis detection gaps: Evidence-based linguistic markers for underrepresented populations. Based on peer-reviewed research documenting population-specific crisis language patterns that standard Western clinical models miss. Covers: strength schema breakdown (African American), collectivist burden/shame framing (South Asian, Pacific Islander), religious prohibition conflict (MENA/Muslim), documentation fear (immigrant/refugee), and LGBTQ+ subgroup-specific patterns.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
70.2%
47
May 7
progressive_disengagement
Progressive disengagement patterns in crisis conversations. Research shows message shortening, withdrawal language, and emotional flattening often precede disconnection and potential harm. These multi-turn patterns require active outreach. Based on Althoff et al. (2016) crisis counselor effectiveness research and Crisis Text Line trajectory analysis.
anthropic-mod +8
70.3%
37
May 7
verbatim_crisis_starter_pack
Short verbatim excerpts from public reports, court decisions, and published research. Seed set with full provenance - not the full validation corpus. Sources: ACL Anthology, NJ DV Near-Fatality Report, Correctional Service Canada, NZ Women's Refuge, Delaware Supreme Court, UK child neglect research, Swiss refugee study, UCLan trafficking thesis, UCT eating disorder thesis.
anthropic-mod +8
70.6%
17
May 7
fp_research_driven-v2
[v1] False positive test cases derived from academic research on mental health classification systems. Tests negation handling, humor markers, physical pain idioms, internet slang, minimal-context slang variants, and other documented FP patterns.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +9
71.1%
38
May 7
screen-healthcare-workers
Crisis patterns specific to healthcare workers (physicians, nurses, veterinarians) during high-stress periods. Research shows unique linguistic markers including workplace demoralization, learned helplessness from systemic barriers, help-seeking barriers (no time, feeling irresponsible), and loss of compassion as distress signal. Veterinarians have highest occupational suicide rate; 49% of veterinarians with ideation cite work problems.
anthropic-mod +8
71.4%
14
May 7
edge-v9-known-weaknesses
Comprehensive test suite capturing Edge minime v9 known limitations discovered through testing (2026-02-18). These cases document false positive issues with short inputs, semantic misunderstanding, abstract descriptions, and serialization edge cases. Used for regression testing and future model improvement tracking.
anthropic-mod +8
72.7%
55
May 7
core_transparency_v2
Evidence-based core test suite for current taxonomy. Based on dual independent clinical research review. 26 cases covering full risk taxonomy with clinically-informed expectations.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
73.1%
26
May 7
screen-ambiguous-gray-area
Ambiguous presentations where even trained clinicians disagree on severity. Inter-rater reliability among psychotherapists is AC1 = 0.44 (psychology students AC1 = 0.35), with middle-range cases showing lowest agreement. These cases test the system's ability to handle uncertainty and borderline severity, where binary classification is inappropriate and conservative flagging is warranted.
anthropic-mod +8
73.3%
15
May 7
literary_metaphorical
Literary, poetic, and metaphorical expressions of suicidal ideation. Inspired by Li et al. (2025) 'Can Large Language Models Identify Implicit Suicidal Ideation?' which found LLMs struggle with abstract, metaphorical expressions. Tests detection of: cinematic metaphors ('fading out'), isolation metaphors (walls, barriers), existential/philosophical death framing, and artistic/literary references that encode suicidal ideation.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +9
73.3%
15
May 7
youth-false-positives
Youth hyperbole vs genuine crisis differentiation. NLP tools achieve only 15% accuracy in detecting sarcasm. Adolescent expressions like 'I'm literally dying,' 'kill me now,' and 'I'm dead' are ubiquitous in casual communication. Key markers: gaming/entertainment context, no emotional escalation, humor indicators, quick emotional recovery, social engagement patterns.
anthropic-mod +8
75.0%
12
May 7
unknown
Expanded stalking pattern test cases. Tests detection of stalking risk type across ex-partner, workplace, stranger, and tech-enabled contexts. Includes physical following, cyberstalking, surveillance, and threats.
anthropic-mod +8
75.0%
20
May 7
multilingual_crisis
Multilingual crisis expressions to ensure equitable detection across languages. Based on research into culturally-specific idioms, somatic expressions, and indirect crisis communication patterns. Each case includes original language, translation, and cultural context notes.
anthropic-mod +8
75.0%
72
May 7
chat_register_false_negatives
False-negative focused test suite for chat-register crisis patterns. Tests whether NOPE catches crisis expressed in Gen-Z/millennial slang, minimized language, and indirect patterns that might be missed. Inverse of fp-research-driven.json.
ocular-risk-H_c49_1_hybridseeds_v1_vllm +11
75.0%
20
May 9
screen-substance-use
Tests integration of substance use context with suicidal ideation. Research shows 45.6% of opioid overdose survivors reported some desire to die, revealing a spectrum from 'using to cope' through 'passive indifference' to 'active desire to die'. Critical for detection systems as substance use is a major co-occurring risk factor.
anthropic-mod +8
76.5%
17
May 7
panel-probe
Probe cases for panel consensus behavior. Tests edge cases where model variance matters: ambiguous severity, subtle signals, false positive traps, and explicit cases (control).
oai-oss-safeguard, ocular-prod-H_c49_1_hybridseeds_v1_vllm
78.6%
14
May 5
screen-reddit-non-suicidal-fp-minime
False positive cases from Reddit Non-Suicidal dataset that minime incorrectly flagged. All are clearly benign posts labeled Non-Suicidal in source dataset. Tests model ability to avoid false positives on content with superficially triggering phrases.
anthropic-mod +8
80.0%
45
May 7
protective_factors-v2
[v1] Test suite focused on protective factor DETECTION. Validates that protective factors are correctly identified from conversation content. Severity/imminence expectations are deliberately wide - this suite tests PF detection, not severity calibration.
anthropic-mod +10
80.0%
20
May 7
filter_router-v2
[v1] Tests that the classifier correctly identifies risk subjects and types. Validates orthogonal subject (self/other) and type (suicide/abuse/violence/etc) classification. Maps v1 domain routing to v1 subject/type structure.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
80.0%
15
May 7
cultural-crisis-gaps-2
Cultural crisis detection gaps part 2: Additional evidence-based linguistic markers including Cool Pose masking, culture-bound syndromes (falling out, ataque de nervios, musu), theological lethality markers, asexual/trans-specific patterns, and intergenerational debt framing. Based on comprehensive research synthesis of 160+ sources.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
80.0%
20
May 7
acute_real_world_v2
Acute real-world scenarios for current taxonomy. Covers immediate danger, psychosis, mania, AI attachment, safeguarding, and third-party risk. 30 evidence-aligned test cases from clinical literature and real-world patterns.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +9
80.0%
30
May 7
screen-reddit-false-positives
False positive regression tests from Reddit Suicidal Ideation Dataset. These are clearly benign posts (labeled Non-Suicidal) that should NOT trigger crisis resources. Covers: mundane daily life, accomplishments, celebrations, questions, humor, positive mental health recovery.
anthropic-mod +8
81.8%
55
May 7
somatization_patterns
Crisis presentations through somatic (physical) symptoms rather than emotional language. 'Somatization is the single most common reason why psychiatric illness goes undetected in general medical settings.' Physical complaints are the acceptable way to express distress in many cultures where mental illness is stigmatized.
anthropic-mod +8
82.1%
28
May 7
dangerous_content
Tests for detecting content that ENCOURAGES harm to others (not expressing personal distress). Includes dangerous challenges, pro-self-harm content, method sharing, and death validation. Key distinction: subject='other' (the person being encouraged) not 'self' (unless speaker is also at risk).
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
82.1%
28
May 7
hedged_severity_calibration
Hedged disclosure severity calibration tests. Research shows hedging (hypothetical framing, 'asking for a friend', dismissive qualifiers like 'lol') is a NORMATIVE disclosure pattern, not evidence of reduced intent. Content indicators (means, method, timeline) should override framing. Based on De Choudhury et al. (2016), Coppersmith et al. (2018), and Crisis Text Line research.
nope +11
82.4%
17
May 18
youth-developmental-stages
Age-specific crisis expressions across developmental stages. Research shows children's death vocabulary and crisis communication evolve significantly: preschoolers use concrete death language, pre-teens use indirect expressions and somatic complaints, young teens exhibit emerging abstract thinking with burden language, and older teens employ adult-like articulation.
anthropic-mod +8
83.3%
12
May 7
benign-idioms-&-colloquialisms---short-form
Tests false positive resistance on short statements using death/crisis-adjacent language in clearly benign contexts. Includes colloquialisms, hyperbole, academic references, and dramatic expressions that should NOT trigger crisis resources.
anthropic-mod +8
83.3%
72
May 7
neurodivergent_presentations
Crisis presentations from neurodivergent individuals (autism, ADHD, etc.) that may be missed by classifiers trained on neurotypical communication patterns. Autistic people are 3-9x more likely to die by suicide (Hirvikoski 2016); autistic women have ~13x suicide mortality risk vs non-autistic women; ideation is ~4x general population (Cassidy 2014). Key patterns: flat affect, literal expression, communication shutdown, perseverative thoughts.
anthropic-mod +9
84.4%
32
May 7
screen-chronic-illness-disability
Crisis patterns specific to chronic illness, chronic pain, and disability populations. These populations express crisis through unique linguistic markers including perceived burdensomeness related to dependency, treatment non-adherence as passive suicide method, and conditional survival language.
anthropic-mod +8
84.6%
13
May 7
harassment_severe
Tests for detecting severe, explicit harassment cases that can be identified from single messages (not pattern-based). Includes doxxing threats, sexual harassment, targeted degradation, and online pile-on indicators. Note: most harassment detection requires conversational context NOPE cannot provide - these tests cover explicit/severe cases only.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
85.2%
27
May 7
substance_crisis
Tests for substance abuse crisis detection, including substance+suicidal ideation combinations (known high-risk), overdose risk severity, and severe withdrawal states. These combinations are clinically established risk escalators.
anthropic-mod +9
85.7%
28
May 7
platform-context-calibration
Tests how platform/situational context affects classification. Gaming slang in gaming context should be benign. Same phrases without context may flag. For v12+ models, context is passed via [[begin system context]] tags.
anthropic-mod +8
85.7%
21
May 7
litmus-v2
Litmus test suite using orthogonal subject/type taxonomy. Each risk has subject (who) + type (what) + features. speaker_severity derived from risks where subject='self'. Comprehensive coverage of suicide, IPV, safeguarding, violence, and edge cases.
nope +14
86.4%
44
May 18
screen-post-discharge-transitions
Post-discharge and high-risk transition periods - linguistic markers for crisis detection. Research shows 300x elevated suicide risk in first week post-psychiatric discharge, with Day 2 showing highest absolute deaths. 51.5% of inpatients conceal SI. 49% of 2-week deaths occur before first follow-up.
anthropic-mod +8
86.7%
30
May 7
screen-research-derived
Test cases derived from academic research on crisis communication patterns, algospeak, cultural idioms, and forensic linguistics.
anthropic-mod +8
88.9%
45
May 7
benzodiazepine-withdrawal-crisis
Benzodiazepine withdrawal crisis patterns with suicidal ideation and impulsive self-harm. Research shows 54.4% of individuals who discontinued benzodiazepines experienced suicidal thoughts or attempted suicide. Case report documents 62-year-old male who, during rapid taper (60mg→7mg diazepam equivalent), within 36 hours became agitated and twice inflicted serious stab wounds requiring emergency surgery (Neale et al., 2007). Short-acting benzodiazepines (alprazolam, lorazepam) carry highest risk due to abrupt offset. Withdrawal mechanisms include GABA receptor dysregulation, paradoxical disinhibition, and medical invalidation of protracted symptoms. Cases cover acute taper, protracted withdrawal, paradoxical reactions, and false positives.
anthropic-mod +8
88.9%
18
May 7
filter_none_v2
False positive regression suite for current taxonomy. Tests that benign content, humor, idioms, and non-personal discussions should NOT trigger risk detection. Critical for preventing over-flagging in production.
ocular-prod-H_c49_1_hybridseeds_v1_vllm +8
89.8%
49
May 7
stalking_sam
Tests for stalking detection using SAM (Stalking Assessment & Management) framework. Includes ex-intimate partner stalking (highest risk), escalation patterns, violence history, victim fear, and proxy harassment. SAM-informed risk levels: severe, elevated, standard.
anthropic-mod +8
90.0%
20
May 7
pippa-genuine-crisis-multi-turn
Real-world crisis disclosures from Character.AI (Pippa dataset). Users disclosing genuine suicidal ideation and self-harm within roleplay conversations. These should ALL be flagged - the roleplay context does NOT negate real distress. Source: edge/testbeds/pippa/
nope +10
90.0%
10
May 18
c-ssrs-escalation
C-SSRS (Columbia-Suicide Severity Rating Scale) escalation prompts from academic research on mental health chatbot safety. Based on: Pichowicz W, Kotas M, Piotrowski P. 'Performance of mental health chatbot agents in detecting and managing suicidal ideation.' Scientific Reports. 2025;15:31652. https://doi.org/10.1038/s41598-025-17242-4
nope +10
90.0%
10
May 18
alcohol-withdrawal-crisis
Alcohol withdrawal crisis patterns with suicidal ideation. Research shows individuals with AUD have 40% lifetime prevalence of suicide attempt and 10-14× higher suicide risk vs general population. Withdrawal symptoms peak at 24-72 hours post-cessation with depressive symptoms most severe during detoxification. Cognitive constriction (narrowing of attention to present emotional state) is validated mechanism linking acute alcohol use and withdrawal to suicidal behavior. Cases cover acute withdrawal, protracted withdrawal syndrome, comorbidity, relapse patterns, and false positives.
anthropic-mod +8
90.0%
20
May 7
Implicit Crisis Expressions
Test cases targeting documented blind spots in competitor moderation APIs. OpenAI shows 33.2% false negative rate on implicit content, Azure 63.9% on implicit expressions, Perspective 75.4% overall. These cases test passive ideation, context-dependent statements, coded language, and cultural expressions that major APIs systematically miss.
nope +10
90.0%
20
May 18
high-risk-occupational-crisis
Crisis patterns in high-risk occupations: farmers (3.5x general rate), construction (75% higher), first responders (police 58% of FR suicides), lawyers (2x ideation rate), active military (28.2/100k), and dentists (PMR 2.01). Research-derived linguistic markers from qualitative studies, crisis hotline research, and occupational health literature. Citations in rationale.
anthropic-mod +8
90.9%
33
May 7
literary-metaphorical
Tests detection of implicit suicidal ideation expressed through literary, poetic, and metaphorical language. Based on Li et al. (2025) 'Can Large Language Models Identify Implicit Suicidal Ideation?' arXiv:2502.17899. These patterns use abstract language rather than explicit crisis statements - a known blind spot for keyword-based moderation.
anthropic-mod +8
91.7%
12
May 7
algospeak-comprehensive
Comprehensive algospeak test suite covering suicide euphemisms, self-harm depth indicators, character substitutions, emoji patterns, and false positive contexts. Based on 2026 research into TikTok, Reddit, Discord, and Tumblr crisis communication patterns.
nope +10
92.2%
51
May 18
youth-lgbtq-minority-stress
LGBTQ+ youth crisis patterns reflecting minority stress accumulation. Trevor Project 2024 survey: 39% seriously considered suicide, 12% attempted. Those experiencing 4 types of minority stress face 12x greater odds of attempt. Age 13-17 shows higher risk (46%/16%) than 18-24 (33%/8%). Conversion therapy exposure doubles attempt likelihood.
nope +10
92.3%
13
May 18
v2_population_context
Tests for population context extraction. Validates that demographic/identity indicators are captured as features to enable population-specific resource matching.
anthropic-mod +8
92.3%
13
May 7
screen-core
Core test suite for /screen endpoint. Tests suicide/self-harm detection using C-SSRS framework, covering active crisis, passive ideation, method-seeking, self-harm (NSSI), idioms, multi-turn conversations, and false positive prevention.
anthropic-mod +8
92.4%
92
May 7
subject_attribution-v2
Tests for correct subject attribution (self vs other). Probes edge cases where speaker IS the victim but mentions others in their situation.
nope +11
92.6%
27
May 18
older-adult-supplemental
Supplemental crisis patterns for adults 65+, covering patterns from second research document: firearm euphemisms (cleaning gun), medication hoarding (insurance/peace of mind framing), VSED patterns, financial ruin triggers, completed life rhetoric, spousal reunion/pact patterns, instructional farewells, past-tense narratives, and affective neutrality. Complements older-adult-crisis-patterns.json (42 cases).
anthropic-mod +8
92.6%
27
May 7
pregnancy-reproductive-loss
Pregnancy loss and reproductive health crisis detection. Based on evidence that suicide is leading cause of maternal death 6 weeks to 1 year postpartum (MBRRACE-UK), with stillbirth conferring 5.2x elevated risk. Covers miscarriage, stillbirth, infertility/IVF, birth trauma, NICU, TFMR, partner grief, TTC community language, medical terminology trauma, financial entrapment, obstetric violence, and reunion motivation patterns. Sources: Weng et al. 2018 (BJOG, DOI: 10.1111/1471-0528.15105), Lewkowitz et al. 2019 (AJOG), Tommy's National Centre, Bailey et al. 2019 (BMJ Open), Shani et al. 2016, 1001 Critical Days study.
anthropic-mod +8
92.7%
55
May 7
validated-clinical-expressions-screen
Crisis language patterns from validated clinical instruments (C-SSRS, ASQ, PHQ-9), official public health warning sign lists (NIMH, AFSP, JED Foundation), and crisis service training materials. Organized by 7 thematic categories from suicidal ideation detection research.
anthropic-mod +8
92.9%
28
May 7
screen-reddit-v9-detection-gaps
Detection gap cases from Reddit dataset that minime-v9 failed to flag. All are labeled Suicidal in source dataset but v9 returned 'none'. Tests critical detection capabilities: algospeak ('unalive'), past-tense disclosures, and third-party concern.
anthropic-mod +8
93.3%
15
May 7
screen-immigrant-refugee-crisis
Immigrant and refugee crisis patterns including asylum detention, deportation fear, family separation trauma, professional deskilling, and climate refugees. Addresses complete gap (0 existing cases) where immigrants/refugees represent high-risk population. Based on 2020 ICE detention suicide rate of 17.4 per 100,000 (5.3x the 2010-2019 average) and Hispanic suicide rate increase of 26.6% (2015-2020).
anthropic-mod +8
93.3%
15
May 7
eating-disorders
Pro-ED content, thinspo, and eating disorder terminology - Koko has 498 high-confidence patterns here
nope +10
93.3%
15
May 18
chat-register-crisis
Crisis detection in informal chat register: text speak, emoji, hedging with dismissive qualifiers ('lol its dumb but...'), and specific medication/means references masked by casual framing. Based on De Choudhury et al. (2016) computational discourse analysis, Coppersmith et al. (2018) lexical markers, and Crisis Text Line emoji research.
nope +10
93.3%
15
May 18
algospeak-obfuscation
Obfuscated crisis terms using leetspeak, spacing, symbols - tests pattern matching sophistication
nope +10
93.3%
15
May 18
Cultural Identity False Positives
Test cases targeting documented over-flagging by competitor moderation APIs. CHI 2025 audit found all major APIs over-flag content containing LGBTQ+, Black, Jewish, and Muslim identity terms. These cases test trauma processing, recovery narratives, educational content, and identity discussions that should NOT be flagged as crisis.
anthropic-mod +8
93.3%
15
May 7
pippa-roleplay-benign
Roleplay and fictional scenarios from Character.AI (Pippa dataset) that should NOT flag as crisis. Tests model's ability to distinguish fictional death/harm language from genuine crisis. Source: edge/testbeds/pippa/
nope +10
93.8%
16
May 18
long-form-multi-turn-conversations
Tests crisis detection in extended conversations with many turns, long messages, and needle-in-haystack scenarios. Validates that models can maintain context and detect signals embedded in large amounts of benign content.
anthropic-mod +8
93.8%
16
May 7
realistic_prevalence_v2
Realistic prevalence test suite mimicking real-world deployment conditions. 200 messages with ~1% crisis rate (2 true crises, 198 non-crises). Tests PPV, specificity, and false positive rate under realistic class imbalance. Based on Swaminathan et al. (2023) findings: prospective deployment had 0.55% crisis prevalence. (Converted to subject/type taxonomy with subject/type risks.)
ocular-risk-H_c49_1_hybridseeds_v1_vllm +13
93.8%
97
May 9
screen-economic-distress-advanced
Advanced economic distress patterns covering the 'Transactional Self' (commodification of existence), somatic manifestations, high-velocity ruin (crypto/trading), agrarian stewardship failure, construction industry stoicism, and housing deadline triggers. Based on the 'Deaths of Despair' framework (Case & Deaton) and Joiner's Interpersonal Theory of Suicide.
anthropic-mod +8
94.4%
36
May 7
ai-companion-dependency-crisis-patterns
User-side detection of crisis risk from AI companion dependency, isolation, and harmful AI encouragement. Based on 2024-2026 research including Sewell Setzer case and clinical dependency frameworks.
nope +10
95.0%
20
May 18
opioid-overdose-suicidal-intent
Opioid overdose cases with varying levels of suicidal intent. Research shows 58.5% of opioid overdose patients had at least some desire to die, 36% had strong desire, and 21% reported 'I definitely wanted to die' (Connery et al., 2019). Yet few overdoses are coded as suicide attempts, creating critical detection gap. Uses C-SSRS framework to distinguish passive intent (wish to be dead) from active ideation (thoughts with method) and preparatory behaviors.
anthropic-mod +8
95.8%
48
May 7
tiktok-algorithmic-harm-spirals
Detection of crisis risk from TikTok algorithm-driven exposure to self-harm, suicide, and harmful content. Based on 2024-2025 research on For You feed escalation, engagement loops, and vulnerable user targeting.
azure +8
96.0%
25
May 7
speaker_third_party_v2
Tests for speaker vs third-party risk disambiguation. Uses orthogonal subject/type: speaker_severity derived from risks where subject='self', third-party risks have subject='other'.
nope +12
96.0%
25
May 18
sextortion-crisis-patterns
Sextortion crisis patterns for suicide detection. Tests the critical linguistic shift from external problem-focus ('I'm being blackmailed') to internal defeat ('I can't survive this'). Based on documented cases showing victims dying within 27 minutes to 6 hours of first contact. Sources: FBI 2024 sextortion data, NCMEC 36+ documented suicide cases, Thorn financial sextortion research 2024, Sadath et al. 2024 humiliation-suicide meta-analysis.
anthropic-mod +8
96.4%
28
May 7
cryptocurrency-romance-scam-crisis-patterns
Detection of crisis risk from cryptocurrency investment fraud (pig butchering scams). Based on FBI Operation Level Up data (80 suicide interventions) and financial strain research showing 20-fold increased suicide risk with cumulative stressors.
anthropic-mod +8
96.7%
30
May 7
screen-economic-distress-crisis
Economic distress and financial crisis suicide patterns. Based on research showing problem debt creates 7.96x suicide attempt risk (Naranjo et al. 2021), combined financial strains create 20x increase (Elbogen et al. 2020), and 79% of foreclosure suicides occur BEFORE actual housing loss (Houle & Light 2014). Tests three primary pathways: perceived burdensomeness, provider identity collapse, and escape reasoning.
anthropic-mod +8
97.5%
40
May 7
screen-reddit-longform-narratives
Long-form (1500-5000 char) Reddit posts from Suicidal Ideation Detection Dataset. Focus on extended narratives that require understanding context across multiple paragraphs.
azure +8
98.0%
50
May 7
youth-school-academic
School-based crisis patterns including academic pressure, discipline, college rejection, and exam stress. Research shows 14% prevalence of suicide risk among students with academic pressure, and the MARIS study found academic failure was the only predictor of suicide after one month. Suicide rates are notably lowest during school closures.
azure +8
100.0%
12
May 7
youth-romantic-relationship
First romantic relationship and breakup crisis patterns. Research shows adolescents with dissolution stage issues (breakup concerns) were significantly more likely to present with suicide/self-harm than other relationship stages. Study of 370 students found 19.5% had major breakup, 34.3% had suicidal ideation with significant association. First breakup intensity reflects developmental factors: limited experience creates catastrophic thinking.
azure +8
100.0%
12
May 7
youth-family-conflict
Family conflict and dysfunction as primary risk pathway. Meta-analysis found adolescents with family dysfunction have 1.93x higher suicide risk. TORDIA study demonstrated family conflict significantly predicted suicidal adverse events. Child maltreatment, particularly sexual abuse, carries highest attempt risk (OR 11.7-49.3).
anthropic-mod +8
100.0%
13
May 7
youth-cyberbullying
Cyberbullying crisis patterns distinct from traditional bullying. NIH/CHOP study found cybervictims are 4x more likely to report suicidal thoughts/attempts, independent of in-person bullying. Key distinguishing factor: inability to escape - harassment follows victims home, can be anonymous, spreads virally, reaches wider audiences.
anthropic-mod +8
100.0%
12
May 7
youth-contagion
Suicide contagion and cluster patterns. Teens who know someone who attempted are approximately 3x more likely to attempt themselves (OR 2.8-11.0). 1-5% of teen suicides occur in clusters, more common in under-25s. Notably, 'less close friends who knew the deceased have the highest risk.' Celebrity suicides increase rates significantly - Robin Williams linked to 1,800+ additional deaths, '13 Reasons Why' caused 28.9% increase in 10-17yo suicides.
anthropic-mod +8
100.0%
12
May 7
stimulant-psychosis-crisis
Methamphetamine-induced psychosis crisis patterns with violence risk and suicidal ideation. Research shows in meth users with psychosis: 85.5% have delusions of persecution, 75.6% violence behavior, 51.3% auditory hallucinations, and 40% formication (delusional parasitosis/'meth mites'). Violence risk escalation documented: 37% obtained weapon, 11% used weapon, 15% attacked another person. Onset can occur 1-5 days after initiation or within less than a week with prior history. Hallucinations typically resolve 1-2 days, delusions 2-3 weeks, but 30% have symptoms persisting 6+ months. Cases cover paranoid delusions, formication self-injury, command hallucinations, disorganized states, and false positives.
anthropic-mod +8
100.0%
18
May 7
screen-victimization
Tests victimization detection in expanded /screen. Victimization cases (abuse, stalking, trafficking, etc.) should show_resources=true with correct risk type detection. SI/SH should only flag when speaker also expresses suicidal ideation or self-harm.
anthropic-mod +8
100.0%
18
May 7
screen-veteran-military-crisis
Veteran and military crisis patterns including moral injury, institutional betrayal, military sexual trauma, TBI comorbidity, and transition crisis. Addresses gap where only 3 veteran cases existed. Based on VA data showing veterans with moral injury have higher suicide rates and 50% of veteran suicide deaths had received VA services.
azure +8
100.0%
18
May 7
screen-postpartum-transitions
Crisis patterns during major life transitions including postpartum period, motherhood adjustment, and acute care-seeking urgency. Research shows mothers hide suicidal feelings to adhere to cultural expectations of motherhood, with unique linguistic markers around loss of control, overwhelm, and incongruence between expectations vs reality.
anthropic-mod +8
100.0%
10
May 7
screen-minority-cultural
Crisis patterns from racial/ethnic minority populations showing distinct linguistic markers. Research shows 'hidden ideation' in Asian American populations (less likely to explicitly state suicidal thoughts), intergenerational trauma framing in Indigenous populations, and shame-limited disclosure in Latino populations. Critical for ensuring detection systems work across demographic groups.
anthropic-mod +8
100.0%
11
May 7
screen-lgbtq-adult-crisis
LGBTQ+ adult crisis patterns distinct from youth coverage. Includes coming out later in life (30s-60s), trans healthcare denial, elder LGBTQ+ isolation/re-closeting, HIV/AIDS crisis, religious trauma in adulthood, workplace discrimination, and conversion therapy aftermath. Addresses gap where existing coverage focused on youth (13 cases) with minimal adult representation (5-7 cases).
anthropic-mod +8
100.0%
15
May 7
screen-indigenous-global-patterns
Indigenous crisis patterns globally including intergenerational trauma (residential/boarding schools), land dispossession, cultural genocide, MMIW, substance misuse linked to historical trauma, youth suicide clusters, forced removal, environmental destruction, colonial violence legacy, and cultural disconnection. Addresses complete gap (0 existing Indigenous-specific cases). Based on CDC data showing Indigenous suicide rate 3.5x higher than general population, Canadian TRC documentation, and global Indigenous health disparities.
anthropic-mod +8
100.0%
12
May 7
screen-bias-stability
Bias stability testing: semantically equivalent cases where only identity markers differ. Each group contains variants that should produce identical classification results. Tests for gender, relationship, name, and age bias. Inspired by research showing AI systems can exhibit differential treatment based on demographic markers.
anthropic-mod +8
100.0%
109
May 7
multilingual-crisis
Crisis terms in non-English languages - tests international keyword coverage
anthropic-mod +8
100.0%
15
May 7
method-seeking
Tests for implicit method-seeking patterns (distress + location/means queries)
anthropic-mod +8
100.0%
17
May 7
kms-hyperbole-calibration
Calibration suite for 'kms' (kill myself) detection. Tests the boundary between hyperbolic internet slang and genuine masked ideation. Key principle: trivial stressors + humor markers = no flag; significant stressors or isolation language = flag even with humor.
nope +10
100.0%
19
May 18
keyword-blind-spots
Cases requiring semantic understanding - keyword matchers will struggle here, LLMs should excel
anthropic-mod +8
100.0%
16
May 7
indigenous-crisis-patterns
Crisis patterns in Indigenous/Native populations including historical trauma, intergenerational effects, cluster/contagion contexts, and Two-Spirit/Indigenous LGBTQ+ intersections. Based on Brave Heart (2003), Bombay et al. (2014), and SAMHSA cluster guidance. Includes critical false positive guidance for cultural spiritual expressions.
nope +10
100.0%
15
May 18
explicit-keywords
Direct crisis terminology - keyword matchers should perform well here
anthropic-mod +8
100.0%
16
May 7
elder-self-neglect-crisis
Elder self-neglect crisis patterns with depression and suicidal ideation. Research shows 51% of older adults with self-neglect have depression (vs 28% without); depression more than doubles self-neglect risk (OR: 2.38). When both dementia and depression present, risk is 8.6× higher. Diogenes syndrome (severe self-neglect) has 46% 5-year mortality rate. Among those with suicidal ideation, hoarding is most prevalent self-neglect phenotype (33.3%). Executive dysfunction - capacity to make decisions remains intact but capacity to implement self-care is diminished. Elder Self-Neglect Assessment (ESNA) domains: personal care, nutrition, environment/housing, financial stewardship, socialization. Cases cover medical non-compliance, nutritional neglect, hygiene neglect, environmental squalor, executive dysfunction, hoarding, and false positives.
nope +10
100.0%
23
May 18
correctional-crisis
Crisis patterns in correctional populations: booking/intake, pre-trial detention, and post-release periods. Based on BJS suicide data, Binswanger et al. (2007) post-release mortality research, and NCCHC guidelines.
anthropic-mod +8
100.0%
15
May 7