Model Leaderboard

Aggregate performance across all test suites

16/16
Models Selected
128
Total Suites
0
Fair Suites
0
Cases

Fair Comparison - Aggregate Performance

Only includes 0 suites run on ALL selected models

ModelSuitesCasesAccuracyScorePrecisionRecall ↓F1TPTNFPFN
anthropic-mod BEST 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
azure 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
baseten-v14f-mini 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
llamaguard 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
nope 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
evaluate-prod 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
oai-oss-safeguard 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
ocular-prod-H_c49_1_hybridseeds_v1_vllm 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
ocular-prod-H_c51_precision_hybrid_v1_vllm 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
ocular-prod-hybridseeds-v1.1-lockdown 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
ocular-risk-H_c49_1_hybridseeds_v1_vllm 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
ocular-risk-H_c51_b1b2_hybrid_v1_vllm 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
ocular-risk-H_c51_precision_hybrid_v1_vllm 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
openai 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
smod 0 0 0.0% N/A 0% 0% 0% 0 0 0 0
zentropi 0 0 0.0% N/A 0% 0% 0% 0 0 0 0

Metrics Guide

Accuracy - Binary pass/fail (all checks must pass)
Score - Graduated quality (0-100, partial credit)
Precision - When flagged, how often correct?
Recall - Of actual crises, how many caught?
F1 - Harmonic mean of precision and recall
FP - False Positive (flagged benign)
FN - False Negative (missed crisis)
SATURATED - Uncalibrated heads firing at max on all inputs; severity unreliable
Internal NOPE engineering tool · CF Access (@nope.net) · data is the full corpus, not the curated public surface