Model Leaderboard
Aggregate performance across all test suites
16/16
Models Selected
128
Total Suites
0
Fair Suites
0
Cases
Fair Comparison - Aggregate Performance
Only includes 0 suites run on ALL selected models
| Model | Suites | Cases | Accuracy | Score | Precision | Recall ↓ | F1 | TP | TN | FP | FN | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| anthropic-mod BEST | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| azure | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| baseten-v14f-mini | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| llamaguard | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| nope | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| evaluate-prod | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| oai-oss-safeguard | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| ocular-prod-H_c49_1_hybridseeds_v1_vllm | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| ocular-prod-H_c51_precision_hybrid_v1_vllm | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| ocular-prod-hybridseeds-v1.1-lockdown | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| ocular-risk-H_c49_1_hybridseeds_v1_vllm | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| ocular-risk-H_c51_b1b2_hybrid_v1_vllm | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| ocular-risk-H_c51_precision_hybrid_v1_vllm | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| openai | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| smod | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 | |
| zentropi | 0 | 0 | 0.0% | N/A | 0% | 0% | 0% | 0 | 0 | 0 | 0 |
Metrics Guide
Accuracy - Binary pass/fail (all checks must pass)
Score - Graduated quality (0-100, partial credit)
Precision - When flagged, how often correct?
Recall - Of actual crises, how many caught?
F1 - Harmonic mean of precision and recall
FP - False Positive (flagged benign)
FN - False Negative (missed crisis)
SATURATED - Uncalibrated heads firing at max on all inputs; severity unreliable