Eval scorecard
Public quality metrics for AuditAI's classify-workflow AI route. All 13 AI routes are scored nightly; this page surfaces the headline ground-truth eval.
Every classify-workflow call is scored against a curated ground-truth suite (5 cases spanning MINIMAL through HIGH_RISK EU AI Act tiers). The pass threshold is 0.7 on a weighted blend: tier accuracy 30%, NIST recall 20%, OWASP recall 15%, ISO recall 15%, human-in-loop match 10%, shadow-IT match 10%. The other 12 routes (finding drafting, executive summary, threat modelling, etc.) run structural shape evals with binary pass / fail per case.
No eval runs yet.
Data source: EvalRun table, refreshed nightly via the Inngest ai-evals-nightly job at 03:30 UTC. The same scorer runs against operator-trigger ad-hoc requests too.