Which model is really behind your API relay or agent IDE? Behavioral fingerprinting + anytime-valid sequential tests (FPR<=1%). LLMs can't be random - measured on 9 frontier models.
-
Updated
Jul 23, 2026 - Python
Which model is really behind your API relay or agent IDE? Behavioral fingerprinting + anytime-valid sequential tests (FPR<=1%). LLMs can't be random - measured on 9 frontier models.
Sequential comparison of probabilistic forecasts and anytime-valid inference in R.
Anytime-valid A/B test analyzer that holds the false-positive rate under 1.5% while you peek at the dashboard continuously, where naive fixed-horizon testing leaks to 23% at 15 looks. mSPRT with confidence sequences, CUPED variance reduction (50% on the demo, SE 0.223 to 0.135), a peeking guard, and a reproducible A/A simulation.
A GxP-purpose-built LLM evaluation framework: anytime-valid validation, judge qualification, ALCOA+ audit, human plus AI workflow validation, drift monitoring, listener hooks, and a high-level facade on Inspect AI. Enables responsible AI use under risk-based assurance.
Add a description, image, and links to the anytime-valid topic page so that developers can more easily learn about it.
To associate your repository with the anytime-valid topic, visit your repo's landing page and select "manage topics."