Budget-aware certification of which agent-eval tasks clear a reliability floor. At 5 repeats per task a uniform eval certifies zero tasks at 95 percent simultaneous confidence; a thresholding-bandit allocation certifies 79 of 200 on the same budget, and cuts wrong reliability verdicts from 14.9 to 9.5.
flaky-tests budget-allocation llm-evaluation agent-evaluation confidence-sequence thresholding-bandit
-
Updated
Oct 7, 2026 - Python