One-pass residual readouts vs fair closed-set AR scoring (BoolQ, RuleTaker, ARC)
-
Updated
Aug 12, 2026 - Python
One-pass residual readouts vs fair closed-set AR scoring (BoolQ, RuleTaker, ARC)
Confidence calibration toolkit for LLM verbalized-probability outputs. Real benchmark on 998 BoolQ questions with Llama-3.1-8B: ECE 0.148 -> 0.030, log-loss 3.9 -> 0.41.
Add a description, image, and links to the boolq topic page so that developers can more easily learn about it.
To associate your repository with the boolq topic, visit your repo's landing page and select "manage topics."