Evaluation-driven autonomous development: a harness for agent loops whose acceptance criterion is a metric, not a test suite. Held-out metrics, sealed scoring, integrity checks.
verification autonomous-research llm-agents reward-hacking agent-harness evaluation-driven-development loop-engineering held-out-validation
-
Updated
Jul 31, 2026 - Python