Skip to content

Create evaluation and calibration harness for decision packs #6

Description

@Sunny-commit

Goal

Make decision-pack quality measurable and claims reproducible.

Scope

  • Labeled fixtures and deterministic evaluation runner.
  • Accuracy, confusion counts, per-class metrics, Brier score, and calibration/ECE where probabilities are available.
  • Separate mock-provider tests from optional live-Laya evaluation.
  • Machine-readable and human-readable reports.

Acceptance criteria

  • Dev and held-out fixtures are clearly separated.
  • Reports include model/checkpoint, configuration, sample size, hardware, and methodology.
  • Loading, routing, network, and inference latency are distinguished.
  • Negative results and known limitations are retained.
  • CI prevents unsupported benchmark claims from being presented as verified.

Dependency

Decision-pack pull requests should add fixtures compatible with this harness.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions