Skip to content
#

ml-observability

Here are 22 public repositories matching this topic...

Autonomous multi-agent system for ML infrastructure reliability. Detects, diagnoses, and remediates production ML incidents through a coordinated Sentry → Sleuth → Medic → Scribe agent pipeline. Built with Google ADK + Gemini, FastAPI, and a live failure-injection demo harness.

  • Updated Jul 6, 2026
  • Python

CPU-first framework to compare ML decision-system versions, identify changed decisions, and attribute shifts to features, models, thresholds, rules, and interactions.

  • Updated Sep 25, 2026
  • Python

Sentinel — Production ML Observability Platform MLOps · Model Drift Detection · Real-time Anomaly Monitoring · Full-stack A production-grade platform that monitors deployed ML models in real time — detecting data drift, concept drift, and prediction anomalies — visualised through a live dashboard with alerting. No hardware needed.

  • Updated Jun 15, 2026
  • Python

Add this topic to your repo

To associate your repository with the ml-observability topic, visit your repo's landing page and select "manage topics."

Learn more