Skip to content
#

model-evaluation

Here are 1,583 public repositories matching this topic...

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

  • Updated Aug 3, 2026
  • Python

AI-powered NBA game outcome predictor that uses advanced team stats and trend-based features to forecast winners and track model performance

  • Updated Aug 1, 2026
  • Jupyter Notebook

Customers in the telecom industry can choose from a variety of service providers and actively switch from one to the next. With the help of ML classification algorithms, we are going to predict the Churn.

  • Updated Dec 29, 2021
  • Jupyter Notebook

An in-depth analysis of audio classification on the RAVDESS dataset. Feature engineering, hyperparameter optimization, model evaluation, and cross-validation with a variety of ML techniques and MLP

  • Updated Nov 5, 2020
  • Jupyter Notebook
model-serving-minefield

Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.

  • Updated Aug 3, 2026
  • Python

OpenLLM Monitor is a plug-and-play, real-time observability dashboard for monitoring and debugging LLM API calls across OpenAI, Ollama, OpenRouter, and more. Tracks tokens, latency, cost, retries, and lets you replay prompts — fully open-source and self-hostable.

  • Updated Jul 7, 2026
  • JavaScript

Improve this page

Add a description, image, and links to the model-evaluation topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the model-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more