The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.
-
Updated
May 23, 2025 - Python
The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.
NitroML is a modular, portable, and scalable model-quality benchmarking framework for Machine Learning and Automated Machine Learning (AutoML) pipelines.
Benchmarking the ability of large language models to detect semantic conflicts across domains, documents, and evolving knowledge bases.
Open-source self-hosted AI model testing, LLM evaluation, and API compatibility for Claude, GPT, Gemini, custom models, REST API, SQLite, and Docker.
Adversarial Testing Lab for Agentic Safeguards (ATLAS). A synthetic multi-agent eval environment for adversarial fraud decisioning inspired by Anthropic's Project Deal. Measures how model quality, tool access, and agent orchestration affect attack discovery & defensive recovery, with deterministic evals and realistic customer-friction limits
Agent 降智检测与自愈公评网络 — an immune system for the AI agent society
A model predicts whether a patient will be diagnosed with diabetes
Intro to Machine Learning Project from TripleTen
Model drift, delayed-label quality, and monitoring-evidence readiness gates with CLI/API parity, Prometheus, Docker, Kubernetes, and Terraform.
MLOps ticket triage with quality, routing, and auditable artifact promotion gates
Add a description, image, and links to the model-quality topic page so that developers can more easily learn about it.
To associate your repository with the model-quality topic, visit your repo's landing page and select "manage topics."