#
lighteval
Here are 5 public repositories matching this topic...
Wrapper of common LLM evaluation frameworks
evaluation artificial-intelligence llm lm-evaluation-harness vllm lighteval openai-compatible swebench
-
Updated
Apr 2, 2026 - Python
Rank-targeted nested sequential design for LLM evaluation: reach the same ranking conclusion for less, and see which comparisons the data never supported.
statistics variance-reduction experimental-design sequential-testing llm-evaluation lighteval anytime-valid
-
Updated
Aug 15, 2026 - Python
Evaluate and compare language models using lighteval for benchmarks and custom tasks, with tools for flexible, efficient analysis.
-
Updated
Jan 25, 2025 - Jupyter Notebook
Add this topic to your repo
To associate your repository with the lighteval topic, visit your repo's landing page and select "manage topics."