Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

2026 LLM Evaluation 논문 스터디

Index 주제 논문/자료 발표자 발표자료 & 영상
1 LLM 평가 개괄 A Survey on Evaluation of Large Language Models [링크]
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations [링크]
김기범 📄 🎥
2 Long-Context Needle in a Haystack
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
InfiniteBench: Extending Long Context Evaluation Beyond 100K Tokens
RULER: What’s the Real Context Size of Your Long-Context Language Models?
A Controllable Examination for Long-Context Language Models
조동헌 📄 🎥
2-1 Long-Context(Satellite) LongBench pro [링크] 김기범 📄 🎥
3 지표 붕괴, Goodhart's law The Leaderboard Illusion [링크]
Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models [링크]
한완규 📄 🎥
4 Software Engineering Evaluating Large Language Models Trained on Code (HumanEval)[링크]
SWE-bench[링크]
SWE-bench Verified[링크]
Multi-SWE-bench[링크]
SWE-Bench Illusion[링크]
SWE-rebench[링크]
SWE-bench Pro[링크]
박진우 📄 🎥
4-1 Tuning Coding Agents Improving Deep Agents with harness engineering[링크] 김기범 📄 🎥
5 Agents - End to End GAIA: A Benchmark for General AI Assistants [링크]
WebArena: A Realistic Web Environment for Building Autonomous Agents [링크]
An Illusion of Progress? Assessing the Current State of Web Agents(Online-Mind2Web) [링크]
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering[링크]
김동현
6 Knowledge GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks [링크]
GPQA: A Graduate-Level Google-Proof Q&A Benchmark [링크]
김정훈
7 Textual & Content Safety HarmBench: A Standardized Evaluation Framework for Automated Red Teaming[링크]
A StrongREJECT for Empty Jailbreaks[링크]
TrustLLM: Trustworthiness in LLMs [링크]
홍소현 🎥
_ Pi-mono, harness engineering Pi Monorepo: Tools for building AI agents.[링크] Sigrid Jin 🎥
8 Agentic & Behavioral Safety Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents [링크]
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents[링크]
sudo rm -rf agentic_security [링크]
이동건 📄🎥
9 LLM-as-a-Judge Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models [링크]
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena [링크]
JudgeBench: A Benchmark for Evaluating Judges [링크]
조성국 📄🎥
10 Agents - Tool Use AGENTBENCH: Evaluating LLMs as Agents[링크]
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models[링크]
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers[링크]
김강민 📄
- Twenty Questions Benchmark Twenty Questions Benchmark[링크] 김기범 🎥
11 Thinking Process & Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning[링크]
Evaluating Mathematical Reasoning Beyond Accuracy[링크]
박진형 🎥
12 Multimodal Reasoning MMMU Benchmark [링크]
Humanity's Last Exam [링크]
최동혁 📄🎥

About

2026 llm evaluation study

Resources

Stars

26 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors