| Index | 주제 | 논문/자료 | 발표자 | 발표자료 & 영상 |
|---|---|---|---|---|
| 1 | LLM 평가 개괄 | A Survey on Evaluation of Large Language Models [링크] A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations [링크] |
김기범 | 📄 🎥 |
| 2 | Long-Context | Needle in a Haystack LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks InfiniteBench: Extending Long Context Evaluation Beyond 100K Tokens RULER: What’s the Real Context Size of Your Long-Context Language Models? A Controllable Examination for Long-Context Language Models |
조동헌 | 📄 🎥 |
| 2-1 | Long-Context(Satellite) | LongBench pro [링크] | 김기범 | 📄 🎥 |
| 3 | 지표 붕괴, Goodhart's law | The Leaderboard Illusion [링크] Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models [링크] |
한완규 | 📄 🎥 |
| 4 | Software Engineering | Evaluating Large Language Models Trained on Code (HumanEval)[링크] SWE-bench[링크] SWE-bench Verified[링크] Multi-SWE-bench[링크] SWE-Bench Illusion[링크] SWE-rebench[링크] SWE-bench Pro[링크] |
박진우 | 📄 🎥 |
| 4-1 | Tuning Coding Agents | Improving Deep Agents with harness engineering[링크] | 김기범 | 📄 🎥 |
| 5 | Agents - End to End | GAIA: A Benchmark for General AI Assistants [링크] WebArena: A Realistic Web Environment for Building Autonomous Agents [링크] An Illusion of Progress? Assessing the Current State of Web Agents(Online-Mind2Web) [링크] MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering[링크] |
김동현 | |
| 6 | Knowledge | GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks [링크] GPQA: A Graduate-Level Google-Proof Q&A Benchmark [링크] |
김정훈 | |
| 7 | Textual & Content Safety | HarmBench: A Standardized Evaluation Framework for Automated Red Teaming[링크] A StrongREJECT for Empty Jailbreaks[링크] TrustLLM: Trustworthiness in LLMs [링크] |
홍소현 | 🎥 |
| _ | Pi-mono, harness engineering | Pi Monorepo: Tools for building AI agents.[링크] | Sigrid Jin | 🎥 |
| 8 | Agentic & Behavioral Safety | Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents [링크] AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents[링크] sudo rm -rf agentic_security [링크] |
이동건 | 📄🎥 |
| 9 | LLM-as-a-Judge | Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models [링크] Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena [링크] JudgeBench: A Benchmark for Evaluating Judges [링크] |
조성국 | 📄🎥 |
| 10 | Agents - Tool Use | AGENTBENCH: Evaluating LLMs as Agents[링크] StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models[링크] MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers[링크] |
김강민 | 📄 |
| - | Twenty Questions Benchmark | Twenty Questions Benchmark[링크] | 김기범 | 🎥 |
| 11 | Thinking Process & Reasoning | Measuring Faithfulness in Chain-of-Thought Reasoning[링크] Evaluating Mathematical Reasoning Beyond Accuracy[링크] |
박진형 | 🎥 |
| 12 | Multimodal Reasoning | MMMU Benchmark [링크] Humanity's Last Exam [링크] |
최동혁 | 📄🎥 |
Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|