Researcher at the ITMO University AI Center. I work on whether automated monitors can actually read what an LLM agent did.
Two numbers, if you only read one line: established LLM judges report a failure on clean agent traces between 34% and 78% of the time, and 84% of a trace by volume contributes nothing to locating a failure.
- MASeval: automated evaluation of multi-agent systems. Contributor, on judge and verifier implementation and ablation tooling.
- What Must You Log?: what an execution trace must retain for failure attribution. First author.
- TraceJudgeBench: 877 agent traces unified from ten judge-validation datasets. Main contributor.
- AutoJudge: generating an LLM judge per trace. FAGEN @ ICML 2026.


