Skip to content
View barakhsin's full-sized avatar
  • ITMO University
  • Saint Petersburg

Highlights

  • Pro

Block or report barakhsin

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
barakhsin/README.md

Grigorii Barakhsin

Researcher at the ITMO University AI Center. I work on whether automated monitors can actually read what an LLM agent did.

Two numbers, if you only read one line: established LLM judges report a failure on clean agent traces between 34% and 78% of the time, and 84% of a trace by volume contributes nothing to locating a failure.

  • MASeval: automated evaluation of multi-agent systems. Contributor, on judge and verifier implementation and ablation tooling.
  • What Must You Log?: what an execution trace must retain for failure attribution. First author.
  • TraceJudgeBench: 877 agent traces unified from ten judge-validation datasets. Main contributor.
  • AutoJudge: generating an LLM judge per trace. FAGEN @ ICML 2026.

Website · Google Scholar · barakhsin@gmail.com

Pinned Loading

  1. ITMO-NSS-team/AutoJudge ITMO-NSS-team/AutoJudge Public

    Framework for Evaluating LLM-based Agentic Systems

    Python 10

  2. sb-ai-lab/MASeval sb-ai-lab/MASeval Public

    A Python library for automated evaluation of Multi-Agent Systems using pydantic AI

    Python 9 1