Hi, congrats on the benchmark release!
I wonder if there's a plan to open-source your GitHub task-mining/discovery pipeline?
Specifically, your paper mentions “expert-exploratory” tasks - I wonder if these "experts" are just LLMs, or actual human experts?
Hi, congrats on the benchmark release!
I wonder if there's a plan to open-source your GitHub task-mining/discovery pipeline?
Specifically, your paper mentions “expert-exploratory” tasks - I wonder if these "experts" are just LLMs, or actual human experts?