Skip to content

About

Reproducible coherence failures in real AI agents. One row per framework (LangGraph, CrewAI, Letta, OpenHands, Phoenix, DBOS), each with a trace that replays in one command.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

The coherence census

Every agent and framework the Right Rudder read has run on, what it caught, and a trace that reproduces the catch in one command.

Fathom is now Right Rudder, by Embedded Risk Analytics. The rows, traces and findings stay as they were, and the commands use the right-rudder package that replaces fathom-read.

A long-running agent renames a field, then keeps writing the old name. It reports twenty records written when it wrote nineteen. It places an order, then places it again. The run reports success, and the contradiction ships. The Right Rudder read folds the actions that succeeded into the agent's committed state and names the step that contradicts it. This repository is the census: one row per agent or framework, with the study the row comes from and a small trace in the shape of the failure the study found.

pip install right-rudder
git clone https://github.com/RightRudderAI/coherence-census && cd coherence-census
right-rudder read traces/langgraph_history.json --format langgraph
Agent Trace the read consumes What the read caught in the study The bundled trace returns Reproduce Study
LangGraph Checkpoint lineage (graph.get_state_history) The channel reducer merged both writes with no consistency check and left a record citing the key the agent had already renamed away superseded_value x1 right-rudder read traces/langgraph_history.json --format langgraph study
CrewAI Event bus (tool_usage_finished, tool_usage_error, task_completed) A hierarchical crew reported all twenty sub-tasks complete while one record never got written; the read recovered the dropped record from the event stream alone residual x1 right-rudder read traces/crewai_events.json --format crewai --supersede guest_id=customer_id study
CrewAI (the write_a_book_with_flows example) Event bus (tool_usage_finished, task_completed), mapped by the row's book_flow_ops.py Under llama-3.3-70b at 14 chapters one chapter crew's researcher issued the identical search 25 times and three crews re-ran queries the outline crew had already committed; another run saved a 14-chapter book as finished with an outlined chapter missing behind a placeholder title duplicate_commit x28 right-rudder read traces/crewai_book_flow.json study
deepagents (the deep_research example) The tool trace across the orchestrator and its research sub-agents (a LangChain callback handler), mapped by the row's deep_research_ops.py Under llama-3.3-70b the report's only citation was a URL no search returned, the real article's slug rewritten to match its title; in another run the orchestrator delegated the same question twice, the second sub-agent reopened with the first one's exact query, and the run held one unreachable article five times before reporting that nothing was found stale_reference x1 right-rudder read traces/deepagents_deep_research.json study
DBOS (the Hacker News research agent example) The DBOS step stream (every checkpointed step with its arguments and result), mapped by the row's hn_agent_ops.py As published the agent repeated 14 of 50 searches, all from iteration 4 on, and re-read 42 percent of the threads it fetched; with the repair in front of the one step where it proposes its next queries, repeats fell to 0 of 50, re-reads to 12 percent, and distinct threads covered rose 35 percent at the same model and iteration count duplicate_commit x2, stale_reference x1 right-rudder read traces/dbos_hn_agent.json study
Prime Agent (an orchestrator with rlm child agents) The orchestrator's session with the agent messages its children delivered, mapped by right-rudder-prime-agent Under deepseek/deepseek-chat-v3-0324 with 8 children per round the orchestrator wrote child-format messages for children that had not replied and committed from them; across 11 runs with no failure induced, 39 of 329 commits rested on reports not yet delivered, 27 of them in one run stale_reference x1 right-rudder read traces/prime_agent_fanin.json study
Letta Core blocks, archival passages, and the memory-edit tool calls The agent renamed the six blocks it could see and left the archival copy on the old key; the read recovered the stale passage from the persisted memory and the edit stream duplicate_commit x6, residual x1 right-rudder read traces/letta_memory.json --format letta --supersede guest_id=customer_id study
Arize Phoenix (OpenInference) TOOL spans as Phoenix stores them The starved run's platform signals read tests-pass-and-done while the tool spans showed no successful edit; the read named the files still carrying the old key residual x1 right-rudder read traces/phoenix_spans.json --format openinference --supersede guest_id=customer_id study
OpenHands (coding agent) The edit log (str_replace_editor calls with success flags) The small-model run made no successful edit, ran the suite against unchanged code, and reported the task complete with tests passing; the read was the one check that separated the reported success from the actual one residual x5 right-rudder read traces/rename_starved.json --format edits --supersede guest_id=customer_id study
Agent-E (web agent) The browser action stream With the window starved, a frontier model duplicated the order (8 lines and $106.20 against the correct 3 lines and $44.10) while every conventional success signal stayed green; the read recovered all five duplicates duplicate_commit x1, post_commit_mutation x1 right-rudder read traces/order_duplicate.json study
Agent Zero Memory (provenanced memory, on LongMemEval) Retrieved items and the answer, with the citation lock's verdict When the update fell outside the retrieval window the reader cited the superseded item and answered with the superseded value; across 83 stale answers the citation lock rejected none, and the read fired on every one superseded_value x1 right-rudder read traces/knowledge_update.json study
ContextPilot (context management, on LongMemEval) The folded history and the answer Once the update was folded, 55 to 60 percent of questions came back with the value the agent had already replaced; the read produced no verified false positives across more than a hundred runs superseded_value x1 right-rudder read traces/knowledge_update.json study

The bundled traces are minimal: each one carries the failure's shape in a dozen ops so you can read it in a minute and run it without a model. The study column holds the full run, with the framework, the models, and the counts. python scripts/build.py runs every row against the read and rewrites this table; CI does the same on every push.

Rows in progress

HAL's τ-bench airline traces across seven frontier models. The CrewAI book-flow, deepagents and DBOS rows above carry their full runs under rows/crewai-book-flow/, rows/deepagents-deep-research/ and rows/dbos-hn-agent/. Each lands as a row with its trace and its command when the run is done.

Add a row

Run the read on an agent the census does not cover, and open a pull request with three things: a trace in traces/ in one of the formats the read accepts (right-rudder formats lists them), the command that reproduces the finding, and a row in census.json naming the agent, the trace it consumes, and what the read caught. A row whose trace returns a finding on the hosted read passes CI. If the read stayed silent on a run you expected it to flag, open an issue with the trace; a miss is worth as much as a catch.

Send us a trace

If you would rather not run it yourself, send a trace to contact@embeddedriskanalytics.com and get a readout back. The read is deterministic and needs no model access; the trace is the only thing it receives.

Right Rudder is a program of Embedded Risk Analytics. The research behind the read: embeddedriskanalytics.com/research. Paper: SSRN 6683578.


If a row here matches a failure in your own agents, a star on this repository helps other teams find the census.

About

Reproducible coherence failures in real AI agents. One row per framework (LangGraph, CrewAI, Letta, OpenHands, Phoenix, DBOS), each with a trace that replays in one command.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages