-
Notifications
You must be signed in to change notification settings - Fork 0
feat(ruby_llm): correlate evaluation traces and control delivery #5
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
6 commits
Select commit
Hold shift + click to select a range
dd41106
feat(ruby_llm): correlate evaluation traces and control delivery
TonsOfFun ebbfec5
chore(release): 0.3.0
TonsOfFun 7610bda
fix(ruby_llm): keep a turn's scope and honour sampling on synchronous…
TonsOfFun 8e6585e
fix(ruby_llm): capture a turn's scope once and describe acceptance ho…
TonsOfFun 37404a2
fix(core): deliver a synchronous batching call's traces without racin…
TonsOfFun 175a182
fix(core): report a synchronous batching call whose payload cannot be…
TonsOfFun File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -3,7 +3,7 @@ | |
| module ActiveAgents | ||
| module Telemetry | ||
| module RubyLLM | ||
| VERSION = "0.2.0" | ||
| VERSION = "0.3.0" | ||
| end | ||
| end | ||
| end | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,120 @@ | ||
| # frozen_string_literal: true | ||
|
|
||
| require "test_helper" | ||
|
|
||
| class TestEvaluationContext < Minitest::Test | ||
| include RubyLLMTelemetryTestHelpers | ||
|
|
||
| def test_correlates_the_actual_trace_and_separates_judge_identity | ||
| ids = [] | ||
| Adapter.with_agent("EvaluationJudge", action: "score", | ||
| attributes: { "eval.run_id" => "run-1", "eval.result_id" => "result-1", "agent.class" => "Wrong" }, | ||
| on_trace: ->(trace) { ids << trace.trace_id }) do | ||
| instrument("chat.ruby_llm", chat_payload) { nil } | ||
| end | ||
|
|
||
| trace = traces.fetch(0) | ||
| root = spans_of(trace, "root").fetch(0) | ||
| assert_equal [ trace["trace_id"] ], ids | ||
| assert_equal "EvaluationJudge.score", root["name"] | ||
| assert_equal "EvaluationJudge", root["attributes"]["agent.class"] | ||
| assert_equal "run-1", root["attributes"]["eval.run_id"] | ||
| assert_equal "result-1", root["attributes"]["eval.result_id"] | ||
| end | ||
|
|
||
| def test_context_is_restored_after_nested_judge_and_exception | ||
| Adapter.with_agent("Support", action: "respond", attributes: { "eval.run_id" => "outer" }) do | ||
| assert_raises(RuntimeError) do | ||
| Adapter.with_agent("Judge", action: "score", attributes: { "eval.run_id" => "inner" }) do | ||
| instrument("chat.ruby_llm", chat_payload) { nil } | ||
| raise "synthetic error" | ||
| end | ||
| end | ||
| instrument("chat.ruby_llm", chat_payload) { nil } | ||
| end | ||
| instrument("chat.ruby_llm", chat_payload) { nil } | ||
|
|
||
| roots = traces.map { |trace| spans_of(trace, "root").first } | ||
| assert_equal [ "Judge.score", "Support.respond", "RubyLLM::Chat.chat" ], roots.map { |root| root["name"] } | ||
| assert_equal [ "inner", "outer", nil ], roots.map { |root| root["attributes"]["eval.run_id"] } | ||
| end | ||
|
|
||
| def test_synchronous_scope_delivers_in_the_calling_thread_without_mutating_configuration | ||
| subscribe(async: true) | ||
| caller_thread = Thread.current | ||
| delivery_threads = [] | ||
| captured = posted | ||
| Adapter.reporter.define_singleton_method(:deliver) do |body| | ||
| delivery_threads << Thread.current | ||
| captured << body | ||
| end | ||
|
|
||
| Adapter.with_agent("Support", synchronous: true) do | ||
| instrument("chat.ruby_llm", chat_payload) { nil } | ||
| assert_equal 1, posted.size | ||
| end | ||
|
|
||
| assert_equal [ caller_thread ], delivery_threads | ||
| assert Adapter.configuration.async? | ||
| end | ||
|
|
||
| def test_synchronous_scope_still_honours_sampling_and_skips_the_callback | ||
| configuration = ActiveAgents::Telemetry::Configuration.new | ||
| configuration.sample_rate = 0.0 | ||
| subscribe(configuration: configuration) | ||
| ids = [] | ||
|
|
||
| Adapter.with_agent("Support", synchronous: true, on_trace: ->(trace) { ids << trace.trace_id }) do | ||
| instrument("chat.ruby_llm", chat_payload) { nil } | ||
| end | ||
|
|
||
| assert_empty posted | ||
| assert_empty ids | ||
| end | ||
|
|
||
| def test_an_unscoped_turn_never_adopts_the_scope_whose_chat_flushes_it | ||
| ids = [] | ||
| instrument("chat.ruby_llm", chat_payload(tool_call: true)) { nil } | ||
| assert_empty posted, "the unscoped turn stays open on its pending tool call" | ||
|
|
||
| Adapter.with_agent("Judge", action: "score", attributes: { "eval.run_id" => "run-1" }, | ||
| on_trace: ->(trace) { ids << trace.trace_id }, synchronous: true) do | ||
| instrument("chat.ruby_llm", chat_payload(chat: Object.new)) { nil } | ||
| end | ||
|
|
||
| roots = traces.map { |trace| spans_of(trace, "root").first } | ||
| assert_equal [ "RubyLLM::Chat.chat", "Judge.score" ], roots.map { |root| root["name"] } | ||
| assert_equal [ nil, "run-1" ], roots.map { |root| root["attributes"]["eval.run_id"] } | ||
| assert_equal [ traces.fetch(1)["trace_id"] ], ids, "only the judge's own trace reaches its callback" | ||
| end | ||
|
|
||
| def test_a_turn_keeps_the_scope_it_started_under_when_flushed_later | ||
| ids = [] | ||
| pending = chat_payload(tool_call: true) | ||
| Adapter.with_agent("Judge", action: "score", attributes: { "eval.run_id" => "run-1" }, | ||
| on_trace: ->(trace) { ids << trace.trace_id }) do | ||
| instrument("chat.ruby_llm", pending) { nil } | ||
| end | ||
| assert_empty posted, "a turn with a pending tool call stays open" | ||
|
|
||
| Adapter.flush!(pending) | ||
|
|
||
| trace = traces.fetch(0) | ||
| root = spans_of(trace, "root").fetch(0) | ||
| assert_equal "Judge.score", root["name"] | ||
| assert_equal "run-1", root["attributes"]["eval.run_id"] | ||
| assert_equal [ trace["trace_id"] ], ids | ||
| end | ||
|
|
||
| def test_callback_failure_does_not_discard_the_trace_or_expose_its_message | ||
| _, stderr = capture_io do | ||
| Adapter.with_agent("Support", on_trace: ->(_) { raise "private callback content" }) do | ||
| instrument("chat.ruby_llm", chat_payload) { nil } | ||
| end | ||
| end | ||
|
|
||
| assert_equal 1, posted.size | ||
| assert_includes stderr, "on_trace failed: RuntimeError" | ||
| refute_includes stderr, "private callback content" | ||
| end | ||
| end |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,6 @@ | ||
| # Branch | ||
|
|
||
| - Branch: `codex/evaluation-trace-context`. | ||
| - Base: `main` at `ccbab0ffd2b26722685d8d2f7f7ff74a45466373`. | ||
| - Scope: backward-compatible RubyLLM trace context and delivery control for evaluations. | ||
| - Release: consumers can pin this branch's reviewed commit before the next adapter release. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,17 @@ | ||
| # Evaluation trace identity and completion | ||
|
|
||
| Tool-less evaluation judge calls previously inherited an application's generic | ||
| tool-less identity. Applications also had no supported callback for retaining the | ||
| actual emitted trace ID in a scenario result, and short-lived commands inherited | ||
| asynchronous delivery that could outlive the process. | ||
|
|
||
| `RubyLLM.with_agent` now accepts scope-local correlation attributes, an `on_trace` | ||
| callback and a synchronous delivery option. Existing calls remain compatible. | ||
| Nested/raising scopes restore their previous context; callback failures do not | ||
| drop traces or log callback message content. Body capture remains opt-in. | ||
|
|
||
| Validation: core 39 tests / 100 assertions and RubyLLM adapter 20 tests / 71 | ||
| assertions pass on Ruby 4.0.2. New tests cover actual trace-ID correlation, judge | ||
| identity, nested context restoration, synchronous delivery under sampling, a | ||
| turn keeping the scope it started under, and callback failure. | ||
| Raw logs live in gitignored `tmp/`. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,8 @@ | ||
| # Milestone: evaluation correlation | ||
|
|
||
| - [x] Separate candidate and judge trace identity. | ||
| - [x] Expose actual completed trace IDs for report correlation. | ||
| - [x] Support blocking delivery attempts for short-lived commands. | ||
| - [x] Test restoration, callback failures, and existing adapter compatibility. | ||
| - [x] Bump both gems to 0.3.0 with a changelog entry, so the merge is releasable. | ||
| - [ ] Merge the PR, then tag `v0.3.0` to publish both gems through the release workflow. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,12 @@ | ||
| # Draft PR: evaluation trace context | ||
|
|
||
| Branch `codex/evaluation-trace-context` proposes scoped correlation attributes, | ||
| completed trace callbacks, and blocking delivery attempts for RubyLLM calls. | ||
| It keeps existing agent resolvers and ordinary asynchronous application tracing | ||
| compatible. Consumers can distinguish evaluation judges from application agents | ||
| and link report results to their actual response and judge traces. | ||
|
|
||
| Core and adapter suites pass: 59 tests, 171 assertions, zero failures. Changed | ||
| Ruby files pass the Omakase lint configuration. No provider or collector network | ||
| requests are made by the tests. Versions and the changelog are bumped to 0.3.0 | ||
| on the branch; tagging `v0.3.0` after the merge publishes both gems. |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.