Skip to content

chore(main): release trg 0.10.0 - #112

Open
sht-bot wants to merge 1 commit into
mainfrom
release-please--branches--main--components--trg
Open

chore(main): release trg 0.10.0#112
sht-bot wants to merge 1 commit into
mainfrom
release-please--branches--main--components--trg

Conversation

@sht-bot

@sht-bot sht-bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

An automated release has been created for you.

0.10.0 (2026-09-13)

Features

  • trg: A case needs a directory before it needs more fields (#146) (dd3d8d3)
  • trg: A harness must declare what it cannot do (#145) (ea5485a)
  • trg: A judge asked once cannot say whether it is sure (#138) (6624297)
  • trg: A pass that names no scenario gets a baseline to read against (#137) (e850e7b)
  • trg: Bound what one eval pass may spend (#135) (d8329d3)
  • trg: Declare checks as data and report what a harness cannot answer (#109) (f4261dd)
  • trg: Draw every eval cell more than once by default (#129) (7202f91)
  • trg: Execute more than one eval run at a time (#128) (8c2659c)
  • trg: Let a case ask whether the skill gets reached for at all (#123) (6be74b2)
  • trg: Let a case say what state it is asking about (#134) (f0c7a5a)
  • trg: Let a case state that a tool must not be reached for (#130) (3656f18)
  • trg: Let the judge address any OpenAI-compatible endpoint (#108) (ef07863)
  • trg: Observe every harness's tool calls and report what left the workspace (#116) (728d658)
  • trg: Read every harness transcript through one event vocabulary (#107) (a347293)
  • trg: Say which command a case expected, not just which tool (#131) (7cd19d2)
  • trg: Scope graders to the arm they can answer in (#122) (10caa5a)
  • trg: Select part of an eval suite per run (#125) (78a369c)

Bug Fixes

  • trg: A cached run must answer for the draw that asked (#127) (0f7b5e5)
  • trg: A cached run must not answer for an eval case it never ran (#113) (872a9f2)
  • trg: A failed run must not read as a skill that failed its assertions (#115) (8f23d65)
  • trg: A manifest that names no version should not get the narrowest one (#142) (082f368)
  • trg: A pass nobody can price should not report a total of zero (#139) (20312d8)
  • trg: A pass-rate gate nothing was measured against must not report green (#132) (66e7f68)
  • trg: A reused run must answer for the arm that asked (#126) (b568c8e)
  • trg: A run inherits nothing by accident (#119) (8d49745)
  • trg: A run must not be able to read its own answer key (#118) (aac9bcd)
  • trg: A run where nothing could be scored has no pass rate (#114) (39af3b4)
  • trg: A run's permissions must come from the eval, not the machine (#144) (7cc7526)
  • trg: A score that could have been copied says so (#124) (0ca1c52)
  • trg: A timed out run must not leave agents running (#120) (a8188e9)
  • trg: A version nobody reads is not a contract (#141) (9cfaa4d)
  • trg: Docs must not tell consumers to gate on a version that is gone (#143) (0cb60b0)
  • trg: Stop eval verify and run reuse from misreporting results (#106) (46304a5)
  • trg: Whether a bundle conforms cannot depend on how the suite scored (#133) (c11036a)

This PR was generated with Release Please. See documentation.

@cursor

cursor Bot commented Sep 12, 2026

Copy link
Copy Markdown

PR Summary

Low Risk
The diff is version and changelog bookkeeping only; no application source changes in this PR.

Overview
Automated Release Please PR that cuts trg 0.10.0: it bumps the crate version in crates/trg/Cargo.toml, syncs Cargo.lock and .github/release-please-manifest.json, and adds a 0.10.0 section to crates/trg/CHANGELOG.md.

The new changelog entry documents what landed since 0.9.0—eval execution (parallel runs, multi-draw cells, suite subset selection, spend caps), richer case/check declarations (skills, tools, commands, state, harness capabilities), unified transcript events and tool-call observability, OpenAI-compatible judge endpoints, and many fixes around cached/reused runs, pass-rate and cost reporting, run isolation (permissions, answer keys, timeouts), and bundle conformance vs scoring.

Reviewed by Cursor Bugbot for commit 861ed26. Bugbot is set up for automated code reviews on this repo. Configure here.

@coderabbitai

coderabbitai Bot commented Sep 12, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 1d5a427b-e1cc-4f62-afc6-2715629ec5d2


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sht-bot
sht-bot force-pushed the release-please--branches--main--components--trg branch 22 times, most recently from dae95fe to 33043d4 Compare September 13, 2026 04:49
Signed-off-by: SHT Bot <61149376+sht-bot@users.noreply.github.com>
@sht-bot
sht-bot force-pushed the release-please--branches--main--components--trg branch from 33043d4 to 861ed26 Compare September 13, 2026 06:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant