Skip to content

test: add provider usage-normalization, compaction-telemetry, prompt-prefix coverage - #962

Merged
agentforce314 merged 1 commit into
agentforce314:mainfrom
fxinfo24:test/add-uncovered-provider-and-compaction-tests
Oct 4, 2026
Merged

agentforce314 merged 1 commit into
agentforce314:mainfrom
fxinfo24:test/add-uncovered-provider-and-compaction-tests

Conversation

@fxinfo24

@fxinfo24 fxinfo24 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Three test modules that were written against existing behaviour but never committed — so they have not been running in CI. All pass on Python 3.10 and 3.11.

Module Tests Covers
tests/providers/test_usage_normalization.py 30 normalizeUsage() across every provider
tests/test_compaction_telemetry.py 17 Cache-hostile compaction detection
tests/prompt/test_prefix_stability.py 15 Prompt-prefix stability

Why the usage-normalization matrix matters

The tests pin each provider's normalizeUsage() to the Anthropic wire convention:

  • input_tokens — cache MISS, priced at the input rate
  • cache_read_input_tokens — cache HIT, priced at the cache-read rate
  • cache_creation_input_tokens — cache WRITE, priced at the cache-write rate

This is the audit that catches an OpenAI-wire-format provider folding prompt_tokens_details.cached_tokens back into prompt_tokens. That mistake is silent — the request still succeeds and token totals still look plausible — but it overstates cache-miss spend, which is precisely the number prompt-caching economics depend on.

One fix beyond adding files

TestOpenAICompatibleProvider is a helper base class, not a test class. pytest collects by the Test prefix, so it emitted a PytestCollectionWarning on every run (it also inherits an __init__ from MockMessage).

Marked it __test__ = False. Nothing subclasses it, so this suppresses no real tests — verified by confirming zero subclasses before applying.

Verification

  • These three files: 62 passed on both Python 3.11 (CI's pinned version) and 3.10 (the declared floor).
  • Full suite on this branch: 10889 passed, 15 skipped, 4 failed.

The 4 full-suite failures are pre-existing, not from this PR

Confirmed by running the same tests on an untouched fork/main checkout.

Three failures in the full-suite run, all in test_init_integration.py, are the SIGINT/SIGTERM startup race that #961 fixes:

  • TestSigtermPathRunsCleanups::test_sigterm_triggers_drain
  • TestSigintDuringPrefetch::test_sigint_during_prefetch_clean_exit
  • TestSigintBeforePrefetchStarted::test_sigint_before_any_prefetch

(The third passes in some runs — the race is timing-dependent, so it appears when that file is run on its own more reliably than in a full-suite run. Same root cause.)

Two further failures reproduce on pristine fork/main with none of my changes applied, and are unrelated to this branch:

Failing test Cause
workflow/test_agent_concurrency.py::test_concurrent_agents_run_in_parallel Timing assertion (9.76s < 1.0s)
bridge/test_bridge_main.py::test_spawn_worktree_creates_isolated_dir_and_cleans_up Spawner race (assert 0 == 1)

Note

No production code is touched here — this is tests only, all passing against current behaviour on a clean checkout.

…prefix coverage

Three test modules written against existing behaviour but never committed, so
they have not been running in CI. All pass on Python 3.10 and 3.11.

- `tests/providers/test_usage_normalization.py` (30 tests) pins
  `normalizeUsage()` across every provider against the Anthropic convention:
  input_tokens for cache MISS, cache_read_input_tokens for HIT, and
  cache_creation_input_tokens for WRITE. This is the audit that catches an
  OpenAI-wire-format provider folding `prompt_tokens_details.cached_tokens`
  back into `prompt_tokens`, which silently overstates cache-miss spend.
- `tests/test_compaction_telemetry.py` (17 tests) covers cache-hostile
  compaction detection.
- `tests/prompt/test_prefix_stability.py` (15 tests) covers prompt-prefix
  stability.

Marks `TestOpenAICompatibleProvider` with `__test__ = False`. It is a helper
base rather than a test class, and pytest collects by the `Test`` prefix, so
it emitted a PytestCollectionWarning on every run. Nothing subclasses it, so
this suppresses no real tests.
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Test Results

     5 files   1 021 suites   21m 3s ⏱️
16 166 tests 16 144 ✅ 22 💤 0 ❌
32 302 runs  32 231 ✅ 71 💤 0 ❌

Results for commit 26f7737.

@agentforce314
agentforce314 merged commit 17c950b into agentforce314:main Oct 4, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants