Skip to content

Add SayGM provider - #4626

Open
markdavison wants to merge 11 commits into
anomalyco:devfrom
taostat:add-saygm-provider-final
Open

markdavison wants to merge 11 commits into
anomalyco:devfrom
taostat:add-saygm-provider-final

Conversation

@markdavison

@markdavison markdavison commented Aug 13, 2026 •

Copy link
Copy Markdown

Summary

Adds SayGM to the provider catalogue with 17 reviewed models and hourly price synchronization from its public model endpoint.

SayGM can route a request to any currently eligible miner, so neither the cheapest offer nor published retail is the best single catalogue number:

  • price_range.floor can understate the actual bill when a more expensive eligible route serves the request.
  • price_range.ceiling is the published-retail absolute cap, but can materially overstate every route currently in the pool.
  • price_range.route_ceiling is the shallowest-discount (highest-priced) currently eligible route. The synchronizer publishes this conservative live bound; an individual request may settle lower.

Synchronization remains review-only: it updates the 17 reviewed models, does not automatically add IDs or remove temporarily unavailable models, retains authored optional cost fields when omitted, and falls back to the explicit published-retail ceiling when a response has no live route ceiling. It never uses the optimistic headline pricing block as the catalogue cost.

Pricing sources

  • Authoritative model and price source: GET https://api.saygm.com/v1/models — unauthenticated and live.
  • Catalogue field: price_range.route_ceiling.dimensions, with basis: "highest_eligible_offer". It is computed from the shallowest-discount route in the same eligibility snapshot as availability and the floor, using SayGM's actual discount, buyer-markup, integer-rounding, and retail-clamp settlement arithmetic.
  • Absolute fallback: price_range.ceiling.dimensions, with basis: "published_retail". This is used only when the live route ceiling is absent.
  • Human-readable model metadata: https://saygm.com/models/<id>, for example https://saygm.com/models/claude-sonnet-5.

The checked-in costs were resynchronized on 2026-09-03 from the live public route_ceiling. That refresh includes Claude Sonnet 4.6's post-promotion $15/Mtok output basis and GPT-5.5 Pro's 272K long-context tier. The endpoint exposes the field for every currently listed model, and a live sync dry run leaves all 17 reviewed entries unchanged.

Verification

  • bun validate
  • 7 focused SayGM synchronization tests
  • bun run saygm:sync --dry-run: 17 reviewed models unchanged against the live endpoint
  • 28 live completions across five models and both API shapes: every settled bill was at or below the route ceiling read at the start and end of the run
  • The containing sync test file has one unrelated existing DeepInfra failure (DeepInfra preserves live modalities for new base models)

Tip: I will respond to comments that @ mention @cyrusagent on this PR/MR. You can also submit a review with all your feedback at once, and I will automatically wake up to address each comment.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [low] [possible mistake] providers/saygm/models/claude-sonnet-4-6.toml:19 - Check: Data-changing provider additions should cite first-party pricing/docs/API references mapped to claims. Why: Several SayGM prices diverge from lab list rates used elsewhere in the catalog (e.g. Claude Sonnet 4.6 output 10 vs Anthropic 15; Claude Sonnet 5 input 3 vs Anthropic 2). The PR describes a published-retail ceiling from the live models API but does not cite that endpoint or other sources tied to those figures, so the host-specific deltas cannot be reviewed from the diff alone. Action: In the PR body, cite the public models/pricing sources (for example https://api.saygm.com/v1/models and any pricing docs) and state what each supports, especially every cost that does not match the corresponding lab list price.

@markdavison

Copy link
Copy Markdown
Author

Added a "Pricing sources" section to the PR body citing the sources and addressing the two examples above.

Summary: GET https://api.saygm.com/v1/models (unauthenticated) is the authoritative source; each model also has a human-readable page at https://saygm.com/models/<id>. The API returns both a price_range.ceiling (basis: "published_retail", what buyers are charged — this is what the TOMLs store) and a price_range.floor (basis: "cheapest_eligible_offer", the current wholesale price paid to the winning miner). The floor moves continuously with miner bidding; the ceiling is SayGM's own retail rate card, stable enough to check in, and kept current by the hourly in-repo synchroniser (packages/core/src/sync/providers/saygm.ts).

On the specific deltas: SayGM's retail rate card is set per model/dimension, not pegged line-by-line to each lab's list, so it can land above or below list depending on the model. claude-sonnet-4-6 output is below Anthropic's list (10.00 vs 15.00); claude-sonnet-5 input is currently above Anthropic's list (3.00 vs 2.00) even though its wholesale floor (2.0301) sits close to that list price — the gap to the 3.00 ceiling is SayGM's margin, not a wholesale increase. So this isn't a uniform "always below list" relationship; it varies by model and dimension, and the PR body now says so explicitly rather than generalizing.

@markdavison
markdavison force-pushed the add-saygm-provider-final branch from 64f255d to 18bfbff Compare August 14, 2026 13:26
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [violation] packages/core/src/sync/providers/saygm.ts:229 - Check: Sync must preserve cost fields the API is not currently asserting, matching peer providers (e.g. xAI/OpenRouter tier retention, ofox cost spread). Why: buildSaygmModel rebuilds cost from scratch whenever price_range.ceiling exists. Optional dimensions (cache_*, audio, long_context_*) are omitted from the new object when absent, so stripUndefined + formatToml drop previously reviewed cache_read / cache_write / [[cost.tiers]] on the next hourly run even though input/output still validated. That can silently strip long-context tiers from models like gpt-5.4 / gpt-5.5 / gpt-5.6-* and Claude cache prices whenever the ceiling payload is partial. Action: When building cost from a ceiling, spread the existing authored cost (or at least retain tiers / optional cache/audio keys) unless the API explicitly supplies replacements; only clear tiers when long-context pricing is affirmatively absent (xAI-style), and add a regression test for “ceiling present, long-context/cache fields omitted → keep last reviewed optional cost fields.”

@markdavison

Copy link
Copy Markdown
Author

Addressed the sync-bot finding on buildSaygmModel in 572dfc9: the cost object was rebuilt as a fresh literal each sync, so any optional pricing field (cache_read, cache_write, audio rates, tiers) that the live API happened to omit was silently dropped from the reviewed catalog entry. It now spreads existing.cost and only overwrites a field when the API dimensions supply a value, so an omitted dimension keeps the authored value — the same omitted-means-keep-authored convention as packages/core/src/sync/providers/xai.ts (and the authored.cost spread in ofox.ts). input/output remain unconditionally taken from the live ceiling. Added a regression test covering a ceiling that omits cache and long-context dimensions.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [possible mistake] providers/saygm/models/claude-sonnet-5.toml:1 - Check: Toggle wire comments must document this host’s real request field/values (AGENTS.md Reasoning options; audit-reasoning-options). Why: The file claims thinking.type = enabled|disabled, but SayGM serves this model on the Anthropic Messages path ([provider].npm = "@ai-sdk/anthropic"), and same-surface peers for claude-sonnet-5 document adaptive thinking as thinking.type = adaptive|disabled (e.g. Neon, Azure). A wrong on-state value misleads clients configuring the toggle. Action: Confirm the live SayGM Messages wire for this model and update the leading comment to the exact accepted values (likely adaptive|disabled if it mirrors Anthropic); keep toggle + effort only if that on/off control is real on SayGM.

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added reviewer: ready Automated review found no actionable items and removed reviewer: ready Automated review found no actionable items labels Aug 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [possible mistake] packages/core/src/sync/providers/saygm.ts:1014 - Check: Catalog cost must be the USD/MTok price API buyers are charged. Why: Patch 5 publishes pricing.basis = "cheapest_eligible_offer" for all 17 models, but the PR body still defines that basis as the miner wholesale floor and published_retail as what buyers are charged. If that original split is still true, the catalog will systematically understate end-user price. Action: Confirm against SayGM’s billing//v1/models docs which basis is the buyer-facing rate; keep or switch the sync + seeded TOML costs to that basis only; update the PR pricing section so it matches the code.
  • [low] [violation] .pr-review/pull-request.json (PR body) - Check: Data-changing PRs should cite the price source actually used. Why: The body still documents price_range.ceiling / published_retail as authoritative, while the final diff reads pricing / cheapest_eligible_offer and rejects retail as a publishable price. Reviewers cannot map claims to the shipped behavior. Action: Rewrite the pricing sources section for the final policy (field path, basis, retention on published_retail/missing pricing, and what each model page supports).

@roykollensvendsen

Copy link
Copy Markdown
Contributor

I ran a small measurement against the question the bot left open on 2026-08-23 (which basis the buyer is charged), because I have an interest in this entry landing.

Method: 28 completions on 2026-09-01 (10:19 to 10:45 UTC), five models, both API shapes (/v1/chat/completions with Bearer, /v1/messages with x-api-key), max_tokens 8, one-line prompts. For each response I took usage.cost_nano_usd (the docs call it "the amount actually settled against the account's prepaid balance, not an estimate") and compared it with the tokens priced at price_range.floor (basis cheapest_eligible_offer, which is also the pricing block the sync reads) and at price_range.ceiling (basis published_retail), both read from /v1/models at the start of the run. Per-dimension integer division reproduces the floor bills to the nano-dollar (13 x 102794000 / 1e6 = 1336, 4 x 642462500 / 1e6 = 2569, sum 3905).

model n settled = floor settled above floor above-floor bill / floor above-floor bill / ceiling
gpt-5.4-nano 8 4 4 1.77, 1.79 0.91, 0.92
claude-haiku-4-5 6 4 2 2.58 0.93
glm-5.3-flash 4 3 1 2.38 0.98
kimi-k3 5 4 1 1.16 0.82
deepseek-v4-flash-0731 5 4 1 2.73 0.41

So neither basis is the buyer's price. The buyer pays the price of the miner that served the request (x-gm-model-served names the upstream, x-gm-cascade-index was 0 every time). In 19 of 28 responses that was the cheapest eligible offer, so pricing / floor was exactly right; in 9 of 28 a second, more expensive route served, at 0.41 to 0.98 of retail, and the bill was 1.16 to 2.73 times what the floor predicts. The two price points recur per model within minutes of each other (gpt-5.4-nano alternated 3905 and 6821 nano-USD for the same prompt), so this is routing between miners, not a price change over time.

For a catalogue that needs one number per model, the floor understates the bill about a third of the time by up to 2.7x; the ceiling never understates it and overstates it by up to 6.6x (deepseek at 0.15 of retail). Whichever the PR keeps, the header and the body should say which it is and that the settled price sits between the two; today the code publishes the floor and the body describes the ceiling.

Two smaller things from the same reads: the seeded TOML costs no longer match /v1/models on any of the 17 models (the sync would fix them on the first run, but the bot compares the diff against the body), and the live list has grown to 45 models, 28 of them not in this PR (Gemini, and the open and -tee models: glm, kimi, qwen, deepseek among them).

Raw samples and the script are in a public record if useful: #.

@markdavison
markdavison force-pushed the add-saygm-provider-final branch from 1c6c2bf to 723dead Compare September 1, 2026 11:54
@markdavison

Copy link
Copy Markdown
Author

@roykollensvendsen Thanks — your measurement identified the missing bound correctly. The cheapest eligible offer understates requests that land on another route, while retail can substantially overstate every route currently available.

I've updated this PR to publish a new conservative live bound: price_range.route_ceiling (basis: "highest_eligible_offer"). SayGM computes it from the shallowest-discount currently eligible route in the same router snapshot as the floor and availability, using the same settlement arithmetic as billing. A request may settle between floor and route_ceiling; published retail remains the absolute cap and compatibility fallback.

The 17 TOMLs are now seeded from the current production offer snapshot, the synchronizer and tests read the new field, and the PR body matches the shipped policy. The GM change needs to deploy before this PR merges so hourly syncs do not temporarily fall back to retail.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 1, 2026
@markdavison

Copy link
Copy Markdown
Author

I repeated the same 28-completion shape against the now-live route_ceiling on 2026-09-01 (14:22 to 14:40 UTC): the same five models and request counts, both /v1/chat/completions with Bearer auth and /v1/messages with x-api-key, one-line prompts, and max_tokens: 8. I read price_range.floor, price_range.route_ceiling, and published-retail price_range.ceiling from GET https://api.saygm.com/v1/models at the start and end; none changed during the run. There were no retries.

model n settled above floor settled at/below route ceiling max bill / floor max bill / route ceiling served route
gpt-5.4-nano 8 2 8 1.79 0.99 openai:gpt-5.4-nano
claude-haiku-4-5 6 3 6 2.56 0.99 anthropic:claude-haiku-4-5
glm-5.3-flash 4 2 4 2.38 1.00 zai:glm-5.3-flash
kimi-k3 5 3 5 1.42 1.00 moonshot:kimi-k3
deepseek-v4-flash-0731 5 0 5 1.00 0.21 deepseek:deepseek-v4-flash-0731
total 28 10 28 2.56 1.00

This reproduces the original problem: 10/28 requests settled above the cheapest-offer floor, by up to 2.56x. The new catalogue number held for all 28/28 requests, including the more expensive routes, while allowing cheaper requests to settle below it. That is the intended meaning of route_ceiling: a conservative bound over the currently eligible pool, not a prediction that every request will cost exactly that amount.

For arithmetic at this tiny token count, six floor-route settlements were one nano-dollar below a per-dimension reconstruction because of integer rounding; they are not counted as above-floor calls.

@roykollensvendsen

Copy link
Copy Markdown
Contributor

Confirmed independently, two days later and on a different price snapshot. Same request shape as before (max_tokens: 8, one-line prompts, four models, both API shapes, price_range read before and after and unchanged during the run): 15 successful completions on 2026-09-03 06:26 to 06:31 UTC, all 15 at or below route_ceiling, 5 of them above the floor by up to 2.73x (deepseek-v4-flash-0731, 2529 nano-USD against a floor of 1065). Highest settlement was 0.927 of route_ceiling on glm-5.3-flash, lowest 0.359 on deepseek. One SSL EOF, no retries. Spend $0.00009. So the bound holds on a run neither of us designed around it.

price_range.route_ceiling is live on all 50 rows of the public list this morning, including the five ids added since 09-01.

Two observations from the same read, both arguments for the field rather than against it:

route_ceiling equals published retail on 14 of 50 models, among them glm-5.2, glm-5.3, glm-5.3-flash, kimi-k3, deepseek-v4-flash-0731 and six of the -tee ids. For those the catalogue number is retail whichever field the sync reads, so the change is a no-op there and the seeded values stay comparable with the lab entries. On the other 36 it sits below retail, which is where it earns its keep.

The spread it has to cover is wide on a few models: route_ceiling / floor is 6.6 on deepseek-v4-flash-0731, 2.8 on qwen3.6-35b-a3b, 2.5 to 2.6 on the Claude family and glm-5.2. My deepseek settlements came in at 0.36 to 0.41 of the bound. A conservative bound is the right call for a catalogue that has one number, and worth one line in the body so a reader knows the published cost is an upper bound over the eligible pool rather than a typical bill.

Thanks for turning this into an API field rather than a note in a PR body. It makes the number checkable by anyone.

@roykollensvendsen

roykollensvendsen commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Separate from the pricing thread, in case it is useful while this is open: I mapped what the live list carries against this PR, since I had the catalogue open anyway.

GET /v1/models returned 50 ids this morning (2026-09-03 06:26 UTC), against the 17 here. The 33 outside the PR:

group n ids
confidential (-tee, chat.completions) 14 deepseek-v3.2-tee, deepseek-v4-flash-0731-tee, gemma-4-31b-turbo-tee, glm-5.1-tee, glm-5.2-tee, kimi-k2.6-tee, kimi-k3-tee, mistral-nemo-instruct-2407-tee, nemotron-3-nano-omni-30b-tee, qwen3-235b-a22b-thinking-2507-tee, qwen3-32b-tee, qwen3.5-397b-a17b-tee, qwen3.6-27b-tee, qwen3.8-27b-tee
open (chat.completions) 9 deepseek-v4-flash-0731, glm-5.2, glm-5.3, glm-5.3-flash, kimi-k3, ornith-1.5-397b, qwen3.6-35b-a3b, qwen3.8-27b, qwen3.8-flash-next
frontier (generateContent) 9 gemini-3.1-flash-image, gemini-3.1-flash-lite, gemini-3.1-flash-lite-image, gemini-3.1-pro-preview, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash
frontier (messages) 1 claude-fable-5-1

Five of those appeared between 09-01 and today (claude-fable-5-1, gemini-3.7-flash, gemini-3.8-flash and the two -image ids), so the list is moving faster than the PR.

What I checked while mapping them, in case any of it saves you time:

Lab files. 31 of the 33 point at an existing models/<lab>/ entry. Three need a name that is not the obvious one, with precedent in providers/chutes: gemma-4-31b-turbo-tee to google/gemma-4-31b-it, mistral-nemo-instruct-2407-tee to mistral/mistral-nemo, nemotron-3-nano-omni-30b-tee to nvidia/nemotron-3-nano-omni-30b-a3b-reasoning. Two have no lab file at all and would need one added rather than inlined: deepreinforce/ornith-1.5-397b (GM is the catalogue's first host of that size) and alibaba/qwen3-235b-a22b-thinking-2507 (18 providers serve it today and all 18 inline it in full). Every family slug needed is already in family.ts. All 14 -tee ids have a one-for-one counterpart under providers/chutes/**/*TEE.toml, which is the closest same-surface peer set for reasoning options.

One thing that would make an entry hard to write honestly. On glm-5.3-flash the reasoning behaviour depends on which upstream serves the request, and the response headers show at least four in the pool at x-gm-cascade-index: 0 (an engy build via x-engy-version, an in-cluster SGLang behind litellm, Chutes on the TEE ids, and DeepSeek native). With chat_template_kwargs.enable_thinking: false, the litellm route moved the thinking into content behind a stray </think> and dropped reasoning_content 5 times out of 5; the engy route kept reasoning_content and merely shortened it, 0 leaks out of 3. Same split for reasoning_effort: "none". A reasoning_options value is one number per model, so whichever is authored is wrong on one route. Worth knowing regardless of the catalogue, since a client that turns thinking off gets </think> in its output on some requests.

Two smaller things from the same reads. capabilities.vision disagrees with the lab metadata on six ids in both directions: five -tee builds of text-only weights report vision: true (including deepseek-v4-flash-0731-tee, where GM's own non-TEE row for the same weights reports false), and qwen3.6-35b-a3b reports false where its lab entry is multimodal. And available flaps: three reads of the same endpoint on 09-01 showed 0, then 3, then 5 unavailable ids, with the basis falling back to published_retail while a model is unavailable.

None of this is a request to change this PR. If you would rather keep the scope where it is, say so and I will leave it alone; if the mapping is useful I can send it as a file.

Correcting myself on the value of the rest, since I checked after writing the above: I had assumed the confidential tier was the part of the list nothing else covers, and that is wrong. All 14 of the -tee ids already have a counterpart under providers/chutes and 10 of them under providers/nano-gpt as well. Of the 33 ids outside this PR, ornith-1.5-397b is the only one no other provider entry in the catalogue serves. So the case for adding the rest is convenience for SayGM users rather than coverage the catalogue lacks, which makes it your call about your own provider rather than a gap worth someone else filling.

One note on the state here: the bot has reviewer: ready on this, and the only thing showing red is the conflict. It is 148 commits behind dev and the collision is the groups object in packages/core/src/sync/index.ts, one line, keep both sides. sync.test.ts auto-merges.

markdavison and others added 6 commits September 3, 2026 15:26
SayGM serves this model on the Anthropic Messages surface, which uses
adaptive|disabled for claude-sonnet-5 (see azure and
azure-cognitive-services peer entries, and gm gateway's own
claude-opus-4-8 fixture at gateway/src/api/anthropic.rs:144).
gm's discount over each lab's own price is the pitch for routing through
it. Publishing the published-retail ceiling made every SayGM model look
identical to going direct, since the ceiling equals list price for
several of them. Switch the sync adapter to read pricing.dimensions
(basis cheapest_eligible_offer) instead of price_range.ceiling, and
seed every SayGM model TOML with the current live values.

The basis is still schema-validated as a literal: an unrecognized basis
fails the sync loudly, and the known retail-fallback basis (returned
when nothing is currently routable) is parsed but deliberately not
adopted as a price, so a temporary supply gap can't republish the cap
under a stale "current price" label. Both cases retain the last
reviewed cost instead, per the existing update-only philosophy.

This accepts drift between hourly syncs in exchange for a number that
is usually competitive rather than one that is stably wrong.
@markdavison
markdavison force-pushed the add-saygm-provider-final branch from 723dead to d6e9c01 Compare September 3, 2026 14:26
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 3, 2026
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Action items

  • [high] [possible mistake] providers/saygm/models/claude-sonnet-4-6.toml:14 - Check: Provider costs should reflect this host’s real USD/MTok settlement (and, for a retail-derived bound, the lab’s published list when that is the ceiling basis). Why: Final SayGM costs are input = 2.781 / output = 9.27 (≈0.927× a 3/10 base). Anthropic first-party and essentially every peer catalog use Sonnet 4.6 list 3 / 15, so the same 0.927 factor would be ≈2.781 / 13.905. The 9.27 output tracks the PR’s earlier hand-seeded output = 10, not Anthropic’s standard output list price, so the catalogue can systematically understate this model’s bill. Action: Re-check price_range.route_ceiling (and ceiling retail) for claude-sonnet-4-6 on GET https://api.saygm.com/v1/models; if retail/route output is wrong relative to Anthropic’s $15/MTok list, correct the TOML (and any sync mapping) so output matches the live SayGM bound off the real retail.
  • [medium] [possible mistake] providers/saygm/models/gpt-5.5-pro.toml:8 - Check: Long-context [[cost.tiers]] should be authored when this host exposes a context price band (same as sibling GPT entries and OpenAI). Why: OpenAI’s gpt-5.5-pro and other SayGM GPT files (gpt-5.4, gpt-5.5, gpt-5.6-*) carry a 272k context tier, but gpt-5.5-pro only has base cost with no tier. If SayGM’s model payload includes long-context dimensions for this id, the catalogue understates price above the threshold. Action: Confirm whether the live SayGM model object for gpt-5.5-pro includes long-context price dimensions; if yes, add the matching [[cost.tiers]] (and ensure the sync writes them), or document in a leading comment if this host truly has no long-context band for Pro-only Responses.

@markdavison

Copy link
Copy Markdown
Author

Resolved both findings in 29ff35266.

  • Sonnet 4.6: corrected SayGM's curated retail output from the expired $10 promotion to Anthropic's standard $15/Mtok, published it to mainnet, then resynced the live route_ceiling output as 13.905. Upstream fix: https://github.com/taostat/gm/pull/1191
  • GPT-5.5 Pro: added SayGM's published 272K tier and resynced the live route ceiling as 55.62 input / 250.29 output.

Fresh public GET https://api.saygm.com/v1/models shows both models available with those route-ceiling dimensions. The SayGM dry run is now 17 unchanged; its seven focused tests and bun validate pass. The full sync test file remains 188 pass / the same unrelated DeepInfra modality failure already noted in the PR body.

@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 3, 2026
…inal

# Conflicts:
#	packages/core/src/sync/index.ts
#	packages/core/test/sync.test.ts
The live SayGM list now includes a per-image SKU (flux.2-klein-4b) and a
coming-soon id with null pricing; either one failed the whole response
parse. Skip per-image SKUs, which models.dev cannot price, and accept
null pricing on unpriced ids.

Also map the price dimensions newer SayGM models publish: OpenAI's
unqualified cache-write rate, long-context cache read/write tiers, and
the image-output rate that Gemini image models bill generated images
at (the value lab catalogs publish as output). Stop publishing
SayGM's audio-output rate, which it lists on text-only models.
Add the 38 servable ids SayGM's public /v1/models now lists beyond the
original 17: Claude Fable 5.1 and Opus 5.5, GPT-6 Astra/Luna/Sol and
gpt-oss-20b, nine Gemini models on the Gemini surface, the open
DeepSeek/GLM/Kimi/MiMo/Qwen/Ornith models, and the 14 confidential
(-tee) ids. Costs for all 55 are resynced from price_range.route_ceiling.

Route gpt-5.6-* and gpt-6-* over the Responses API, since OpenAI
rejects function tools with reasoning effort on chat completions for
those models. Add lab entries for Ornith 1.5 397B and Qwen3 235B-A22B
Thinking 2507, which had none.
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 26, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/saygm/models/deepseek-v4-flash-0731.toml:13 - Check: Relay reasoning baseline must copy lab + same-surface peer effort set. Why: This file authors effort ["high","max"] while first-party providers/deepseek/models/deepseek-v4-flash.toml and relay peer providers/openrouter/models/deepseek/deepseek-v4-flash-0731.toml both expose ["low","high","max"], and the sibling deepseek-v4.1-flash.toml in this same PR uses ["low","high","max"]; dropping low understates caller control without host evidence. Action: Change to ["low","high","max"] to match lab/peers, or provide SayGM docs/live proof that this ID rejects low.
  • [medium] [violation] packages/core/src/sync/providers/saygm.ts:82 - Check: Sync must not retain stale fields it claims to stop publishing. Why: buildSaygmModel spreads ...existing.cost but no longer sets output_audio, so any existing output_audio survives indefinitely, contradicting the code comment and PR claim that only input_audio is published and audio-output is stopped. Action: Explicitly clear it (e.g. output_audio: undefined with stripping, or delete after spread) so the hourly sync removes stale audio-output rates instead of preserving them.
  • [medium] [possible mistake] packages/core/src/sync/providers/saygm.ts:68 - Check: Per-image SKU filter must cover all price shapes that can fail parsing. Why: isPerImagePriced only inspects pricing.dimensions for output_per_image_ndollars, but PriceRange.ceiling.dimensions can also carry per-image dimensions; a per-image ID with null pricing (like the coming-soon case this patch adds) would evade the filter and fail SaygmResponse.parse for the whole list because PriceRange requires input/output_per_mtok. Action: Also check price_range.floor/route_ceiling/ceiling.dimensions for output_per_image_ndollars in the filter, with a regression test for per-image + null pricing.
  • [low] [possible mistake] providers/saygm/models/ornith-1.5-397b.toml:12 - Check: Hand-authored limits on new provider models need verifiable sources. Why: This sets [limit] context 262_144 output 262_144 (output equals full context) while new lab models/deepreinforce/ornith-1.5-397b.toml has no output to inherit, and the header cites only the HF chat template for the toggle, not limits; output-equals-context mirrors a fallback pattern, not a measured SayGM/Chutes cap. Action: Verify output against SayGM/chutes serving metadata and cite the source in a leading header, or correct the value.

Detect per-image SKUs from any pricing or price_range block, so one
listed without headline pricing cannot fail the whole parse. Clear a
stale audio-output rate instead of retaining it. Match DeepSeek V4
Flash 0731's effort levels to its lab and relay peers, and take Ornith
1.5 397B's output limit from its model card.
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 26, 2026
@roykollensvendsen

roykollensvendsen commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

@markdavison Some data that might help this one land. engy (#5910) was stuck the same way: bot-clean, reviewer: ready, a month without a maintainer looking at it.

I counted new-provider PRs opened since late August that the bot had marked reviewer: ready. 25 of the 79 that added only the provider were merged, and 0 of the 9 that also included a sync module. Over the same weeks, sync modules for providers already in the catalogue were merged as their own PRs within 0 to 3 days (#7040, #7112, #7006, #6990, #6130). The 12 new providers merged in September ran from +29 to +499 lines.

So on 10-02 I cut #5910 down to the provider directory: one commit, 10 files, +228 lines. It was merged 2.5 hours later. The sync module then went up as #8715, unchanged.

Here that split would be providers/saygm/ plus the two new lab files under models/ (59 files, about +1,000) first, and then saygm.ts, its registration and the sync.test.ts cases (+553) once the provider is in.

Also note the stale closer, which closes a PR after 30 days without an update. This comment counts as one, so the next deadline is around 11-04.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants