Repository navigation
Add SayGM provider - #4626
Add SayGM provider#4626markdavison wants to merge 11 commits into
Conversation
Action items
|
|
Added a "Pricing sources" section to the PR body citing the sources and addressing the two examples above. Summary: On the specific deltas: SayGM's retail rate card is set per model/dimension, not pegged line-by-line to each lab's list, so it can land above or below list depending on the model. |
64f255d to
18bfbff
Compare
Action items
|
|
Addressed the sync-bot finding on |
Action items
|
|
No actionable findings. |
Action items
|
|
I ran a small measurement against the question the bot left open on 2026-08-23 (which basis the buyer is charged), because I have an interest in this entry landing. Method: 28 completions on 2026-09-01 (10:19 to 10:45 UTC), five models, both API shapes (
So neither basis is the buyer's price. The buyer pays the price of the miner that served the request ( For a catalogue that needs one number per model, the floor understates the bill about a third of the time by up to 2.7x; the ceiling never understates it and overstates it by up to 6.6x (deepseek at 0.15 of retail). Whichever the PR keeps, the header and the body should say which it is and that the settled price sits between the two; today the code publishes the floor and the body describes the ceiling. Two smaller things from the same reads: the seeded TOML costs no longer match Raw samples and the script are in a public record if useful: #. |
1c6c2bf to
723dead
Compare
|
@roykollensvendsen Thanks — your measurement identified the missing bound correctly. The cheapest eligible offer understates requests that land on another route, while retail can substantially overstate every route currently available. I've updated this PR to publish a new conservative live bound: The 17 TOMLs are now seeded from the current production offer snapshot, the synchronizer and tests read the new field, and the PR body matches the shipped policy. The GM change needs to deploy before this PR merges so hourly syncs do not temporarily fall back to retail. |
|
No actionable findings. |
|
I repeated the same 28-completion shape against the now-live
This reproduces the original problem: 10/28 requests settled above the cheapest-offer floor, by up to 2.56x. The new catalogue number held for all 28/28 requests, including the more expensive routes, while allowing cheaper requests to settle below it. That is the intended meaning of For arithmetic at this tiny token count, six floor-route settlements were one nano-dollar below a per-dimension reconstruction because of integer rounding; they are not counted as above-floor calls. |
|
Confirmed independently, two days later and on a different price snapshot. Same request shape as before (
Two observations from the same read, both arguments for the field rather than against it:
The spread it has to cover is wide on a few models: Thanks for turning this into an API field rather than a note in a PR body. It makes the number checkable by anyone. |
|
Separate from the pricing thread, in case it is useful while this is open: I mapped what the live list carries against this PR, since I had the catalogue open anyway.
Five of those appeared between 09-01 and today (claude-fable-5-1, gemini-3.7-flash, gemini-3.8-flash and the two What I checked while mapping them, in case any of it saves you time: Lab files. 31 of the 33 point at an existing One thing that would make an entry hard to write honestly. On Two smaller things from the same reads. None of this is a request to change this PR. If you would rather keep the scope where it is, say so and I will leave it alone; if the mapping is useful I can send it as a file. Correcting myself on the value of the rest, since I checked after writing the above: I had assumed the confidential tier was the part of the list nothing else covers, and that is wrong. All 14 of the One note on the state here: the bot has |
SayGM serves this model on the Anthropic Messages surface, which uses adaptive|disabled for claude-sonnet-5 (see azure and azure-cognitive-services peer entries, and gm gateway's own claude-opus-4-8 fixture at gateway/src/api/anthropic.rs:144).
gm's discount over each lab's own price is the pitch for routing through it. Publishing the published-retail ceiling made every SayGM model look identical to going direct, since the ceiling equals list price for several of them. Switch the sync adapter to read pricing.dimensions (basis cheapest_eligible_offer) instead of price_range.ceiling, and seed every SayGM model TOML with the current live values. The basis is still schema-validated as a literal: an unrecognized basis fails the sync loudly, and the known retail-fallback basis (returned when nothing is currently routable) is parsed but deliberately not adopted as a price, so a temporary supply gap can't republish the cap under a stale "current price" label. Both cases retain the last reviewed cost instead, per the existing update-only philosophy. This accepts drift between hourly syncs in exchange for a number that is usually competitive rather than one that is stably wrong.
723dead to
d6e9c01
Compare
Action items
|
|
Resolved both findings in
Fresh public |
|
No actionable findings. |
…inal # Conflicts: # packages/core/src/sync/index.ts # packages/core/test/sync.test.ts
The live SayGM list now includes a per-image SKU (flux.2-klein-4b) and a coming-soon id with null pricing; either one failed the whole response parse. Skip per-image SKUs, which models.dev cannot price, and accept null pricing on unpriced ids. Also map the price dimensions newer SayGM models publish: OpenAI's unqualified cache-write rate, long-context cache read/write tiers, and the image-output rate that Gemini image models bill generated images at (the value lab catalogs publish as output). Stop publishing SayGM's audio-output rate, which it lists on text-only models.
Add the 38 servable ids SayGM's public /v1/models now lists beyond the original 17: Claude Fable 5.1 and Opus 5.5, GPT-6 Astra/Luna/Sol and gpt-oss-20b, nine Gemini models on the Gemini surface, the open DeepSeek/GLM/Kimi/MiMo/Qwen/Ornith models, and the 14 confidential (-tee) ids. Costs for all 55 are resynced from price_range.route_ceiling. Route gpt-5.6-* and gpt-6-* over the Responses API, since OpenAI rejects function tools with reasoning effort on chat completions for those models. Add lab entries for Ornith 1.5 397B and Qwen3 235B-A22B Thinking 2507, which had none.
Action items
|
Detect per-image SKUs from any pricing or price_range block, so one listed without headline pricing cannot fail the whole parse. Clear a stale audio-output rate instead of retaining it. Match DeepSeek V4 Flash 0731's effort levels to its lab and relay peers, and take Ornith 1.5 397B's output limit from its model card.
|
No actionable findings. |
|
@markdavison Some data that might help this one land. engy (#5910) was stuck the same way: bot-clean, I counted new-provider PRs opened since late August that the bot had marked So on 10-02 I cut #5910 down to the provider directory: one commit, 10 files, +228 lines. It was merged 2.5 hours later. The sync module then went up as #8715, unchanged. Here that split would be Also note the stale closer, which closes a PR after 30 days without an update. This comment counts as one, so the next deadline is around 11-04. |
Summary
Adds SayGM to the provider catalogue with 17 reviewed models and hourly price synchronization from its public model endpoint.
SayGM can route a request to any currently eligible miner, so neither the cheapest offer nor published retail is the best single catalogue number:
price_range.floorcan understate the actual bill when a more expensive eligible route serves the request.price_range.ceilingis the published-retail absolute cap, but can materially overstate every route currently in the pool.price_range.route_ceilingis the shallowest-discount (highest-priced) currently eligible route. The synchronizer publishes this conservative live bound; an individual request may settle lower.Synchronization remains review-only: it updates the 17 reviewed models, does not automatically add IDs or remove temporarily unavailable models, retains authored optional cost fields when omitted, and falls back to the explicit published-retail ceiling when a response has no live route ceiling. It never uses the optimistic headline
pricingblock as the catalogue cost.Pricing sources
GET https://api.saygm.com/v1/models— unauthenticated and live.price_range.route_ceiling.dimensions, withbasis: "highest_eligible_offer". It is computed from the shallowest-discount route in the same eligibility snapshot as availability and the floor, using SayGM's actual discount, buyer-markup, integer-rounding, and retail-clamp settlement arithmetic.price_range.ceiling.dimensions, withbasis: "published_retail". This is used only when the live route ceiling is absent.https://saygm.com/models/<id>, for examplehttps://saygm.com/models/claude-sonnet-5.The checked-in costs were resynchronized on 2026-09-03 from the live public
route_ceiling. That refresh includes Claude Sonnet 4.6's post-promotion $15/Mtok output basis and GPT-5.5 Pro's 272K long-context tier. The endpoint exposes the field for every currently listed model, and a live sync dry run leaves all 17 reviewed entries unchanged.Verification
bun validatebun run saygm:sync --dry-run: 17 reviewed models unchanged against the live endpointDeepInfra preserves live modalities for new base models)