Skip to content

feat(engy): add engy provider - #5910

Open
roykollensvendsen wants to merge 1 commit into
anomalyco:devfrom
roykollensvendsen:add-engy
Open

roykollensvendsen wants to merge 1 commit into
anomalyco:devfrom
roykollensvendsen:add-engy

Conversation

@roykollensvendsen

@roykollensvendsen roykollensvendsen commented Aug 31, 2026 •

Copy link
Copy Markdown

Summary

engy (api.engy.ai) relays eight open models over one OpenAI-compatible chat-completions endpoint (@ai-sdk/openai-compatible). Each model file is override-only against its lab entry. The sync module that was in this PR before will follow as its own PR once the provider is in.

engy's maintainer has endorsed the entry: #5910 (comment). That covers the figures that can't be checked from outside, including the input/output split, the modality narrowing, and kimi-k3 honouring temperature.

Models

Model $/MTok in / out / cached Context / Input / Output reasoning_options
qwen3.6-35b-a3b 0.045 / 0.30 / 0.015 208,192 / 200,000 / 8,192 toggle
qwen3.8-27b 0.045 / 0.32 / 0.015 1,001,536 / 936,000 / 65,536 toggle + low|medium|xhigh
glm-5.2 0.68 / 1.50 / 0.18 262,144 / 229,376 / 32,768 toggle + high|max
glm-5.3 0.98 / 3.08 / 0.18 327,680 / 294,912 / 32,768 low|high|max
glm-5.3-flash 0.135 / 0.45 / 0.027 262,144 / 229,376 / 32,768 low|high|max
kimi-k3 1.95 / 9.75 / 0.195 1,113,088 / 1,047,552 / 65,536 toggle
deepseek-v4-flash-0731 0.045 / 0.09 / 0.009 1,048,576 / 920,576 / 128,000 toggle + low|high|max
deepseek-v4.1-flash 0.04 / 0.08 / 0.008 327,680 / 262,144 / 65,536 toggle + low|high|max
  • engy's context is max_input + max_output, so both halves are authored.
  • Image input was checked by sending an image. qwen3.6-35b-a3b is text-only on engy, so it sets attachment = false. qwen3.8-27b, glm-5.3-flash, kimi-k3 and deepseek-v4.1-flash take text and image, and the input list is narrowed wherever the lab lists more.
  • glm-5.3-flash bills 0.135 / 0.45 / 0.027 while the page shows its 0.15 / 0.50 / 0.03 list price struck through. "(Ox Alpha)" is a codename, so status is unset.

Reasoning

engy is a relay, so each model's baseline is its lab entry plus peers on the same kind of API. The toggle is chat_template_kwargs.enable_thinking for glm-5.2 and Qwen, and chat_template_kwargs.thinking for DeepSeek and Kimi. Effort is reasoning_effort. There is no budget field, so no budget_tokens.

The effort sets follow engy's published per-template mapping (engy.ai/docs). engy renders the effort level into the prompt as a sentence instead of enforcing a budget. Each toggle was checked live: it suppresses reasoning on every sample. Each file's header gives the sample sizes and p-values. Two exceptions:

  • glm-5.3 pair. Three rungs measured (n=40, low < high p=3.4e-06), and nothing switches reasoning off, so there is no toggle.
  • kimi-k3. Effort is inert (n=40, all pairs p>=0.61), so the file authors the toggle alone.

Sources

  • api.engy.ai/v1/models, last read 2026-10-02: ids, prices, context, modalities.
  • engy.ai/api/v1/models (authenticated), 2026-10-02: the input and output limits. kimi-k3's max_input rose from 983,040 (08-30) to 1,047,552, so its context is now 1,113,088 and is authored above the lab's 1M.
  • engy.ai/pricing, 2026-10-02: the doc link. engy's docs list no models.
  • engy.ai/docs: the reasoning_effort mapping and chat_template_kwargs.

Test plan

  • bun validate. Each resolved engy model checked against both model lists on 2026-10-02.
  • bun run test in packages/sdk
  • Rebased on dev c81646e
  • opencode: listed only with ENGY_API_KEY, and the cost matches engy's bill

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 1, 2026
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 3, 2026
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 3, 2026
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 4, 2026
@crabbytt

crabbytt commented Sep 10, 2026 •

Copy link
Copy Markdown

Endorsing this — I'm the owner of engy (api.engy.ai), and it's accurate.

I checked all seven files against what we serve. The prices, context/limit split, modalities, and reasoning controls are all correct — including the parts you can't verify externally: the authed input/output split (context = max_input + max_output), the per-deployment modality narrowing (qwen3.6 text-only; qwen3.8 / glm-5.3-flash / kimi-k3 text+image), kimi-k3 honoring temperature, the glm-5.3-flash launch-discount billing (0.135/0.45), and the recent glm-5.3 context bump + kimi-k3 30% price raise.

The sync safeguards are exactly right for our unauthenticated list: skipCreates (avoids limit.output = 0 files), deleteMissing: false (an empty/truncated 200 can't wipe hand-measured files), and price rounding so repricing doesn't churn.

engy.ai/pricing is the right doc link, since our /docs lists no models.

Great work — the probing and fixed-point tests are well above the bar.

@roykollensvendsen

Copy link
Copy Markdown
Author

Thanks for checking it line by line, and for confirming the parts that cannot be verified from outside. That settles the two questions the draft carried: the doc link stays on engy.ai/pricing, and the entry goes in from this fork rather than yours. I have updated the description to point at your comment.

Two things from the last few days that the entry now reflects, in case they are useful on your side: the sync module picked up glm-5.3's window going to 327,680 and kimi-k3's 30% raise on its own, which is what it is there for, and I re-checked all 21 prices, the seven limits and the modality lists against the live list this morning at build 0.2.19 with no drift.

One measurement you may want regardless of this PR: on qwen3.6-35b-a3b, when max_tokens ends the generation inside the thinking block, the response comes back with the partial thinking as content and finish_reason: "stop" rather than "length" (6/6 at cap 16, 5/5 at cap 200, across ten miners; at cap 3000 the answer is clean). A client that trusts finish_reason reads that as a completed turn. Same shape as the parser faults your own MINER.md documents for tool calls. Happy to send the request ids.

@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 12, 2026
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 12, 2026
@roykollensvendsen

Copy link
Copy Markdown
Author

@crabbytt deepseek-v4.1-flash is on the entry now (head 828aa15, bot clean). Same sources as before: 0.04 / 0.08 / 0.008 from /v1/models, 262,144 in and 65,536 out from the authenticated list, text+image checked by sending an image.

Two things only you can settle:

  • Effort. On 0.2.20, low, medium and high measure as one rung and xhigh and max as another (n=40 per level, high vs max paired p=1e-04), the same shape as 0731, so the file authors high|max. Your docs' DeepSeek V4 column sends low and high to different levels; if v4.1-flash is meant to get that mapping, low should separate from high, and at n=40 it does not.
  • The toggle. On v4.1-flash both chat_template_kwargs.thinking and enable_thinking switch reasoning (40/40 each way); on 0731 enable_thinking still looked inert in the one check I ran today. The header names thinking, which both honour. Say if you would rather it named the other.

Two observations from the run, both on 0.2.20, in case they are useful: the x-engy-miner and x-engy-request-id headers are absent on every response slower than about 15 s (114 of 480; the body's x_engy block had them on the one slow response I checked), and completion_tokens_details.reasoning_tokens is 0 on responses with a full reasoning_content, while completion_tokens and the bill include them.

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 13, 2026
@andy-engy

Copy link
Copy Markdown

Answering from the engy side.

Thanks — all four are answerable. Measurements are from 2026-09-15.

Effort: your measurement is right, the conclusion needs one correction

The ladder is four real tiers, not two — low=25, high=50, xhigh=75, max=100, plus none to turn thinking off. But it is rendered into the prompt as a sentence ("Reasoning Effort: N (range 1-100, the higher the value, the more thorough the reasoning)"), not enforced as a token budget. So:

  • medium genuinely collapses onto high — there is no medium tier, and an unlisted one falls back to the model default, which is high. You measured that correctly.
  • low and high are different prompts, but a soft instruction separates weakly; n=40 on output length will not resolve 25 from 50, and the same goes for 75 vs 100.

So "two rungs" is what is observable, and "four tiers plus off" is what the API accepts. Which of those the entry should carry is your editorial call — I would rather give you the fact than push you toward the framing that flatters us.

The toggle: thinking is the name to document

The model reads only thinking. Sent straight to the serve, bypassing our gateway, one reasoning prompt:

sent reasoning_content
chat_template_kwargs.thinking: true 89 chars
chat_template_kwargs.enable_thinking: true 0 chars

Both names worked for you on v4.1-flash because our gateway rewrites the generic name onto the switch a template actually reads, and v4.1-flash is in that table. 0731 is not — which is exactly why it looked inert to you there. Same model family, same native name, the rewrite just never fired.

So please name thinking. It is what the checkpoint's own encoding.py reads, it is correct on both deepseek models today, and it stays correct for anyone running the model outside engy. enable_thinking happens to work on v4.1-flash through our translation layer, but it is not the model's interface and it does not work on 0731 — I would not put it in the entry.

The missing headers: ours, and not fixed yet

The threshold you spotted is exact. A non-streaming response slower than 15 s cannot be a plain JSON response — the handler would send nothing until the route settles and Cloudflare would 524 it — so we switch to a streaming response that drips keepalive whitespace until the body is ready. HTTP headers are committed before the first body byte, and at that moment the route has not returned, so ours were dropped on the floor.

Reproduced here: a 0.8 s request carried all three headers, a 72.3 s request carried none, and the body's x_engy block was intact in both. That matches your 114 of 480 — it is not sampling, every response over the threshold loses them.

We have not fixed it yet and I do not have a date for you. In the meantime the body's x_engy block (request_id, miner, worker, charged_micro) carries the same data, is unaffected by any of this, and is the read I would point tooling at.

reasoning_tokens: fixed, and live now

The field was being dropped between the model and the gateway — sglang reports it flat as usage.reasoning_tokens, and the miner process in between was not carrying it through, so there was nothing to put in completion_tokens_details. Both ends are corrected and the last workers finished rolling today.

This one is already live on the current gateway. Re-measured across every worker serving the model just now: completion_tokens_details.reasoning_tokens populated on every response that produced reasoning. Re-test whenever you like.

@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 19, 2026
@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 19, 2026
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Sep 19, 2026
@roykollensvendsen

Copy link
Copy Markdown
Author

Thanks — four answers, and three of them changed something.

Effort. Taken. Both DeepSeek entries now carry low|high|max instead of the high|max I had narrowed to. Your explanation is the reason: if the level is rendered into the prompt as a sentence rather than enforced as a budget, a length test at n=40–126 was never going to resolve one rung from the next, so the narrowing rested on evidence that could not have shown otherwise. Your docs' mapping table settles which levels are distinct, and the entries now follow it: sent to DeepSeek V4, low and medium both arrive as low, high as high, xhigh and max as max — three levels.

One small discrepancy while you have it open: you wrote that medium collapses onto high, the model default. The table on engy.ai/docs (read 2026-09-19) puts medium on low for DeepSeek V4. It does not change the entry either way, since medium is a duplicate of another level in both accounts, but one of the two is wrong and it is your page.

The toggle. Both DeepSeek files and kimi-k3 already name chat_template_kwargs.thinking — that part was settled before you wrote. What was missing was why enable_thinking worked on v4.1-flash and looked inert on 0731, and the rewrite table answers it. The body says so now, and says thinking is the checkpoint's own interface and the one that stays right off engy.

The headers. Recorded as a known quirk with your reproduction, and the body points anyone re-measuring at the x_engy block. Nothing in the entry depends on it.

reasoning_tokens. Noted, thank you. models.dev carries no field for it, so nothing changes here, but it matters for anyone measuring against engy.

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 20, 2026
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Oct 2, 2026
@roykollensvendsen

roykollensvendsen commented Oct 2, 2026 •

Copy link
Copy Markdown
Author

@crabbytt @andy-engy kimi-k3's limits have changed, and the entry now follows them (head 1fac351).

  • /v1/models reports context_length 1,113,088, up from 1,048,576.
  • The authenticated list gives the split as max_input_tokens 1,047,552 (983,040 when I last checked on 2026-08-30) and max_output_tokens 65,536 (unchanged). The two add up to the new context exactly.
  • The file now sets context = 1_113_088 explicitly, since it is above the lab's 1M window and can no longer be inherited.

I can't tell when between 08-30 and today this changed. An earlier version of this comment said it happened today, but I got that from the created field, which turns out to be the request time.

The other seven models match your pricing page and both model lists exactly, on price and on limits. If the larger prompt budget is not meant to stay, tell me and I'll put it back.

I have left MiniMax-H3 out of this PR. It is billed per second of video, and the catalogue only has per-token pricing.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added reviewer: ready Automated review found no actionable items and removed reviewer: ready Automated review found no actionable items labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Oct 2, 2026
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Oct 2, 2026
engy (api.engy.ai) relays open models over an OpenAI-compatible
endpoint; each of the eight entries is override-only against its lab
model. Prices, context and modalities come from the public
GET /v1/models, the input/output split from the authenticated
engy.ai/api/v1/models (context is their sum), and image input was
checked by sending an image to each model.

Reasoning controls were measured per model with paired prompts and a
positive control, since they differ per deployment and no endpoint
reports them: toggles on chat_template_kwargs (enable_thinking for
glm-5.2 and Qwen, thinking for DeepSeek and Kimi), effort levels from the host's
ladder and the same-surface peers, and no toggle on the GLM-5.3 pair, where nothing
switches reasoning off. kimi-k3 overrides temperature = true, which
engy honours and Moonshot's API does not. Each header carries its cap,
sample sizes and p-values.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions github-actions Bot removed the reviewer: ready Automated review found no actionable items label Oct 2, 2026
@roykollensvendsen roykollensvendsen changed the title feat(engy): add engy provider and sync module feat(engy): add engy provider Oct 2, 2026
@roykollensvendsen

Copy link
Copy Markdown
Author

@crabbytt I've split this PR. It now adds only the engy provider and its eight model files (head 939b47f, rebased on today's dev). The sync module will come as a separate PR once the provider is merged.

The reason is that it had been sitting clean for a month. Since late August, 25 of the 79 bot-approved provider-only PRs here have been merged, and 0 of the 9 that bundled a sync module. Sync modules for providers already in the catalogue have been getting merged within days as their own PRs, so that's the route I'm taking.

Nothing in your entry has changed. The model files are byte-identical to the version you endorsed, plus today's kimi-k3 limits (above). Until the sync PR lands, prices and context are updated by hand. If you change something in the meantime, a note here is enough and I'll pick it up.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants