feat(engy): add engy provider - #5910
Conversation
5cddde7 to
a356e74
Compare
|
No actionable findings. |
a356e74 to
5117faf
Compare
|
No actionable findings. |
5117faf to
8296a50
Compare
|
No actionable findings. |
|
Endorsing this — I'm the owner of engy ( I checked all seven files against what we serve. The prices, context/limit split, modalities, and reasoning controls are all correct — including the parts you can't verify externally: the authed input/output split ( The sync safeguards are exactly right for our unauthenticated list:
Great work — the probing and fixed-point tests are well above the bar. |
|
Thanks for checking it line by line, and for confirming the parts that cannot be verified from outside. That settles the two questions the draft carried: the Two things from the last few days that the entry now reflects, in case they are useful on your side: the sync module picked up glm-5.3's window going to 327,680 and kimi-k3's 30% raise on its own, which is what it is there for, and I re-checked all 21 prices, the seven limits and the modality lists against the live list this morning at build 0.2.19 with no drift. One measurement you may want regardless of this PR: on qwen3.6-35b-a3b, when |
8296a50 to
828aa15
Compare
|
No actionable findings. |
|
@crabbytt deepseek-v4.1-flash is on the entry now (head 828aa15, bot clean). Same sources as before: 0.04 / 0.08 / 0.008 from /v1/models, 262,144 in and 65,536 out from the authenticated list, text+image checked by sending an image. Two things only you can settle:
Two observations from the run, both on 0.2.20, in case they are useful: the |
828aa15 to
5ea0299
Compare
|
Answering from the engy side. Thanks — all four are answerable. Measurements are from 2026-09-15. Effort: your measurement is right, the conclusion needs one correctionThe ladder is four real tiers, not two —
So "two rungs" is what is observable, and "four tiers plus off" is what the API accepts. Which of those the entry should carry is your editorial call — I would rather give you the fact than push you toward the framing that flatters us. The toggle:
|
| sent | reasoning_content |
|---|---|
chat_template_kwargs.thinking: true |
89 chars |
chat_template_kwargs.enable_thinking: true |
0 chars |
Both names worked for you on v4.1-flash because our gateway rewrites the generic name onto the switch a template actually reads, and v4.1-flash is in that table. 0731 is not — which is exactly why it looked inert to you there. Same model family, same native name, the rewrite just never fired.
So please name thinking. It is what the checkpoint's own encoding.py reads, it is correct on both deepseek models today, and it stays correct for anyone running the model outside engy. enable_thinking happens to work on v4.1-flash through our translation layer, but it is not the model's interface and it does not work on 0731 — I would not put it in the entry.
The missing headers: ours, and not fixed yet
The threshold you spotted is exact. A non-streaming response slower than 15 s cannot be a plain JSON response — the handler would send nothing until the route settles and Cloudflare would 524 it — so we switch to a streaming response that drips keepalive whitespace until the body is ready. HTTP headers are committed before the first body byte, and at that moment the route has not returned, so ours were dropped on the floor.
Reproduced here: a 0.8 s request carried all three headers, a 72.3 s request carried none, and the body's x_engy block was intact in both. That matches your 114 of 480 — it is not sampling, every response over the threshold loses them.
We have not fixed it yet and I do not have a date for you. In the meantime the body's x_engy block (request_id, miner, worker, charged_micro) carries the same data, is unaffected by any of this, and is the read I would point tooling at.
reasoning_tokens: fixed, and live now
The field was being dropped between the model and the gateway — sglang reports it flat as usage.reasoning_tokens, and the miner process in between was not carrying it through, so there was nothing to put in completion_tokens_details. Both ends are corrected and the last workers finished rolling today.
This one is already live on the current gateway. Re-measured across every worker serving the model just now: completion_tokens_details.reasoning_tokens populated on every response that produced reasoning. Re-test whenever you like.
5ea0299 to
4a0fe36
Compare
|
No actionable findings. |
4a0fe36 to
f7dbb66
Compare
|
Thanks — four answers, and three of them changed something. Effort. Taken. Both DeepSeek entries now carry One small discrepancy while you have it open: you wrote that The toggle. Both DeepSeek files and kimi-k3 already name The headers. Recorded as a known quirk with your reproduction, and the body points anyone re-measuring at the
|
|
No actionable findings. |
|
@crabbytt @andy-engy kimi-k3's limits have changed, and the entry now follows them (head 1fac351).
I can't tell when between 08-30 and today this changed. An earlier version of this comment said it happened today, but I got that from the The other seven models match your pricing page and both model lists exactly, on price and on limits. If the larger prompt budget is not meant to stay, tell me and I'll put it back. I have left MiniMax-H3 out of this PR. It is billed per second of video, and the catalogue only has per-token pricing. |
|
No actionable findings. |
|
No actionable findings. |
a0f948d to
1fac351
Compare
|
No actionable findings. |
engy (api.engy.ai) relays open models over an OpenAI-compatible endpoint; each of the eight entries is override-only against its lab model. Prices, context and modalities come from the public GET /v1/models, the input/output split from the authenticated engy.ai/api/v1/models (context is their sum), and image input was checked by sending an image to each model. Reasoning controls were measured per model with paired prompts and a positive control, since they differ per deployment and no endpoint reports them: toggles on chat_template_kwargs (enable_thinking for glm-5.2 and Qwen, thinking for DeepSeek and Kimi), effort levels from the host's ladder and the same-surface peers, and no toggle on the GLM-5.3 pair, where nothing switches reasoning off. kimi-k3 overrides temperature = true, which engy honours and Moonshot's API does not. Each header carries its cap, sample sizes and p-values. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1fac351 to
939b47f
Compare
|
@crabbytt I've split this PR. It now adds only the engy provider and its eight model files (head 939b47f, rebased on today's The reason is that it had been sitting clean for a month. Since late August, 25 of the 79 bot-approved provider-only PRs here have been merged, and 0 of the 9 that bundled a sync module. Sync modules for providers already in the catalogue have been getting merged within days as their own PRs, so that's the route I'm taking. Nothing in your entry has changed. The model files are byte-identical to the version you endorsed, plus today's kimi-k3 limits (above). Until the sync PR lands, prices and context are updated by hand. If you change something in the meantime, a note here is enough and I'll pick it up. |
|
No actionable findings. |
Summary
engy (api.engy.ai) relays eight open models over one OpenAI-compatible chat-completions endpoint (
@ai-sdk/openai-compatible). Each model file is override-only against its lab entry. The sync module that was in this PR before will follow as its own PR once the provider is in.engy's maintainer has endorsed the entry: #5910 (comment). That covers the figures that can't be checked from outside, including the input/output split, the modality narrowing, and kimi-k3 honouring
temperature.Models
attachment = false. qwen3.8-27b, glm-5.3-flash, kimi-k3 and deepseek-v4.1-flash take text and image, and the input list is narrowed wherever the lab lists more.statusis unset.Reasoning
engy is a relay, so each model's baseline is its lab entry plus peers on the same kind of API. The toggle is
chat_template_kwargs.enable_thinkingfor glm-5.2 and Qwen, andchat_template_kwargs.thinkingfor DeepSeek and Kimi. Effort isreasoning_effort. There is no budget field, so nobudget_tokens.The effort sets follow engy's published per-template mapping (engy.ai/docs). engy renders the effort level into the prompt as a sentence instead of enforcing a budget. Each toggle was checked live: it suppresses reasoning on every sample. Each file's header gives the sample sizes and p-values. Two exceptions:
Sources
doclink. engy's docs list no models.reasoning_effortmapping andchat_template_kwargs.Test plan
bun validate. Each resolved engy model checked against both model lists on 2026-10-02.bun run testinpackages/sdkdevc81646eENGY_API_KEY, and the cost matches engy's bill