feat(providers): add iFlytek MaaS (China) provider - #7057
Open
dongjiang1989 wants to merge 1 commit into
Open
dongjiang1989 wants to merge 1 commit into
dongjiang1989 wants to merge 1 commit into
Conversation
Contributor
Action items
|
dongjiang1989
added a commit
to dongjiang1989/models.dev
that referenced
this pull request
Sep 15, 2026
Address review feedback on PR anomalyco#7057: Reasoning options aligned to lab/peer baselines per AGENTS.md relay rules: - DeepSeek V4 Pro: toggle + high/max (was: invented L/M/H) - DeepSeek V4 Flash: toggle + low/high/max (was: invented L/M/H) - DeepSeek V3.2: effort low/medium/high, drop unsupported 'none' - GLM-5.1 / GLM-5: toggle only (was: invented high/max) - GLM-5.2: keep high/max (iFlytek host docs confirm; only GLM model with graded effort) - GLM-4.7-Flash: [] always-on, no caller control (was: invented L/M/H) - Kimi K2.6 / K2.5: toggle only (was: invented L/M/H) - MiniMax M2.5: [] always-on, no caller control (was: invented L/M/H) - Qwen 397B / 3.6-35B / 3.5-35B: toggle + budget_tokens (was: invented L/M/H) - Spark-X2.5-4B / 1.7B: toggle + low/medium/high for OSS models Tool calling set to false on all models except DeepSeek V3.2 and GLM-4.7-Flash (the only two the host docs confirm support tools on this API). Context limits added where host serves a smaller window than lab: - GLM-4.7-Flash: 128K (was: inherited lab 200K) - Kimi K2.5: 128K (was: inherited lab 262K) - MiniMax M2.5: 128K (was: inherited lab 204.8K) - Qwen3.6-35B-A3B / Qwen3.5-35B-A3B: 128K (was: inherited lab 262K) All reasoning controls documented with wire-field references in leading comments. bun validate passes locally. Signed-off-by: dongjiang1989 <dongjiang1989@126.com>
dongjiang1989
force-pushed
the
add-iflytek-provider
branch
from
September 15, 2026 02:05
48684bb to
91d4d84
Compare
Author
|
Fix github-actions review |
dongjiang1989
force-pushed
the
add-iflytek-provider
branch
from
September 15, 2026 02:10
e6bb396 to
6b5118a
Compare
Contributor
Action items
|
dongjiang1989
force-pushed
the
add-iflytek-provider
branch
from
September 15, 2026 02:30
6b5118a to
d2e7281
Compare
Contributor
Action items
|
dongjiang1989
force-pushed
the
add-iflytek-provider
branch
from
September 15, 2026 02:43
d2e7281 to
2bf2bb0
Compare
Contributor
Action items
|
dongjiang1989
force-pushed
the
add-iflytek-provider
branch
from
September 15, 2026 07:41
2bf2bb0 to
87052ee
Compare
Contributor
Action items
|
Adds the iFlytek MaaS Inference API provider (讯飞星辰MaaS推理服务, https://maas.xfyun.cn) with 18 models — 4 first-party Spark and 14 third-party (DeepSeek, GLM, Kimi, MiniMax, Qwen). Provider: providers/iflytek/ - API: https://maas-api.cn-huabei-1.xf-yun.com/v2 (OpenAI-compatible, text-only) - env: IFLYTEK_API_KEY - npm: @ai-sdk/openai-compatible Lab models (models/iflytek/): - spark-x2.5: 293B-A30B MoE flagship, 256K context, tool_call=true - spark-x2-flash: lightweight fast model, 256K context, tool_call=true - spark-x2.5-4b: 4B edge Dense, 1M context, open-weight, tool_call=true - spark-x2.5-1.7b: 1.7B ultra-light edge, 1M context, open-weight Provider models (18 total) — override-only after base_model merge: - Spark (first-party): spark-x2.5, spark-x2-flash, spark-x2.5-4b, spark-x2.5-1.7b (tool_call overridden to false on all 4 — host lacks tool support) - DeepSeek: deepseek-v4-pro, deepseek-v4-flash, deepseek-v3.2 - GLM/Zhipu: glm-5.2, glm-5.1, glm-5, glm-4-7-flash - Kimi/Moonshot: kimi-k2.6, kimi-k2.5 - MiniMax: minimax-m2.5 - Qwen/Alibaba: qwen3.5-397b-a17b, qwen3.6-35b-a3b, qwen3.5-35b-a3b, qwen3-coder-next Reasoning controls (per iFlytek host docs, aligned with lab/peer baselines): - DeepSeek V4 Pro: toggle (thinking.type) + effort high/max - DeepSeek V4 Flash: toggle (thinking.type) + effort low/high/max - DeepSeek V3.2: toggle (thinking.type) only (matches OpenRouter/TokenGo/Novita) - GLM-5.2: effort high/max (only GLM with graded effort on this host) - GLM-5.1/GLM-5: toggle (enable_thinking) only - GLM-4.7-Flash: toggle (enable_thinking) (matches first-party zhipuai baseline) - Kimi K2.6/K2.5: toggle (enable_thinking) only - MiniMax M2.5: [] always-on, no caller control (matches minimax first-party, opencode, alibaba-token-plan-cn, cortecs, ollama-cloud peers) - Qwen 397B/3.6-35B/3.5-35B: toggle (enable_thinking) only — host 'OSS' phrase targets OpenAI OSS gpt-oss, not Qwen; no budget_tokens (host lacks thinking_budget); no graded effort - Qwen3-Coder-Next: non-reasoning - Spark-X2.5/Flash/4B/1.7B: toggle (enable_thinking) + high/max (same surface) All toggle models have leading '# Toggle: <field> = <values>' wire comments. No redundant restated lab values (attachment, modalities, tool_call only overridden when they differ from the lab entry). Context limit overrides (host < lab): - GLM-4.7-Flash: 128K (lab 200K) - Kimi K2.5: 128K (lab 262K) - MiniMax M2.5: 128K (lab 204.8K) - Qwen3.5-397B-A17B: 256K (lab 262K) - Qwen3.6-35B-A3B / Qwen3.5-35B-A3B: 128K (lab 262K) Multimodal override: attachment=false + text-only [modalities] on Kimi and Qwen (base labs are multimodal; iFlytek endpoint is text-only). Pricing: USD per million tokens from Token Plan points at 0.01 CNY/point (1 point ≈ 0.01 CNY per Token Plan doc: 200 CNY / 20000 points for standard members), converted at 7.25 CNY/USD. Spark-X2.5 uses list price (CNY 3.2/0.48/12); 50% promo noted separately. Spark-X2.5-4B/1.7B currently limited-time free (cost = 0). Schema: added 'spark' to ModelFamily enum in packages/core/src/family.ts. Data sources (accessed 2026-09-14): - https://www.xfyun.cn/doc/spark/推理服务-http.html - https://www.xfyun.cn/doc/spark/TokenPlan.html - https://maas.xfyun.cn/modelSquare Excluded: Token Plan API and Coding Plan API (separate billing surfaces), Spark-X2/Spark-X2-Agent (decommissioned). bun validate passes locally. Signed-off-by: dongjiang1989 <dongjiang1989@126.com>
dongjiang1989
force-pushed
the
add-iflytek-provider
branch
from
September 15, 2026 08:09
87052ee to
91b662c
Compare
Contributor
|
No actionable findings. |
Author
|
@rekram1-node PTAL, when you have time |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the iFlytek MaaS Inference API provider (讯飞星辰MaaS推理服务, https://maas.xfyun.cn) with 18 models — 4 first-party Spark and 14 third-party (DeepSeek, GLM, Kimi, MiniMax, Qwen).
providers/iflytek/provider.toml— OpenAI-compatible endpointhttps://maas-api.cn-huabei-1.xf-yun.com/v2, envIFLYTEK_API_KEYproviders/iflytek/logo.svg— squarecurrentColorglyphmodels/iflytek/spark-x2.5— 293B-A30B MoE flagship, 256K context, 200+ languages, code/agentic focus. Released 2026-09-07 (full inline definition; first-party lab)models/iflytek/spark-x2-flash— lightweight fast model, 256K contextmodels/iflytek/spark-x2.5-4b— 4B edge Dense, native 1M context, open-weight. Hybrid attention (1 full + 3 sliding window layers). Released 2026-09-01models/iflytek/spark-x2.5-1.7b— 1.7B ultra-light edge, native 1M context, open-weight. Released 2026-09-01deepseek-v4-pro/deepseek-v4-flash/deepseek-v3.2→base_model = "deepseek/..."glm-5.2/glm-5.1/glm-5/glm-4-7-flash→base_model = "zhipuai/..."kimi-k2.6/kimi-k2.5→base_model = "moonshotai/..."minimax-m2.5→base_model = "minimax/MiniMax-M2.5"qwen3.5-397b-a17b/qwen3.6-35b-a3b/qwen3.5-35b-a3b/qwen3-coder-next→base_model = "alibaba/..."Data sources (accessed 2026-09-14):
/v2/chat/completionsmessage.reasoning_content(interleaved). Spark-X2.5 / Spark-X2-Flash useenable_thinking(bool) + efforthigh/max; DeepSeek V3.2/V4 use effortnone|low|medium|high; GLM-5.x uses efforthigh|maxtoolson this hostSchema: added
"spark"to theModelFamilyenum inpackages/core/src/family.ts.Excluded: Token Plan API (
maas-token-api.../v2) and Coding Plan API (maas-coding-api.../v2) — separate subscription billing surfaces with distinct API keys and base URLs, could be added asiflytek-token-plan/iflytek-coding-planlater. Spark-X2 and Spark-X2-Agent are decommissioned per Token Plan docs (已下线).bun validatepasses locally (validated against currentdev).