Skip to content

feat(zerosignal): add ZeroSignal provider with 12 models - #7262

Open
drichar wants to merge 3 commits into
anomalyco:devfrom
drichar:feat/zerosignal-provider
Open

drichar wants to merge 3 commits into
anomalyco:devfrom
drichar:feat/zerosignal-provider

Conversation

@drichar

@drichar drichar commented Sep 16, 2026 •

Copy link
Copy Markdown

Adds ZeroSignal as an OpenAI-compatible provider. ZeroSignal is a pay-per-use inference network; models are served by independent operators and reached through a local proxy (zs-proxy) at http://127.0.0.1:9376/v1, the same shape as the atomic-chat and lmstudio entries. The proxy does not check the API key (admission comes from the user's wallet), so ZEROSIGNAL_API_KEY can be any non-empty value.

ZeroSignal is a multi-model relay, not a lab, so every entry uses base_model and is override-only. I'm on the ZeroSignal team.

Scope

12 models, all with existing lab metadata under models/. The network serves more (about 50 ids), but I've limited this PR to the ones where the network advertises an explicit effort set that lines up with the lab entry. The rest (ids with no advertised reasoning controls, ids from operators advertising a generic low/medium/high on every model, and a few operator fine-tunes with no lab entry) I'd rather verify on the wire before publishing controls for them. Separate PR if/when that's done.

Data provenance

Everything below was read from GET /v1/models on the local proxy on 2026-09-16 (the proxy aggregates the on-chain operator registry). It is a local endpoint, so the relevant entries are inlined under the fold rather than linked.

  • Cost is the lowest advertised operator rate for the id, USD per 1M tokens, inclusive of the protocol fee (what the caller pays). Operators price independently, so this is a floor; the routed operator may charge more. cache_read is included only where an operator actually discounts it below input; there is no cache-write premium on the network, so cache_write is omitted.

  • reasoning_options were measured on the host for all 12 ids rather than taken from the network's advertised allowed_efforts (that field is declared by each operator in node config). Method: same prompt, max_tokens 3000 to 4000 so nothing truncates, temperature 0 unless the model dictates one, comparing reasoning tokens from usage and the returned trace across every effort value, 2026-09-16 against the local proxy. Where the measured controls differ from the lab entry, the file's leading comment says what was observed.

    id none low medium high xhigh max authored trace field
    glm-5.3 rejected 8 rejected 26 rejected 316 low, high, max reasoning_content
    glm-5.3-flash rejected 21 rejected 24 rejected 236 low, high, max reasoning_content
    glm-5.2 643 (thinks) 454 454 (same output as low) 690 high, max reasoning_content
    kimi-k3 0 85 82 131 none, low, high, max reasoning_content
    grok-4.6 rejected 105 187 110 134 rejected low, medium, high, xhigh reasoning_content
    grok-4.5 rejected 74 103 104 101 rejected low, medium, high reasoning_content
    grok-4.3 0 229 276 216 none, low, medium, high reasoning_content
    google/gemini-3.7-flash rejected 0 147 378 282 399 low, medium, high reasoning
    google/gemini-3.8-flash rejected 0 419 536 418 479 low, medium, high reasoning
    openai/gpt-5.6-luna 0 31 33 53 none, low, medium, high, xhigh, max reasoning
    openai/gpt-5.6-terra 0 24 28 34 30 57 none, low, medium, high, xhigh, max varies
    openai/gpt-6-astra rejected rate-limited rate-limited rate-limited rate-limited rate-limited low, medium, high, xhigh, max not observed

    Cells are reasoning tokens; "rejected" is a host error for that value; blank means not probed. Notes: glm-5.2 returned byte-identical output for low and high and a full trace for none, so only high/max (Z.ai's own set) are authored and off is unavailable. kimi-k3 also enforces Moonshot's temperature rule through the proxy (1 for thinking levels, 0.6 for none). Grok 4.5/4.6 graded counts are within noise, so xAI's sets are kept. Gemini accepts xhigh/max but they do not exceed high, so Google's set is kept; low yields zero reasoning tokens. openai/gpt-6-astra could not be exercised beyond none because every graded call was rate-limited upstream on the day; its authored set is the operator's advertised one, identical to OpenAI's entry, and the comment says so.

  • [interleaved] is declared as field = "reasoning_content" on the seven ids that returned the trace in that field. The two Geminis and gpt-5.6-luna return it in a field named reasoning, which is not one of the schema's values; gpt-5.6-terra returned it in one measurement and not another; gpt-6-astra was not observed. Those five carry no [interleaved] and a comment explaining why.

  • [limit] overrides are only where the operator's advertised context_length or max_completion_tokens differs from the lab file. Where the network advertises no output cap, the lab value is inherited rather than guessed.

id base_model allowed_efforts in / out / cache_read (USD/1M) limit override
glm-5.3 zhipuai/glm-5.3 low, high, max 1.694 / 5.324 / 0.3146 output 32768
glm-5.3-flash zhipuai/glm-5.3-flash low, high, max 0.09075 / 0.3025 / 0.01815 none
glm-5.2 zhipuai/glm-5.2 high, max (measured; see below) 1.694 / 5.324 / 0.3146 output 32768
kimi-k3 moonshotai/kimi-k3 none, low, high, max (measured) 3.432 / 17.16 / 0.3432 context 1000000, output 32768
grok-4.6 xai/grok-4.6 low, medium, high, xhigh 2.42 / 7.26 / 0.363 output 32768
grok-4.5 xai/grok-4.5 low, medium, high 2.42 / 7.26 / 0.363 output 32768
grok-4.3 xai/grok-4.3 none, low, medium, high (measured) 1.5125 / 3.025 / 0.242 context 200000, output 32768
google/gemini-3.7-flash google/gemini-3.7-flash low, medium, high 1.1 / 5.5 / 0.11 none
google/gemini-3.8-flash google/gemini-3.8-flash low, medium, high 1.1 / 5.5 / 0.11 none
openai/gpt-5.6-luna openai/gpt-5.6-luna none, low, medium, high, xhigh, max (measured) 0.319 / 1.936 / 0.033 none
openai/gpt-5.6-terra openai/gpt-5.6-terra none, low, medium, high, xhigh, max (measured) 2.662 / 15.972 / 0.2662 none
openai/gpt-6-astra openai/gpt-6-astra low, medium, high, xhigh, max 12.1 / 60.5 / 1.21 none
Raw /v1/models entries for the 12 ids (2026-09-16)
[
  {
    "id": "glm-5.3",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "canonical_slug": "zai/glm-5.3",
    "context_length": 1000000,
    "pricing": {
      "prompt": "0.000001694",
      "completion": "0.000005324",
      "input_cache_read": "0.0000003146"
    },
    "top_provider": {
      "context_length": 1000000,
      "max_completion_tokens": 32768
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "high",
        "max"
      ]
    },
    "tool_use": true
  },
  {
    "id": "glm-5.3-flash",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "canonical_slug": "zai/glm-5.3-flash",
    "context_length": 1000000,
    "architecture": {
      "input_modalities": [
        "text",
        "image",
        "video"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.00000009075",
      "completion": "0.0000003025",
      "input_cache_read": "0.00000001815"
    },
    "top_provider": {
      "context_length": 1000000,
      "max_completion_tokens": 131072
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "high",
        "max"
      ]
    },
    "tool_use": true
  },
  {
    "id": "glm-5.2",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "hugging_face_id": "zai-org/GLM-5.2",
    "context_length": 1000000,
    "pricing": {
      "prompt": "0.000001694",
      "completion": "0.000005324",
      "input_cache_read": "0.0000003146"
    },
    "top_provider": {
      "context_length": 1000000,
      "max_completion_tokens": 32768
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "none",
        "minimal",
        "low",
        "medium",
        "high",
        "xhigh",
        "max"
      ]
    },
    "tool_use": true
  },
  {
    "id": "kimi-k3",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "hugging_face_id": "moonshotai/Kimi-K3",
    "context_length": 1000000,
    "architecture": {
      "input_modalities": [
        "text",
        "image",
        "video"
      ]
    },
    "pricing": {
      "prompt": "0.000003432",
      "completion": "0.00001716",
      "input_cache_read": "0.0000003432"
    },
    "top_provider": {
      "context_length": 1000000,
      "max_completion_tokens": 32768
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "high",
        "max"
      ]
    },
    "tool_use": true
  },
  {
    "id": "grok-4.6",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "canonical_slug": "xai/grok-4.6",
    "context_length": 500000,
    "architecture": {
      "input_modalities": [
        "text",
        "image"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.00000242",
      "completion": "0.00000726",
      "input_cache_read": "0.000000363"
    },
    "top_provider": {
      "context_length": 500000,
      "max_completion_tokens": 32768
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high",
        "xhigh"
      ]
    },
    "tool_use": true
  },
  {
    "id": "grok-4.5",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "canonical_slug": "xai/grok-4.5",
    "context_length": 500000,
    "architecture": {
      "input_modalities": [
        "text",
        "image"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.00000242",
      "completion": "0.00000726",
      "input_cache_read": "0.000000363"
    },
    "top_provider": {
      "context_length": 500000,
      "max_completion_tokens": 32768
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high"
      ]
    },
    "tool_use": true
  },
  {
    "id": "grok-4.3",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "canonical_slug": "xai/grok-4.3",
    "context_length": 200000,
    "architecture": {
      "input_modalities": [
        "text",
        "image"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.0000015125",
      "completion": "0.000003025",
      "input_cache_read": "0.000000242"
    },
    "top_provider": {
      "context_length": 200000,
      "max_completion_tokens": 32768
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high"
      ]
    },
    "tool_use": true
  },
  {
    "id": "google/gemini-3.7-flash",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "canonical_slug": "google/gemini-3.7-flash",
    "context_length": 1048576,
    "architecture": {
      "input_modalities": [
        "text",
        "image",
        "video",
        "audio"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.0000011",
      "completion": "0.0000055",
      "input_cache_read": "0.00000011"
    },
    "top_provider": {
      "context_length": 1048576,
      "max_completion_tokens": 65536
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high"
      ]
    },
    "tool_use": true
  },
  {
    "id": "google/gemini-3.8-flash",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "canonical_slug": "google/gemini-3.8-flash",
    "context_length": 1048576,
    "architecture": {
      "input_modalities": [
        "text",
        "image",
        "video",
        "audio"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.0000011",
      "completion": "0.0000055",
      "input_cache_read": "0.00000011"
    },
    "top_provider": {
      "context_length": 1048576,
      "max_completion_tokens": 65536
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high"
      ]
    },
    "tool_use": true
  },
  {
    "id": "openai/gpt-5.6-luna",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "hugging_face_id": "openai/gpt-5.6-luna",
    "context_length": 1050000,
    "architecture": {
      "input_modalities": [
        "text",
        "image",
        "file"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.000000319",
      "completion": "0.000001936",
      "input_cache_read": "0.000000033"
    },
    "top_provider": {
      "context_length": 1050000,
      "max_completion_tokens": 128000
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high",
        "xhigh",
        "max"
      ]
    },
    "tool_use": true
  },
  {
    "id": "openai/gpt-5.6-terra",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "hugging_face_id": "openai/gpt-5.6-terra",
    "context_length": 1050000,
    "architecture": {
      "input_modalities": [
        "text",
        "image",
        "file"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.000002662",
      "completion": "0.000015972",
      "input_cache_read": "0.0000002662"
    },
    "top_provider": {
      "context_length": 1050000,
      "max_completion_tokens": 128000
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high",
        "xhigh",
        "max"
      ]
    },
    "tool_use": true
  },
  {
    "id": "openai/gpt-6-astra",
    "object": "model",
    "created": 0,
    "owned_by": "zerosignal",
    "hugging_face_id": "openai/gpt-6-astra",
    "context_length": 1050000,
    "architecture": {
      "input_modalities": [
        "text",
        "image",
        "file"
      ],
      "output_modalities": [
        "text"
      ]
    },
    "pricing": {
      "prompt": "0.0000121",
      "completion": "0.0000605",
      "input_cache_read": "0.00000121"
    },
    "top_provider": {
      "context_length": 1050000,
      "max_completion_tokens": 128000
    },
    "reasoning": {
      "supported": true,
      "allowed_efforts": [
        "low",
        "medium",
        "high",
        "xhigh",
        "max"
      ]
    },
    "tool_use": true
  }
]

Verification

  • bun validate passes locally.
  • Built packages/web and ran OPENCODE_MODELS_PATH="dist/_api.json" opencode; ZeroSignal appears in /connect and the 12 models in /models, and a completion against the running proxy streams.
  • logo.svg is a single-color currentColor mark with a square viewBox and no fixed size.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/zerosignal/models/glm-5.2.toml:9 - Check: Reasoning options must follow lab/same-surface baseline; do not dump the full effort enum or invent levels without distinct host effect. Why: Lab providers/zhipuai/models/glm-5.2.toml catalogs effective high/max and documents that none/minimal skip thinking while low/medium→high and xhigh→max. The ZeroSignal entry lists the entire schema enum (none…max), which mismatches the lab baseline and the “never dump the full enum” rule even if allowed_efforts is advertised. Callers would see fake distinct levels. Action: Replace with the host’s real distinct set (at least align to lab high/max, and only add none/others if this proxy treats them as meaningful, non-aliased controls); cite the wire field and observed behavior.
  • [medium] [possible mistake] providers/zerosignal/models/openai/gpt-5.6-luna.toml:9 - Check: Relay GPT-5.6 effort should match lab/same-surface peers unless this host truly omits a level. Why: Lab and peers (providers/openai, OpenRouter, neon) use none/low/medium/high/xhigh/max; off is effort=none. ZeroSignal omits none on both Luna and Terra, so the catalog implies reasoning cannot be disabled despite the lab surface. Action: Verify GET /v1/models / live requests for gpt-5.6-luna and gpt-5.6-terra; add none if accepted, or document in the leading comment that this host has no off control and keep the narrower list only if that is confirmed.
  • [medium] [possible mistake] providers/zerosignal/models/openai/gpt-5.6-terra.toml:9 - Check: Same GPT-5.6 none baseline as Luna. Why: Same mismatch as Luna: lab/peers include none for disable; this file does not. Action: Same verification/fix as Luna so both GPT-5.6 entries stay consistent with the host’s real effort set.
  • [medium] [possible mistake] providers/zerosignal/models/kimi-k3.toml:9 - Check: Kimi K3 baseline is lab toggle + low/high/max when the host exposes on/off. Why: First-party and many openai-compat peers use toggle (thinking.type / equivalent) plus low/high/max. This entry is effort-only with no none and no toggle, so the catalog provides no way to turn thinking off even though the lab model is toggleable. Action: Confirm whether zs-proxy forwards a thinking on/off field; if yes, add { type = "toggle" } plus a leading wire comment; if off is only via effort, add the real off value (e.g. none); if thinking is always on here, state that in the leading comment and keep effort-only only after that confirmation.
  • [low] [possible mistake] providers/zerosignal/models/grok-4.3.toml:9 - Check: Grok 4.3 lab/peer effort includes none. Why: Lab xAI and peers (OpenRouter, zenmux) use none/low/medium/high; ZeroSignal drops none while still advertising graded effort. That may be intentional host narrowing, but it is a concrete divergence for disable. Action: Confirm whether this id accepts none (or another off control); add it if supported, otherwise note in the leading comment that off is unavailable on this host.

@drichar

drichar commented Sep 16, 2026

Copy link
Copy Markdown
Author

All five findings were fair. The allowed_efforts field I authored from is declared by each operator in their node config, not measured, so I measured the five ids on the host directly and updated the files. Same prompt for every call, temperature 0 (except where the model dictates it), max_tokens 3000 to 4000 so nothing truncated, comparing reasoning tokens from usage and reasoning_content length across levels. 2026-09-16 against the local proxy.

id none low medium high xhigh max now authored
glm-5.2 643 454 454 690 high, max
kimi-k3 0 85 82 131 none, low, high, max
grok-4.3 0 229 276 216 none, low, medium, high
openai/gpt-5.6-luna 0 31 53 none, low, medium, high, xhigh, max
openai/gpt-5.6-terra 0 24 28 34 30 57 none, low, medium, high, xhigh, max

Cells are reasoning tokens. Per item:

  • glm-5.2 — low and high returned byte-identical output at temperature 0, so they alias; max is distinct; none still produced a full trace, so off is not available on this host. Now high/max, matching the Z.ai entry, with the aliasing and the missing off noted in the leading comment.
  • kimi-k3 — none returns no trace, so off is a real effort value rather than a separate field; per AGENTS.md that means effort-with-none and no toggle. The host also enforces Moonshot's temperature rule through the proxy (thinking levels require temperature 1, none requires 0.6), noted in the comment.
  • grok-4.3 — none is off; added. Graded levels all think; counts are within noise of each other, so I kept the lab's low/medium/high rather than claim finer distinctions.
  • gpt-5.6-luna / gpt-5.6-terra — none is off and every graded level is accepted; both now carry the full lab set.

The other seven ids advertise sets that already match their lab entries and are unchanged. PR body updated with the same table.

@github-actions

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [violation] providers/zerosignal/models/glm-5.2.toml:9 - Check: Provider-authored interleaved when this host exposes a reasoning side channel. Why: Measurement notes for this proxy document reasoning_content (and “reasoning trace” on sibling ids). Peers that return that field set [interleaved] field = "reasoning_content" so clients can parse the side channel; none of the ZeroSignal model files do. Action: Add [interleaved] field = "reasoning_content" on every ZeroSignal reasoner that returns it (at minimum the measured ids; the rest if they share the same proxy response shape).
  • [medium] [possible mistake] providers/zerosignal/models/glm-5.3.toml:12 - Check: Relay reasoning_options must reflect this host’s real controls, not only advertised enums. Why: Patch 2 remeasured five ids and every one diverged from the network’s advertised allowed_efforts (e.g. glm-5.2 collapsed to high/max; grok-4.3/kimi/gpt-5.6 gained none). The other seven still ship advertisement-only sets with no host measurement. Action: Probe glm-5.3, glm-5.3-flash, gemini-3.7-flash, gemini-3.8-flash, grok-4.5, grok-4.6, and gpt-6-astra the same way and update reasoning_options (or document that each set was verified on-wire).

@drichar

drichar commented Sep 16, 2026

Copy link
Copy Markdown
Author

Both items addressed.

Remaining seven ids measured. Same method as before (same prompt, max_tokens 3000, every effort value, reasoning tokens from usage), 2026-09-16 on the local proxy. Full 12-row table is now in the PR body; the short version:

  • glm-5.3, glm-5.3-flash: low/high/max accepted with clearly graded traces (8 → 26 → 316; 21 → 24 → 236). none, medium and xhigh are rejected by the host ("always engages in thinking and cannot be disabled"). Advertised set confirmed; off unavailable.
  • grok-4.5, grok-4.6: none and max rejected; graded levels accepted, counts within noise of each other, so xAI's own sets are kept (low/medium/high and low/medium/high/xhigh). Off unavailable.
  • google/gemini-3.7-flash, google/gemini-3.8-flash: none rejected ("Reasoning is mandatory for this endpoint"); low yields zero reasoning tokens, medium/high graded, xhigh/max accepted but do not exceed high. Google's low/medium/high kept. Off unavailable.
  • openai/gpt-6-astra: none rejected (same "mandatory" error). Every graded call was rate-limited upstream, on two attempts a few hours apart, so I could not exercise the levels. The authored set is the operator's advertised one, identical to providers/openai/models/gpt-6-astra.toml, and the leading comment states that the graded levels were not observed. Happy to drop this id to a follow-up PR if you would rather not carry it on that basis.

[interleaved]. Added field = "reasoning_content" on the seven ids that return the trace in that field: both GLM-5.3s, glm-5.2, kimi-k3, and the three Groks. Not added on the other five, with a comment in each file: the two Geminis and gpt-5.6-luna return the trace in a field named reasoning, which is not one of the schema's two values; gpt-5.6-terra returned it in one measurement and omitted it in another while usage still reported reasoning tokens, so the field is not stable across the operators serving that id; gpt-6-astra was not observed.

@github-actions

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant