This page specifies the request fields and response fields that control how long
a reply runs, what happens when the context fills, which thinking levels a model
offers, which sampling defaults it uses, and how CPU fallback is reported or
refused. The HTTP server and the in-process C API (lse_request,
lse_model_info) share the same JSON, so every shape below applies to both. Fields that are not part of the OpenAI schema use LSE's own
names, and OpenAI clients ignore them.
LSE has no output limit of its own. A reply ends at the first of these:
| End | finish_reason |
stop_reason |
|---|---|---|
| The model emits an end-of-turn token | stop (tool_calls when it made a call) |
stop_token |
One of the request's stop strings |
stop |
stop_sequence |
| A token limit (see the order below) | length |
max_tokens |
| The prompt plus the reply fill the context | length |
context_full |
| The client disconnected or cancelled | — | cancelled |
A request's token limit comes from the first of these that is set:
- The request's
max_tokens, ormax_completion_tokens(if both are sent,max_tokensis used). Each must be a positive integer;nullmeans not sent. - The operator's cap,
lse-server --max-tokens N(lse_config.max_tokens). The default is 0, which means no cap. With a cap, a request asking for more than the cap gets HTTP 400 (param: "max_tokens"). max_new_tokensfrom the model'sgeneration_config.json, when that file sets it.- No limit.
If the model's generation_config.json sets max_length, it bounds the prompt
plus the reply for items 2 to 4. Whatever the limit, the prompt plus the reply
never exceed the context (context_length, the engine's --kv-len).
LSE looked for model-defined output limits in generation_config.json
(max_new_tokens, max_length), in config.json (top-level fields and an
embedded generation_config object) and in the chat template. The Qwen3.6 and
Qwen3.8 checkpoints set none of them. tokenizer_config.json's
model_max_length is the trained context length, not an output limit, so it is
not used as one.
When a reply fills the context, it ends normally with HTTP 200. The text
generated so far is kept. The choice reports finish_reason: "length" and
stop_reason: "context_full".
Every completion response carries lse_context, so a client can offer to compact
the conversation before or after the context fills. For a streamed reply, it is
in the final chunk, the one with finish_reason.
{
"object": "chat.completion",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "..."},
"finish_reason": "length",
"stop_reason": "context_full",
"logprobs": null
}],
"usage": {"prompt_tokens": 40, "completion_tokens": 472, "total_tokens": 512,
"prompt_tokens_details": {"cached_tokens": 0}},
"lse_context": {"tokens_used": 512, "context_length": 512, "tokens_remaining": 0},
"timings": {"...": "..."}
}| Field | Type | Meaning |
|---|---|---|
choices[0].stop_reason |
string | stop_token, stop_sequence, max_tokens, context_full or cancelled |
lse_context.tokens_used |
integer | Prompt tokens (whole conversation, including cached ones) plus reply tokens |
lse_context.context_length |
integer | The most tokens the engine holds; the same as /v1/models context_length |
lse_context.tokens_remaining |
integer | context_length - tokens_used, never below 0 |
The final chunk of a stream has the same fields:
{"object": "chat.completion.chunk",
"choices": [{"index": 0, "delta": {}, "finish_reason": "length", "stop_reason": "context_full"}],
"lse_context": {"tokens_used": 512, "context_length": 512, "tokens_remaining": 0},
"timings": {"...": "..."}}/v1/completions uses the same stop_reason and lse_context fields. Its
choices carry text instead of message.
A request whose prompt leaves no room for a reply (prompt_tokens >= context_length)
gets HTTP 400 before any work starts. The error code is the same, context_full:
{
"error": {
"message": "the prompt is 600 tokens and the context holds 512; compact or shorten the conversation",
"type": "invalid_request_error",
"code": "context_full",
"param": "messages",
"lse_context": {"tokens_used": 600, "context_length": 512, "tokens_remaining": 0}
}
}param is messages on /v1/chat/completions and prompt on /v1/completions.
Here, tokens_used is the size of the prompt. Over lse_request, this arrives as
LSE_EVENT_ERROR with status 400 and the same body.
A client can read error.code == "context_full", or stop_reason == "context_full",
and does not need to match message strings. LSE does no compaction itself.
The thinking levels come from the model's own chat template. LSE reads the
template from chat_template.jinja beside config.json. If that file is absent,
it uses tokenizer_config.json chat_template (a string, or the entry named
default in a list), then chat_template.json. LSE renders the template with
probe conversations, using its own Jinja interpreter, and compares the results:
- Toggle: the template variable
enable_thinking, when setting it changes the rendering. - Effort levels: the values of the template variable
reasoning_effortthat the template compares against and that render withoutraise_exception. - Instruction: the text the template inserts before the system prompt for that level.
- Generation prompt: the assistant opening the template appends for that level. LSE uses it to open the reply.
- Default: the level whose rendering matches the template's rendering with neither variable set.
No model has hard-coded strings. A model starts without loading if its template
can't be read: lse_open fails and names the file and the construct.
Each model entry has a thinking object. The native API has the same object in
lse_model_info and GET /v1/lse/model_info (and under served there).
"thinking": {
"supported": true,
"source": "/models/qwen38-27b-q4/chat_template.jinja",
"toggle": "enable_thinking",
"effort": "reasoning_effort",
"default_level": "xhigh",
"levels": [
{"id": "none", "enable_thinking": false, "reasoning_effort": null,
"instruction": null, "opens_reasoning": false, "default": false},
{"id": "xhigh", "enable_thinking": true, "reasoning_effort": "xhigh",
"instruction": "Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.",
"opens_reasoning": true, "default": true},
{"id": "medium", "enable_thinking": true, "reasoning_effort": "medium",
"instruction": null, "opens_reasoning": true, "default": false},
{"id": "low", "enable_thinking": true, "reasoning_effort": "low",
"instruction": "Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.",
"opens_reasoning": true, "default": false}
]
}| Field | Type | Meaning |
|---|---|---|
supported |
boolean | The template defines at least one level |
source |
string or null | The file the template came from; null when the model has none |
toggle |
string or null | "enable_thinking" when the template has the switch |
effort |
string or null | "reasoning_effort" when the template has effort levels |
default_level |
string or null | The level a request that names none gets |
levels[].id |
string | The value to send as reasoning_effort |
levels[].enable_thinking |
boolean or null | What the level passes to the template's toggle; null without one |
levels[].reasoning_effort |
string or null | What it passes to the effort variable; null when none |
levels[].instruction |
string or null | Template text placed first in the system prompt; null when the template adds none |
levels[].opens_reasoning |
boolean | The reply starts inside a <think> block, so its text up to </think> is returned as reasoning_content |
levels[].default |
boolean | This is default_level |
Level ids, in the order they're listed:
nonemeans thinking is off (enable_thinking: false). It is present only when the template has the toggle.- Each effort value the template accepts, in the order the template names them, with thinking on.
onmeans thinking is on, for a template that has the toggle but no effort levels.
| Model template | levels |
default_level |
|---|---|---|
Qwen3.8 (chat_template.jinja) |
none, xhigh, medium, low |
xhigh |
Qwen3.6 (tokenizer_config.json) |
none, on |
on |
| SmolLM2, or a model with no template | (empty), supported: false |
null |
Qwen3.8's template has no high level. Its medium level adds no instruction,
and its xhigh and low levels each add their own.
| Field | Values |
|---|---|
reasoning_effort |
A levels[].id. Also accepted as thinking_level, chat_template_kwargs.reasoning_effort or reasoning.effort |
enable_thinking |
boolean. Also accepted as thinking (a boolean, or {"type": "enabled" | "disabled"}) or chat_template_kwargs.enable_thinking |
Resolution:
- If
reasoning_effortis sent, it must be a level id. The switch, if also sent, must agree with that level'senable_thinking. - If only
enable_thinking: falseis sent, the level isnone. If onlyenable_thinking: trueis sent, the level is the default when the default thinks, otherwise the first level that thinks. - If neither is sent, the level is
default_level. - On a model without thinking controls,
reasoning_effort: "none"andenable_thinking: falseare accepted, since the model already doesn't think. Asking such a model to think is an error.
Errors are HTTP 400, type: "invalid_request_error", and include levels, the
model's level ids:
error.code |
When |
|---|---|
unsupported_reasoning_effort |
The level is not one the template defines, e.g. high or minimal on Qwen3.8. No level is aliased to another |
thinking_unsupported |
Thinking was asked of a model whose template has no controls, or enable_thinking: false was sent to a template that can't switch thinking off |
conflicting_thinking |
Two spellings of the switch or the level disagree, or the level contradicts the switch |
invalid_thinking |
A field has the wrong type |
{"error": {"message": "reasoning_effort 'high' is not defined by this model's chat template; levels: none, xhigh, medium, low",
"type": "invalid_request_error", "code": "unsupported_reasoning_effort",
"param": "reasoning_effort", "levels": ["none", "xhigh", "medium", "low"]}}On /v1/completions the prompt is passed through unchanged. There, the thinking
fields only decide whether <think> markers in the output are split into
reasoning. They are checked for type and agreement, but not against the
template.
The model's own files supply the defaults for any field a request leaves out
(or sends as null). For each field, the first of these that sets it wins:
- The request.
generation_config.jsonbesideconfig.json.config.json's embeddedgeneration_configobject, then its top-level fields.- LSE's neutral defaults.
lse-server --temperature overrides the temperature default and reports
server_option as its source. do_sample: false in the model's files makes the
default temperature 0 (greedy). There's no per-model or per-family table.
Each model entry has a generation_defaults object. The native API has the same
object in lse_model_info and GET /v1/lse/model_info.
"generation_defaults": {
"temperature": 1.0, "top_k": 20, "top_p": 0.95, "min_p": 0.0,
"repetition_penalty": 1.0, "presence_penalty": 0.0,
"max_new_tokens": null, "max_length": null,
"sources": {
"temperature": "generation_config.json", "top_k": "generation_config.json",
"top_p": "generation_config.json", "min_p": "lse_default",
"repetition_penalty": "lse_default", "presence_penalty": "lse_default",
"max_new_tokens": "lse_default", "max_length": "lse_default"
}
}sources values are generation_config.json, config.json, server_option or
lse_default. The example shows Qwen3.8's generation_config.json
(temperature 1.0, top_k 20, top_p 0.95).
These defaults apply to fields the model's files don't set, and LSE reports them
as lse_default. They mean plain sampling from the model's distribution:
| Field | Default | Meaning of the default |
|---|---|---|
temperature |
1.0 | The model's own distribution (0 or less is greedy) |
top_k |
0 | Off |
top_p |
1.0 | Off |
min_p |
0.0 | Off |
repetition_penalty |
1.0 | Off |
presence_penalty |
0.0 | Off |
max_new_tokens, max_length |
null | No limit |
Each of these fields is accepted on /v1/chat/completions and /v1/completions,
and so through lse_request too:
| Field | Type and range |
|---|---|
temperature |
number; 0 or less is greedy |
top_k |
integer; 0 or -1 turns it off |
top_p |
number in [0, 1]; 1 turns it off |
min_p |
number in [0, 1]; drops tokens below min_p times the top token's probability |
presence_penalty |
number; subtracted once from the logit of each distinct token seen in the last 64 tokens |
repetition_penalty |
number > 0; divides a seen token's logit (or multiplies a negative one) |
frequency_penalty |
number; mapped to repetition_penalty = 1 + frequency_penalty when positive (an explicit repetition_penalty wins) |
seed |
nonnegative integer |
Sampling applies the penalties first, then temperature, then top-k, top-p and
min-p, in that order. The lse command line takes --top-k, --top-p,
--min-p, --presence-penalty, --repeat-penalty and --temperature, each
defaulting to the model's value.
An operation with no device kernel runs on the CPU interpreter. This is the only fallback LSE has, and it is reported every time:
- The server log prints
lse: CPU FALLBACK (N so far for this cause): <cause>the first time a cause occurs and again at each power of two. - The response of a request that caused one carries
lse_warnings. For a streamed reply, it is in the final chunk. lse_statuscounts every cause underengine.cpu_fallback, withallowedtelling whether fallback is permitted.
"lse_warnings": [
{"type": "cpu_fallback", "cause": "fused group 12 (3 nodes: mul, silu, add); ... no kernel for silu", "count": 4}
]"cpu_fallback": {"allowed": true, "total": 4,
"events": [{"cause": "fused group 12 (3 nodes: mul, silu, add); ...", "count": 4}]}lse-server --no-cpu-fallback, lse --no-cpu-fallback and
lse_config.disable_cpu_fallback = 1 refuse the CPU interpreter instead:
| Case | Without the option | With the option |
|---|---|---|
No --pool, and no device backend comes up |
The CPU backend runs the model, with a startup warning naming each backend that declined | Startup fails: no device backend came up and CPU fallback is disabled (--no-cpu-fallback), followed by each backend and its reason |
--pool names a CPU device (cpu:0) |
The CPU device is a pool member | Startup fails: device cpu:0 is refused: it has no device code generator ... |
--pool hrx:0 and that device does not come up |
Startup fails, naming the device | The same |
| An operation has no device kernel | It runs on the CPU and is reported as above | The request fails with an error that starts with CPU fallback disabled: and names the operation group |
A failed request gets HTTP 500 with type: "server_error". If the reply is
streamed, the error is sent in the stream. The engine stays loaded, and the next
request runs. lse_open returns NULL when startup fails, and lse exits with
status 1.
LSE_REQUIRE_DEVICE_KERNELS=1 in the environment has the same effect as the
option. It is the older qualification switch, kept for test runs. Leaving out the
option does not override it.