Skip to content

Latest commit

 

History

History
2978 lines (2612 loc) · 146 KB

File metadata and controls

2978 lines (2612 loc) · 146 KB

Workflow Guide

How the catalog is organised

workflows/ holds two trees, and which one a file is in says what it is for.

workflows/templates/ teaches a pattern. One file per capability - a shape (image to video, a multi-shot cut sequence), a mechanism (shared components, sub-workflows, pipeline_reference, typed references), or a reference convention (prompt:, previous_result:). These are what to read and copy. Where several checkpoints run the same pattern through the same pipeline class, one template carries them all and its description spells out the per-checkpoint argument sets, so the variations travel with the file rather than in a document that drifts from it. The templates/ltx2/ and templates/minimax/ subfolders each hold a family whose members build on one baseline.

workflows/models/ records a hardware fact: the quantization, offloading and component placement that make one checkpoint fit a real card. That is knowledge you cannot re-derive from a template, so it is kept runnable - but nobody learns a pattern from the fifth one, so these stay out of the way. Each carries a configures naming the template it is an instance of:

{
    "id": "flux-dev",
    "description": "Text-to-image with FLUX.1 dev - the reference FLUX workflow.",
    "configures": "templates/text-to-image",
    "steps": [ ... ]
}

The distinction exists because a catalog entry that cannot say which of the two it is leaves every reader - and every agent - to guess from the filename.

Structure

Every workflow is a JSON file with an id, optional variables, and a list of steps:

{
    "id": "my_workflow",
    "variables": {
        "prompt": "default prompt text",
        "steps": 25
    },
    "steps": [ ... ]
}

Variables define defaults that can be overridden from the command line:

python -m dw.run my_workflow.json prompt="a cat" steps=50

Variable names must be alphanumeric with underscores or hyphens.

Step Types

Each step has a name and exactly one of four types:

Pipeline Steps

Run a HuggingFace Diffusers model:

{
    "name": "generate",
    "pipeline": {
        "configuration": { "component_type": "FluxPipeline" },
        "from_pretrained_arguments": {
            "model_name": "black-forest-labs/FLUX.1-dev",
            "torch_dtype": "torch.bfloat16"
        },
        "arguments": {
            "prompt": "variable:prompt",
            "num_inference_steps": 25
        }
    },
    "result": { "content_type": "image/jpeg" }
}

A component_type can also name a community pipeline that dw ships in dw/community_pipelines/, by its dotted path; list_pipelines lists them after the diffusers ones. RFInversionFluxPipeline (dw.community_pipelines.pipeline_flux_rf_inversion.RFInversionFluxPipeline) inverts an image into FLUX latents (below). LTX2RefinePipeline (dw.community_pipelines.pipeline_ltx2_refine.LTX2RefinePipeline) is LTX2Pipeline plus a video argument: the clip is VAE-encoded at width x height and renoised at noise_scale over sigmas, a refine at the clip's own size rather than refine-clip's 2x. num_frames (8 * n + 1, no longer than the clip) defaults to the clip's length; video and latents are exclusive.

Pipeline Reference Steps

Re-run an already-loaded pipeline from an earlier step with a fresh set of arguments, instead of loading the model again. This is how a two-pass technique like RF-Inversion works: an invert step loads the pipeline, and a main step reuses it with the inverted latents:

{
    "name": "main",
    "pipeline_reference": {
        "reference_name": "invert",
        "arguments": {
            "prompt": "variable:prompt",
            "inverted_latents": "previous_result:invert.inverted_latents",
            "image_latents": "previous_result:invert.image_latents"
        }
    },
    "result": { "content_type": "image/jpeg" }
}

reference_name must name a step earlier in the same workflow that has a pipeline. Two steps whose pipelines differ only in a LoRA's scale or alpha or a scheduler's shift share one warm model, and a reference always runs at the referenced step's own values; release_pipeline on either step frees the shared model, as it does for two identical steps. See workflows/templates/community-pipeline.json for a full example.

Task Steps

Run utility operations (image processing, QR codes, data gathering):

{
    "name": "preprocess",
    "task": {
        "command": "canny",
        "arguments": {
            "image": { "location": "https://example.com/photo.jpg" }
        }
    },
    "result": { "content_type": "image/jpeg" }
}

A task can take inputs (a plain array) instead of arguments. Each array item becomes its own iteration, the same way multiple previous_result values do. An item that is a previous_result: reference becomes one iteration per result it names, and an object item expands the way an arguments object would:

{
    "name": "prompts",
    "task": {
        "command": "gather_inputs",
        "inputs": ["a marmot on a bicycle", "a bug driving a cycle"]
    }
}

Workflow Steps

Invoke another workflow file:

{
    "name": "expand",
    "workflow": {
        "path": "builtin:h3_context_ir.json",
        "arguments": { "prompt": "variable:prompt" }
    },
    "result": { "content_type": "text/plain" }
}

path is read the way run_workflow's workflow_path is: a catalog name as list_workflows reports it (templates/minimax/reference-to-video), with or without .json; a path relative to the file that names it (../models/x.json); or builtin:name.json for the packaged fragments in dw/workflows/. A name resolves beside the referencing file first, then against the run's own workflows/ directory, then against each read-only source the server lists - so a stored template can be composed without copying it into the workspace. A path that lands outside every source is refused, and one that resolves nowhere is a validation error rather than a run that fails on its first step.

When the composing step declares a result, that is where the composed output is written, once: the child's own last step does not save it a second time under its own name. A composing step that declares no result (or one with no content_type) leaves the saving to the child, as before. The child's other steps write into the same run directory, with the composing step's name leading their file names.

Cross-Step Data Flow

Variable References

Reference workflow variables with variable:name:

"prompt": "variable:prompt"

A variable's declared value is both its default and its type — a value passed in is converted to the type of the default, so declaring 25 and "25" are different things (see the schema note under Variables). Declaring null opts out of that: the variable becomes optional and untyped, taking whatever it is given and staying null when it is given nothing.

"variables": { "image": null }

A value the library already declares does not need a variable at all - see Constant References.

This is how a workflow exposes an argument a caller may pass without inventing a sentinel for its absence — a sub-workflow that behaves differently when handed an image, say. A caller can only set variables the workflow declares, so an optional argument still has to be declared to be passable.

Previous Result References

Pass output from one step to another with previous_result:step_name:

{
    "steps": [
        {
            "name": "preprocess",
            "task": { "command": "canny", "arguments": { "image": { "location": "photo.jpg" } } }
        },
        {
            "name": "generate",
            "pipeline": {
                "arguments": {
                    "control_image": "previous_result:preprocess",
                    "prompt": "a painting"
                }
            }
        }
    ]
}

A reference is resolved wherever it appears in the arguments, not only at the top of them - an argument holding a list or a nested object can reference a step too, which is what lets a constructed object be built from an earlier step.

Multiple previous_result references create a cartesian product: if step A produces 4 images and step B produces 3 masks, a step referencing both will run 12 times.

A step whose result is a dict (a task returning several named outputs, or a pipeline step that returns something like inverted_latents) can be referenced property by property with previous_result:step_name.property_name:

"inverted_latents": "previous_result:invert.inverted_latents"

Media Arguments

Images and videos load automatically for arguments named image/*_image and video/*_video. Any other argument - mask, depth_map, a controlnet's second conditioning image - can load the same way with an explicit form that says what the media is instead of relying on its argument name:

"mask": { "media_type": "image", "location": "mask.png" }

media_type is "image" or "video". location is a path relative to the workflow file, or a URL, exactly like the plain image/video forms.

Authoring a workflow from an agent

For an agent that has read the catalog (list_workflows), found nothing that produces the shape it needs, and is about to write JSON. get_schema says what is well-formed; this section says what the engine does with a well-formed document, which is where a draft that validates still fails.

References

An argument value is a reference when it begins with one of these prefixes. Each resolves before the step runs. variable: and previous_result: names are checked statically, so a bad one is a validation error at the path it sits at; a constant:, asset:, prompt: or output: name in the definition body resolves only when the step runs, and one that is missing fails the run - unless it arrives in the arguments passed to validate_workflow, which checks an asset:, prompt: or output: there for existence.

  • variable: — variable:name is the workflow's own variables entry, overridden by the caller's arguments. A variable declared null is optional and untyped. Schema validation runs before substitution, so a default must already be the JSON type the field expects: 25, not "25".
  • previous_result: — previous_result:step_name is the outputs of an earlier step, named by that step's name. It iterates; see the cartesian rule below. Validation checks it: a literal reference naming no earlier step is an error with the JSON path it sits at, rather than a run-time failure reached after everything before it has generated. A .field suffix (previous_result:invert.inverted_latents) picks one field of a result that is a dict, or a data attribute of a result object.
  • constant: — constant:module.path.NAME is a value declared in Python, read by import rather than copied into JSON. Anything callable is refused.
  • asset: — asset:name is a file in the asset library. Rooted at the library and confined to it, never resolved relative to the workflow file; a path that escapes the library is rejected. upload_asset returns one of these names. One name, three places it can sit: a bare string ("image": "asset:x.png"), an element of a list, or the location of a media dict ({"location": "asset:x.png"}, with or without media_type) - all resolve to the same file. When the file is on neither the MCP session's machine nor reachable as content= (over 4MB, or the agent has no filesystem access to read it into the call), curl it straight into the library the same way the web UI's file picker does — POST /api/uploads, body is the raw bytes, up to 200MB:
    curl -H "Authorization: Bearer $DW_API_TOKEN" --data-binary @portrait.jpg \
      "http://<host>:8765/api/uploads?filename=portrait.jpg&asset_name=cast/portrait.jpg&workspace=<ws>"
    
    It answers 201 with reference - the asset: reference to use in a workflow argument (the bearer token only when the server requires one). The same route is also how a file assembled entirely on the client - a finished cut stitched locally rather than by a workflow step - gets onto the server at all: it never lands in the gallery and export_job has no run to bundle it from, so uploading it as an asset is the only way to hand it back to the engine or to a teammate reading the workspace.
  • output: — output:<workflow identity>/<run id>/<file> is a file an earlier run wrote, under the output root and confined to it. latest in the run-id position picks the newest run that holds that file; v<N> picks the run the gallery labels v<N> (list_gallery's version), and only that run. A run id is not stable against pruning: to depend on a generated file, promote it with keep_output and reference the asset: name instead.
  • prompt: — prompt:name or prompt:folder/name is a stored prompt's text, rooted at the prompt library. That text may not itself begin with any of these prefixes; the engine rejects such a prompt rather than resolving twice. The library is also the worked-example shelf: a template's default prompt is usually a prompt: reference, and the text behind it is a caption written to whatever spec that model was trained on. Before writing a prompt for a family, read the one that is already there — list_prompts(intended_model="ltx-2.5") for the shelf, get_prompt("ltx2/fox_dawn_choir") for a body. The listing leaves the bodies out by default and reports each one's text_chars; asking for all of them at once is more than a client will accept.
  • item: — only inside a step that carries for_each: item: is the entry the member was made for, item:field one field of an object entry, spliced in whole whatever its type — a string, a number, a list of references. See "One step per entry" below.
  • gather: — gather:shot is the result of every member of the for_each step shot, in list order, as one list. Inside a list it splices into it. It is how a step downstream of a fan-out reads the whole group; previous_result:shot naming a for_each step is an error that says so.

After a long inline run that is worth keeping, get_job_workflow(job_id) returns the realized workflow — the definition with the arguments, seed and prompts of that run pinned into it — and save_workflow gives it a name, so the next run is by name rather than by pasting JSON again. export_job(job_id) bundles the whole run (workflow, manifest, job row, the media on both sides) into a directory on the server plus a zip URL, for a run worth committing or handing to someone else.

A reference is resolved wherever it appears in the arguments, including inside a nested object or list — not only at the top level. It is always the whole value: "variable:base_prompt" resolves, "variable:base_prompt, in fog" asks for a variable named base_prompt, in fog and fails the run. Nothing is interpolated around a reference. To vary a fixed prompt across steps, write each full prompt out, or put the shared text in a variable and let a step's argument override it whole. A variable: reference that names nothing the workflow declares is a validation error, not a warning: once a variables block exists the engine refuses an undeclared reference, so it is a run that cannot start.

When several steps share a block of text — a character's description and voice repeated in every shot of a dialogue short — the answer is composition rather than interpolation: write the shared text once as a variable and assemble each step's prompt with a compose_text task, whose parts are whole references joined in order. The shot then references the composed result ("prompt": "previous_result:shot_1_prompt"), so changing the voice changes it in every shot instead of in however many copies were made by hand.

Types and escaping

Any key ending in _type or _dtype, or named dtype, has its string value loaded as a Python object: "FluxPipeline" from diffusers, a dotted name ("torch.bfloat16", "sdnq.SDNQConfig") by full module path. Wrapping a value in braces keeps it a plain string — "{nf4}" is the string nf4. Getting this wrong fails at load time, after validation has already passed, so a value that is meant as text under one of those keys must be braced.

What a variable is allowed to be

A model's own rule about a value belongs in the workflow, not in engine code (CLAUDE.md) and not in a consumer's head. variable_constraints declares it per variable, in the same field names a chain step's frame_snap uses:

"variable_constraints": {
    "num_frames": {
        "modulus": 17,
        "remainder": 5,
        "min_frames": 124,
        "max_frames": 345,
        "snap": "up",
        "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps"
    }
}

The value has to be modulus * n + remainder within min_frames to max_frames. With snap: "up" an off-grid value is rounded to the next one the rule accepts and the run says so - 130 becomes 141, reported as a warning at validation time and again in the job's warnings; without snap it is refused. The bounds are checked against the value the run will use, so they hold for the rounded number: on the rule above 108 is accepted (it becomes 124) and 346 is refused (it would become 362).

Checked three times, for the reasons the task-argument domains are: in validation_errors, so POST /api/validate, validate_workflow and the pre-queue check all refuse a bad value at arguments.<name> or variables.<name> for free; at run time before anything loads, which is the backstop for a value the static pass cannot see (an inline workflow, a value a parent passed down); and in the catalog, where list_workflows and get_workflow(variables_only=true) report the rule beside the default - the half that stops the next caller picking a number the model refuses.

State the rule once. Where a template both declares a constraint and snaps a chain, the chain's frame_snap names it rather than repeating the numbers:

"frame_snap": "constraint:num_frames"

Two limits, both accepted. A constraint cannot express a bound that depends on another variable (a maximum that is fps * seconds where a template exposes fps), and it reaches a top-level variable only - not a field inside a list entry, so a for_each template whose entries each carry their own num_frames is unconstrained and relies on the run-time check.

A workflow takes only the keys the engine reads

The workflow object itself, step, task, workflow, pipeline_reference and result are closed: a property the engine does not read is a validation error naming the object and the key, not a warning. There is no when, no retry, no select - if a draft reaches for one, the shape it wants is a different arrangement of steps, not a flag. The error exists because a plausible invented key used to validate cleanly and then do nothing, so the expensive work ran with the input silently having had no effect - a mistyped sedd left the run unseeded while validation advised setting a seed, and a mistyped subfoldr put the deliverable at the run root rather than in final/.

pipeline is closed to the same rule with one opening: any key whose value is a component definition - an object carrying from_pretrained_arguments - names a component to load, because diffusers grows component names faster than the schema does (latent_upsampler, prompt_enhancer and processor all appear that way in shipped templates). A pipeline key that is not one of those is refused, which is what catches pipeline_type or model_name written a level too high. from_pretrained_arguments stays open - it passes its keys through to from_pretrained.

Where a workflow may read and reach

An argument that names a location is confined, untrusted (the default):

  • a path must resolve inside the workflow's own directory, the asset libraries, or the output root. An absolute path elsewhere is refused whether or not it exists. The remedy is upload_asset (or keep_output) and an asset: reference - which is what those exist for.
  • a glob is confined the same way, and each match re-checked.
  • an http(s) URL may not resolve to an address inside the deployment - loopback, link-local, private ranges. Its host must resolve, the body is capped at 1 GiB and a fetch at 10 minutes end to end; past that, upload_asset is the route.
  • remote_text_encoder.url is https-only, and only a HuggingFace host is sent this machine's token.
  • model_name must be a Hub repo id, or a path inside one of those roots.

All of it is reported by validate_workflow, before anything is queued, so a draft that names a file the server may not read costs nothing to find out. get_server_info's trust_workflows says which posture is in force.

A workflow you wrote inherits the catalog's VRAM ceiling

A template can declare a vram_estimate, and a workflow over it that projects past the card is refused before anything is queued. A workflow you wrote declares none, so validate_workflow (and run_workflow's pre-queue check) match each of its pipeline steps to the catalog by pipeline identity - the step's component_type, from_pretrained_arguments.model_name and from_pretrained_arguments.workflow - and project it against the ceiling a catalog template declared for that same pipeline. Past it, the answer stays valid: true and carries a warning of kind vram_projection_inherited that names the source template and the heaviest step (for a for_each, the heaviest member). It warns rather than refuses: your offload and quantization config may differ from the template's. Cut frames, size or references until it goes quiet, or declare your own vram_estimate, which always wins and is judged exactly as a template's. A pipeline the catalog declares no ceiling for gets no warning - silence there is not a clearance - and H3's t2va and ref2va are different identities with different ceilings.

The projection is base_gb, plus bytes_per_voxel times the product of the voxel_variables, plus gb_per_reference per non-null reference, plus bytes_per_guide_voxel times each guide's frames times the canvas (every voxel variable but num_frames). The guides counted are the step's own guides with a non-null video, at their snapped length, and the one a continuity: "guide" chain appends (guide_frames, 22 or 39). A guide clip is probed at validate, at admission and at run; one that cannot be probed before the run (a previous_result: clip) is charged at the step's num_frames, and the message says it was charged at the worst case. H3's t2va and fl2va templates declare the guide term, so a guided step is judged by what its guides add rather than passing on the clip alone.

Remote code is refused by default

A server started without --trust-workflows refuses any from_pretrained_arguments that sets trust_remote_code or custom_pipeline, at load time, after validation has passed. Use a pipeline diffusers ships: no bundled catalog entry carries either key, and a workflow that does runs only on a server whose operator turned trust on, which get_server_info reports as trust_workflows.

Several previous_result references multiply

When one step carries two or more previous_result references, the engine runs that step once for every combination — a cartesian product. Four images and three masks is twelve iterations, not three pairs. Past 10000 combinations the run is refused outright.

This is deliberate: it is how one prompt fans out over a set. The consequence is that a pairing — shot i with speaker i, prompt i with portrait i — cannot be expressed with two references on one step. Write it as one step per pair, each referencing exactly the two things it pairs, or gather the pairs upstream so each is a single result. A step that seems to need a "zip" is the signal to restructure the workflow, not to add another reference.

One step per entry: for_each

A step that carries for_each runs once per entry of a list — a shot per entry of shots — and the list is a variable the caller supplies, so a six-shot episode is an argument rather than a different file.

{
  "name": "shot",
  "for_each": "variable:shots",
  "pipeline": {
    "arguments": {
      "prompt": "item:prompt",
      "references": "item:references"
    }
  }
}

with

"shots": [
  { "name": "wide_open", "prompt": "the band walks on, wide",
    "references": [{ "reference_type": "…", "from_previous_result": "draw_singer" }] },
  { "name": "closeup", "prompt": "closeup on the singer",
    "references": [{ "reference_type": "…", "from_previous_result": "draw_singer" }] }
]

and downstream

{ "name": "edit",
  "task": { "command": "concat_videos", "arguments": { "videos": "gather:shot" } } }

Before the run starts, the engine replaces the for_each step with one ordinary step per entry, named shot@wide_open, shot@closeup — the entry's name, or its index for an entry without one. Those are the names the manifest, the job's events and the gallery show, and @ is reserved for them: a hand-written step name may not contain it. An entry's name must be unique in its list and match ^[a-zA-Z_][a-zA-Z0-9_-]*$. Give entries names: the step cache keys on the member name, so a shot inserted in the middle of a named list leaves every other shot cached, while an indexed list shifts every later shot onto a different entry and regenerates it.

item:field is the whole value of that field, so an entry can carry anything a step argument can — including a references list whose length differs by shot, with from_previous_result and asset: strings inside it. Nothing is interpolated: "item:prompt" is the field, "shot: item:prompt" is a literal string.

A field an entry may leave out is read with the object form {"item": "singer", "default": "variable:singer_reference"} (the field's bare name, not item:singer). An entry that carries singer gets its value; one that does not gets the default, spliced in whole. An object with no default is the plain required read. The object must have exactly those keys, and it means something only inside a for_each step. music-video uses it so a shot can condition on its own picture (singer) and the rest fall back to singer_reference; the audio slice stays in the template. If every entry supplies the field and the default's step is read nowhere else, that step is elided like any unread one.

An entry may name another variable: "from_file": "variable:character_a_voice" inside a references entry is that variable's value by the time the member exists, so one variable sets a voice in every shot the character speaks in and a caller who supplies the list still writes variable: for the parts the template fixes. Those references are resolved before anything in the entry is loaded, and an undeclared one is a validation error at the entry's path (arguments.shots[2].references[1].from_file when the list is yours, variables.shots[...] when it is the template's). A value may not reference itself, directly or through another variable.

Two for_each steps over the same list are paired by key — the entry's name, or its index for an entry without one: inside shot@closeup, a reference to another for_each step slice over the same shots list resolves to slice@closeup. That is how a shot reads the audio slice cut for it when slicing and generating are two steps. It is the one pairing the engine has; for_each runs over exactly one list, and there is no zip and no loop index.

That auto-pairing is at the step level only, where from_previous_result (or previous_result:) sits directly in the for_each step's own arguments. Inside an item — an entry field the step reads with item:, such as a references entry — nothing pairs it for you, because the entry is data the caller wrote and the engine does not know which list it came from. Spell the member out: {"reference_type": "…", "from_previous_result": "slice@closeup"}, not "slice". Getting this wrong fails the run with a precise error naming the members that do exist, but the guide says it here first.

A member may reference an earlier member of its own list directly ("image": "previous_result:shot@accuse" inside shot@deflect), but it cannot interpose a derivation on that reference — "shot 2 starts where shot 1 ended," a still pulled from shot 1's last frame with get_last_frame. There is no step between two members of the same for_each list: a step meant to derive that still (say, named last) would itself have to sit after shot in the file to read shot@accuse's result, which makes it a step after the whole list, and referencing it from inside the list is a previous_result: naming no earlier step — refused at validate. Split the list at the dependency instead: run the members up to and including the one being derived from as their own for_each (or, for a single entry, an ordinary step), add an ordinary get_last_frame step after it reading previous_result:shot@accuse, and run the remaining members as a second for_each whose dependent entry takes that derived still — as a variable value or, over the API/MCP, folded into the entry itself — rather than a bare previous_result: reference to the video.

Limits: a list has at most 32 entries, and an empty list is a validation error — the step would run nothing. Validation realizes a constant: default before checking it, so a list defaulted to a constant validates the same way it will run. release_pipeline on a for_each step releases after the last member. Each entry is a full generation, so quote the cost before running a list-driven workflow: the listing's lists block names the fields an entry takes and the steps over it, and its cost carries per_entry once one entry has been measured. validate_workflow with your arguments answers with a plan whose estimate already does that arithmetic (basis: per_entry); without per_entry it extrapolates the stored total linearly over your list (basis: derived - an estimate rather than a measurement) and reports the stored total unchanged only when your list is the one it was measured with (basis: catalog). Ahead of all of those it quotes this box's own finished runs of the shape you are about to run when it has any (basis: observed, with runs saying how many) - quote the plan's figure and say which basis it has. An entry key no step reads is a validation warning at the entry's path, so a misspelt field is caught before the run. Then validate_workflow with the arguments you will run with: it expands your list, not the template's default, resolves the variables your entries name, and reports a duplicate name or a missing field at the entry's path.

Every error carries a path in the file you wrote, not in the expanded step list: a bad reference inside a member is reported at the for_each step's own path, with the member it failed in named in the message.

templates/minimax/dialogue-short and templates/minimax/music-video are this shape: each takes one shots list, and get_workflow on either shows the entry an item needs. templates/ltx2/restore-long runs two for_each steps over one windows list - a window_video slice, then a restore of it

  • and gather:s the restores into join_windows; its list length is fixed by the source, and validate_workflow names the entries to add or drop.

The loop

  1. validate_workflow — free and instant. It reports every schema error at once, each with the JSON path it sits at, plus warnings for argument names that do not appear in the real pipeline signature. It also catches a previous_result: (or from_previous_result) that names no earlier step, which is what renaming a step half way through leaves behind. Pass the arguments you are going to run with as well: a name the workflow no longer declares, a value that will not coerce to the declared type, and an asset:, prompt: or output: reference that names nothing in this workspace each come back at arguments.<name>, for free, instead of after the model has loaded. Without them the verdict is about the stored definition and its stock defaults - checked_arguments in the answer says which it was.

  2. Fix everything reported, including the warnings: a passing validation does not mean the pipeline accepts the arguments, and a typo against a real __call__ shows up only as one of those warnings. The server only computes signature warnings for a schema-valid draft — while schema errors remain it returns warnings: [], so validate again after fixing them to see the warnings.

  3. save_workflow — validates again on the way in and returns the catalog metadata the saved draft will carry.

  4. run_workflow with acknowledged_cost=true, after telling the user what it costs. Without the acknowledgement the call is refused. The figure to tell them is the plan on the validate answer - estimate.minutes with its basis, and every downloads_required entry named as its own line item, since weights not on this box are minutes and gigabytes the cost block never counted. basis says where the figure came from, and that is what decides how to quote it: observed is this box's own finished runs of this shape (the cold median over runs of them, preferred over any curated figure) - "about N minutes, measured over M runs"; per_entry is a measured per-entry rate re-priced for your list; catalog is a measured total for a run whose lists are the ones it was measured with; derived is that total extrapolated over a list you changed the length of - say it is an estimate; inherited is an inline workflow priced from this box's runs of the catalog template with the same pipeline (inherited_from) - say it is an approximation, and relay the warning if offload, quantization or frame count differ; other_device is a figure from another accelerator - say so too; unknown is no figure at all. gb on a downloads_required entry is null when the hub could not be asked, and steps/list_entries say how many members the list actually produced. Then pass that plan back: acknowledged_cost={"fingerprint": plan.fingerprint, "minutes": plan.estimate.minutes, "downloads": [...]} - the server refuses with 409 if the run's shape changed since the quote, and the refusal carries the new plan to quote from. true is for a plan that was null. When basis is unknown: a workflow you wrote or copied carries no cost of its own, but the pipeline inside it usually does: list_workflows(include_models=true) finds the models/ entry that loads the same checkpoint, and its per-image figure times the number of images is the number to quote. Say "a few minutes" only when no entry with that pipeline has been measured.

  5. wait_for_job rather than a polling loop, asking for the job's plan.estimate (in seconds, plus a margin) as timeout_seconds, so one call can cover the whole run. The server caps a single call, and the tool's description states the cap: the reply's timeout_applied_seconds and timeout_capped say what you got, and waited_seconds how long it took. Call it again if it returns still_running: true. Without timeout_seconds a call waits 20 seconds. A running job's progress carries the step being run; the phase (loading, generating, decoding, saving) with the model it names in phase_detail; seconds_in_phase, time spent in that phase; and seconds_since_event, time since the last progress event - a number that climbs while denoise_step stays put is the "nothing is happening" read. It also carries denoise_step/denoise_total_steps, null until the denoise loop starts; judge a slow run against a stuck one by whether denoise_step has moved since a poll minutes ago, not by silence past a fixed threshold. A null denoise_step under generating is the pipeline's lead-in - encoding the prompt and every reference - which emits nothing and can run for many minutes when a video reference is among them; gaps between denoise steps are uneven too where a transformer block cache is configured. Both are normal, and the model family's own skill carries the measured figures. denoise_total_steps is the schedule the pipeline actually runs, which is not always the num_inference_steps that was asked for: MiniMax H3's scheduler counts sigma grid points including the terminal zero, so it runs N-1 model evaluations for N (9 reports 8, 20 reports 19) - the vendor's convention, not a dropped step; raising the number still buys the steps it looks like it does.

  6. get_output_image to look at what was actually made, and say whether it answers the request. Nothing before this step establishes that it does. get_output_frames looks at a video and get_output_audio listens to a soundtrack. A text-only client reads levels and clipping from get_gallery_metadata instead (media peak and mean, findings; envelope=true for where).

    To confirm the words a clip speaks - a text-only client can't consume the AudioContent block get_output_audio returns - run templates/check-script against its script rather than reading a transcript by eye. validate_workflow(name="templates/check-script", arguments={"input_audio": "output:<name>", "lines": [...]}) first (free; it takes an audio file or a video's muxed soundtrack directly, and lines is the script in order, H3 markup allowed, {text, shot} to name a line's shot, [] for a take that should be silent), then run_workflow(..., acknowledged_cost={"fingerprint": ..., "minutes": ..., "downloads": [...]}) bound to that plan with wait_seconds=60, then get_output_text on the result, and delete_output(job_id=...) the scratch run afterward. Its findings are the places to listen: a dropped or changed line (line_mismatch, with heard), a markup word spoken aloud, speech in a shot meant to be silent, a last word clipped by the cut. Words Whisper invents over silence or music are under discarded, not findings. This workflow's plan comes back basis: "unknown" with minutes: null - nothing is curated or observed for it - so quote what it actually takes rather than the plan: seconds, not minutes (a few seconds per clip in practice). Four calls and a short wait, not a GPU-spending read tool - keep the normal queue rather than adding one.

  7. Getting the files to the user's machine. download_output and export_job write on the machine running dw.serve, which over a remote --mcp endpoint is the GPU box. The last mile of every deliverable is the url each list_gallery entry carries (or export_job's zip_url), fetched with the same bearer token the MCP connection uses:

    curl -H "Authorization: Bearer $DW_API_TOKEN" \
         -o exports/still.png "http://<box>:8765/outputs/ltx2/Gyre/20260910-.../still.png"
    

    Put the result under exports/ in the session's working directory - it is the user's deliverable, not a temporary file.

Keeping a set consistent

"Four pictures of the same thing" is the commonest shape a request takes that the catalog does not name directly, and what "the same" means decides the workflow.

  • The same style, different subjects or scenes. One prompt per picture, the same seed on the workflow, and a shared style phrase in every prompt. A shared seed does not make the pictures alike; it makes the run reproducible. Consistency here comes from the prompts.
  • The same object, differing in one stated way - four spoons identical but for colour, one mug in four glazes, a product in each colourway. Generate the object once, then run an image-edit pass per variant with the base step's result as its image and an instruction that names only the change ("make the mug red"). Separate generations, seeded or not, draw a different object every time; an edit holds everything the instruction does not mention. templates/consistent-set.json is this shape.
  • The same character in different situations. A reference rather than an edit: an identity-referencing pipeline or IP-Adapter conditioned on one portrait, used by every picture (templates/ip-adapter.json, templates/multi-image-reference.json, and for video the MiniMax reference-to-video and dialogue-short templates). The identity-referenced trait in the listing marks the workflows that take one.

A spoken scene breaking into a song

A musical number: spoken shots, each carrying its own dialogue audio, then shots lip-synced to one song, with the song entering under the last spoken line. A concat_videos join is wrong here - it keeps each sung shot's own audio, the separate slices it was generated against, where the number wants the song unbroken. The task for it is join_into_song (get_guide("tasks", section="join_into_song")), and the recipe is four parts:

  1. Pick cue_seconds: the song time that lands on the first sung shot's frame 0. The song enters that long before the cut, under the last line; 0 starts it exactly at the cut. It may not be longer than the dialogue.
  2. Generate the sung shots against slice_audio slices of the song, as templates/minimax/music-video does (its slice, shot and trim steps; it slices lead_frames early and trims the run-up off). The first slice starts at cue_seconds, and each next one starts where the one before it ended, so the slices tile the song with no gap: start_seconds is cue_seconds plus the length of every sung shot before it (num_frames / fps - 24 fps for H3), duration_seconds its own length. The shots, not the slices, are what the join reads.
  3. join_into_song with the dialogue shots, the sung shots, the unbroken song and the same cue_seconds. It measures the joined dialogue at run time and places the song so cue_seconds lands on the first sung frame, ducks the dialogue under it, and discards the sung shots' own audio. It does not set the output level.
  4. normalize_audio to -3 dBFS, then pair_audio, as every template that muxes to video does - the encoder can add up to ~2 dB, and the written-peak warning reads the file.

The tail, with the shots from earlier runs (output: or asset:) or from earlier steps of the same workflow (previous_result:, gather:):

{
    "name": "joined",
    "task": {
        "command": "join_into_song",
        "arguments": {
            "dialogue": ["asset:scene/line-1.mp4", "asset:scene/line-2.mp4"],
            "song_shots": ["asset:scene/sung-1.mp4", "asset:scene/sung-2.mp4"],
            "song": "asset:scene/song.mp3",
            "cue_seconds": 1.5
        }
    }
},
{
    "name": "balanced",
    "task": {
        "command": "normalize_audio",
        "arguments": { "audio": "previous_result:joined", "peak_dbfs": -3.0 }
    }
},
{
    "name": "number",
    "task": {
        "command": "pair_audio",
        "arguments": {
            "video": "previous_result:joined",
            "audio": "previous_result:balanced",
            "fit": "video"
        }
    },
    "result": { "content_type": "video/mp4", "subfolder": "final" }
}

joined and balanced save nothing and are read by later steps, so they run and write no file of their own; number is the deliverable, and its shot map (get_gallery_metadata's media.shots) names one shot per input. The frame rate comes from the shots themselves - pass fps to the join only when none of its inputs carries one.

Saying which output is the deliverable

A run writes everything into one directory, so a finished episode sits beside the twenty scratch files that went into it. A step's result block can name a subfolder of the run directory for its files:

"result": { "content_type": "video/mp4", "subfolder": "final" }

The convention is two names: final for a step whose output the user will be shown, intermediate for everything else. The engine treats no name specially and applies no default - a step that says nothing writes to the run's root as it always has - but the gallery, get_job and list_gallery all carry the value, so a consumer that follows the convention can tell the deliverable from the scratch without knowing the workflow. Mark every saving step of a multi-step workflow; a one-step workflow needs nothing.

The shipped templates follow it: every template with two or more saving steps marks each one, so a workflow copied from a template starts with the roles in place.

The value is a relative path of any depth (shots/act-1), may be a variable: or, inside a for_each step, an item: reference, and follows the output: segment rule - each segment starts with a letter, digit or underscore; .., a backslash and a leading . are refused - so every subfolder written is one a later workflow can name: output:dialogue-short/latest/final/episode.mp4. A bad value is a validation error at its JSON path. file_base_name is a name, not a path: a separator there is refused, and subfolder is the way to place a file. It replaces the name the engine would derive from the workflow and step rather than prefixing it, so "file_base_name": "episode" in a final subfolder writes final/episode-0.0.mp4 - name each step that sets one differently, or the second collides and picks up a -2.

A step that saves nothing and which no later step reads does not run at all: the engine drops it before the first step executes and warns once per dropped step. That is how a template whose portraits can be supplied as asset: files stops paying for the steps that would have drawn them. It follows from what the definition says, never from a value produced during the run, so it is decided at validate time too - the plan a validate call answers with counts only the steps that will run and lists the rest under elided_steps. Four things keep a step: a result with a content_type and save not false, being the last step, being read by a later step (previous_result:, gather:, a pipeline_reference, a shared component), or being read by a step that is itself kept - elision is transitive. If a step you meant to run is named in the warnings, a reference to it is misspelled somewhere later or it needs a result.

Composing a stored workflow

A step with a workflow block runs another workflow as one step of this one, with arguments handed down as that workflow's variables. Its path is read the way run_workflow's workflow_path is - a catalog name from list_workflows, with or without .json, a path relative to the file that names it, or builtin:name.json - and resolves beside the referencing file first, then in this workspace's workflows/, then in each read-only source the server lists. A stored template is composed by its catalog name; copying it into the workspace to reach it is no longer necessary, and a copy silently stops tracking the original.

Declare a result on the composing step and the composed output is saved there, once, under that step's name and subfolder - the composed workflow's own last step does not write a second copy. Its other steps write into the same run directory, prefixed with the composing step's name.

validate_workflow resolves the path, so a name that reaches nothing is an error at steps[N].workflow.path before anything is queued; it also validates the workflow named, refuses a composition cycle, and warns about an argument the composed workflow declares no variable for.

Being found next time

The catalog derives each entry's shape — one of image, image-set, image-edit, shot, sequence, audio, text, utility — and its traits (has-audio, chained, image-conditioned, identity-referenced, needs-input-media, composes-workflows) from the structure of the definition, and its summary from the first sentence of description.

So write that first sentence to say what the workflow makes and what it needs supplied, in under 120 characters — "H3 video with audio between two supplied stills" — rather than what technique it demonstrates. A first sentence longer than that is truncated with an ellipsis in every listing.

Declare shape, traits or summary at the top level only when derivation gets it wrong; a declaration that merely repeats the derivation is noise that rots when the rules change, and the repo's catalog tests refuse it. cost is never derived — leave it absent until a run has been measured. That makes it the catalog's verified marker as well: an entry carrying cost has been run to completion on the device it names, and one without has only been authored — its description may still say what it has not been able to check (VRAM at a size, whether a format carries what the pipeline returns), and the first run is the one that finds out.

cost_drivers is the other half of saying what a workflow costs, and it is for derivation: the variables that move the wall clock — a frame count, a step count, a segment count, the list a for_each runs over — never a prompt or a seed. The server buckets its own finished runs by those values and reports the result as observed beside the curated cost, so a 345-frame run never informs a 124-frame figure and a list driver buckets on its length. Declaring none is not neutral: the figure then falls back to runs that overrode nothing at all, which most real runs do, so a measured workflow with no drivers keeps answering "unknown". Each name must be a variable the workflow declares — tests/test_observed_cost.py sweeps the catalog for one that is not, since a driver bucketing on nothing looks exactly like a driver that works.

Assessing a run's output

A cut joined from shots can succeed and still be wrong at a seam, and the whole-file numbers get_gallery_metadata reports cannot see inside a join. assess_output(name) (GET /api/gallery/{name}/assess) measures that file on the server. It decodes the file once, runs every assessment probe that applies to it (analyze_shots, analyze_seams, analyze_sync_drift), and returns where to look. It queues nothing: the probes use only the CPU and run in the server process, beside whatever job holds the GPU. name is a gallery name, output: or asset:. The shot boundaries come from what the file's run recorded: the run manifest for an output, and the sidecar keep_output wrote for an asset.

A last shot's num_samples a few dozen samples off round(num_frames * sample_rate / fps) is expected, not a finding - see pair_audio in the tasks guide's Video Processing section for why.

Procedure.

  1. After wait_for_job, call assess_output on the deliverable, which is the file under final/. get_gallery_metadata points at the tool whenever media.shots is set.
  2. Read findings. When the list is empty, no rule crossed its threshold, which is a reason to listen less closely, not a pass.
  3. Drill into each finding at the place its at names. For a seam, use get_output_frames(name, seams=[n]) to see it and get_output_audio(name, start, duration) to hear the second around it. For a shot, look at that shot's span. Judge it against the request.
  4. Fix what you confirmed (see the table below), then assess the new cut.

Pass probe="analyze_seams" (or either of the other two probe names) to get that one probe's full body: every seam's or shot's measurements, not just the ones that crossed a rule. detail=true adds every applicable probe's full body under probes. An unknown probe is refused before anything is read, and the error names the three probes.

The answer.

Field What it holds
findings Every rule a measurement crossed, each {rule, severity, at, value, threshold, says}. severity is warn or info. at names the shot or seam.
rules_applied The rules that were checked, so an empty findings list says which checks came back clean.
rules_skipped {probe, rule, reason} for each rule that could not be measured on this file.
not_applicable {probe: why} for each probe that does not apply to this file. A still has no shots, seams or soundtrack. A mute file has no levels. A file with no recorded shots has no seams.
shots_source Where the boundaries came from: manifest (the run's manifest, or an asset's sidecar), or none.

The thresholds live in one table, dw/assessment_rules.py:

Rule Probe Fires when
shot_level_spread analyze_shots the shots' RMS levels span 6 dB or more
seam_level_step analyze_seams the shots either side of a seam differ by more than 3 dB
seam_click analyze_seams the join peaks more than 12 dB above the audio either side
seam_hole analyze_seams the join's floor drops below -50 dBFS while both sides are voiced (above -30 dBFS)
seam_frame_jump analyze_seams the picture changes more than 25x as much across the seam as inside either shot (info, and skipped at a shot marked hard_cut: true)
shot_dead_air analyze_shots the shot holds a gap more than 0.4 seconds long below its dead-air floor (skipped inside a shot that is quiet throughout)
sync_drift analyze_sync_drift by a shot's end, the audio sits more than 40 ms off the picture
sync_length analyze_sync_drift the soundtrack and the picture differ in length by more than 40 ms

A shots record whose start_frame/num_frames already reaches past the file's own length is not measured against a threshold - it is clipped to the file before any of the above run, and that clip is itself a shot_span_overrun finding on all three probes (analyze_shots, analyze_seams, analyze_sync_drift), with value naming how far past the end it reached. validate_workflow catches the same mistake before the run for a shots argument and an asset:/literal video whose length is knowable ahead of time; it cannot for previous_result:/output: video not yet written, so that case is left to the finding above.

Authority. A finding marks a place to look, not a verdict. Nothing in the engine acts on one, and no run fails because of one. A finding you have checked and accepted is simply left alone. A seam_frame_jump at a cut the story wanted is the cut working, and a level step into a quieter scene can be the scene. Tell the person what you confirmed, not what the probe reported.

Remediation. A recut reruns only the join over the shots the run already made: each entry in videos is output: + the run's intermediate/ shot file. That is cheap, and generates nothing new. A regenerate is a new run, so quote its plan.estimate first.

Finding Fix Kind
shot_level_spread, seam_level_step match_levels: "rms" (with match_levels_dbfs for the target) on the concat_videos / dissolve_videos step recut
seam_click a longer crossfade_ms on the join recut
seam_hole audio_bleed_ms on the join, so the outgoing tail rings on across the seam recut
seam_frame_jump a dissolve_videos join, or regenerate the incoming shot from the outgoing shot's last frame. If the cut was meant, leave it alone recut, or regenerate
shot_dead_air find_loop_bed on the cut names a quiet stretch inside one shot and its gain; cut that bed with slice_audio, loop_audio it to the gap's length, and mix_audio it under the line rather than leaving the drop silent recut
sync_drift regenerate the shot. Drift inside a shot is the model's, not the join's regenerate
sync_length rerun the mux through pair_audio with fit: "video", which cuts or pads the track to the picture recut

Result Configuration

"result": {
    "content_type": "image/jpeg",
    "save": true,
    "file_base_name": "episode",
    "subfolder": "final"
}

Supported content types: image/jpeg, image/png, image/webp, image/gif, video/mp4, audio/wav, audio/flac, audio/mpeg (mp3), audio/ogg, audio/opus, audio/aiff, application/json, text/plain (plus the common aliases audio/x-wav, audio/mp3, audio/vorbis).

A task command's implementation declares what it hands back - most answer an artifact (a file result saves in one of the media content types above), some (judge) answer a bare scalar that cannot be saved at all, and some (the assessment probes in TASKS.md, find_loop_bed, and attribute_voices) answer a json document - every measurement taken, in one dict. A step on a json command must set content_type to application/json, which saves it as one document; a step on a scalar command may not carry a result at all. Both are checked in validation, by the command's own declared kind rather than a name match.

subfolder places the step's files in a subfolder of the run directory - see Saying which output is the deliverable above. file_base_name is the base name the step's files are written under, replacing the name derived from the workflow and step; it may not contain a path separator.

For video, "fps" is the rate the file is written at. It is rarely needed: frames that know their own rate carry it - a video read from a file or an asset:, a concat_videos/dissolve_videos join, an interpolation - and the engine writes them at it. Frames that bring no rate (most generations) fall back to 8, so a workflow that assembles from bare frames should say what they run at. A declared fps always wins over the carried one and warns when the two differ, which is how a deliberate slow motion is written. For audio, add "sample_rate": 44100 when the waveform doesn't already carry a rate of its own (a declared rate always wins). Setting embed_metadata: true on an image result embeds the step's model name and arguments as generation metadata - PNG info chunks for image/png, EXIF UserComment (via piexif) for image/jpeg and image/webp.

A pipeline that generates a video with its own audio track (LTX-2, or a modular pipeline whose output asks for both videos and audio) is muxed into one video/mp4 file with PyAV. audio_sample_rate overrides the rate the pipeline itself reports, for the rare case it needs correcting.

Audio Encoding

Audio is written through soundfile, so both lossless and compressed containers work:

"result": {
    "content_type": "audio/mpeg",
    "sample_rate": 44100,
    "compression_level": 0.3
}
  • subtype — encoding subtype, such as "PCM_24" for wav and flac. Defaults to the container's own default, which is "PCM_16" for wav and flac.
  • compression_level — 0.0 to 1.0 for flac, mp3 and ogg. Higher means smaller files.
  • bitrate_mode — "CONSTANT", "AVERAGE" or "VARIABLE" for compressed formats.

audio/opus writes an Opus stream in an ogg container, and only encodes at sample rates of 8000, 12000, 16000, 24000 or 48000.

Output files are saved as {output_dir}/{base_name}-{result_index}.{artifact_index}.{ext}, where base_name is {workflow_id}-{step_name}.{step_index} unless the step's result sets file_base_name, which replaces it entirely. step_index is the step's position in the workflow, result_index counts the argument-combination iterations the step ran (see cartesian product, above), and artifact_index counts multiple artifacts within one result (num_images_per_prompt > 1, or a dict result saved key by key). The derived name is what makes two steps' files distinct, so when you replace it on more than one step in the same subfolder, give each a different name - otherwise the second one gets a -2 counter.

Pipeline Configuration

A step's configuration is dw's own vocabulary rather than the model's — each key drives a different call — so it is a closed set: a name the schema does not declare fails validation instead of being ignored. That matters most for the keys it would otherwise be quietest about. A misspelled offload used to validate, load, and run with no offloading at all, surfacing as an out-of-memory error with nothing pointing at the spelling; it now fails before the first model loads. Model-side values that are not part of this vocabulary have blocks of their own: from_pretrained_arguments for the constructor, arguments for the call, and configs for a modular pipeline's block configs.

Memory Offloading

Control how models use memory:

"configuration": {
    "component_type": "FluxPipeline",
    "offload": "model"
}
  • "model" — Moves entire models between CPU and GPU. Good balance of speed and memory.
  • "sequential" — Moves individual layers. Slowest but uses least GPU memory (measured at 6x the per-step cost of model offload on FLUX dev and 3.4x on Z-Image; a streamed block-level group_offload of the transformer is the usual better answer - see RECIPES_24GB.md). On MPS it is downgraded to "model" with a warning: with unified memory there is no separate pool to keep small, so the per-layer copies cost speed and save nothing. exclude_from_cpu_offload names components the sweep should leave alone.
  • Omit for no offloading (fastest, requires enough VRAM).

For components the pipeline loads itself — which is all of a modular pipeline's — use components, applied once the pipeline is loaded:

"configuration": {
    "component_type": "ModularPipeline",
    "components": {
        "transformer": {
            "group_offload": {
                "offload_type": "block_level",
                "num_blocks_per_group": 1,
                "use_stream": true
            }
        },
        "text_encoder.model": {
            "group_offload": { "offload_type": "leaf_level", "use_stream": true }
        },
        "vae": { "device": "cuda", "residency": "on_demand" },
        "audio_vae": { "device": "cuda" }
    }
}
  • group_offload — streams the component between system memory and the accelerator a block or a leaf module at a time, which is what fits a component larger than the device. offload_type is required ("block_level" or "leaf_level"); onload_device defaults to the pipeline's device and offload_device to the CPU. Anything else in the block is passed through to apply_group_offloading, so use_stream, num_blocks_per_group, low_cpu_mem_usage and offload_to_disk_path work as diffusers documents them.
  • device — moves a component that is small enough to stay resident.
  • residency — "resident" (the default) leaves the component on its device for the whole run; "on_demand" rests it in system memory and moves it to the device only while one of its own calls runs. See On-demand components.
  • enable_tiling — tiled decode for a decoder not named vae (LTX-2.5's diffusion_decoder, for example).
  • attention_backend — a persistent set_attention_backend on one component, which a compiled component needs (the pipeline-level attention_backend applies per call).
  • attn_processor_type — the attention processor the component runs, constructed with no arguments and handed to set_attn_processor. The unet and transformer blocks cover those two; this covers any other component that carries attention (LTX-2.5's diffusion_decoder, whose default processor is a portable fallback rather than the NATTEN path the decoder was built around).
  • compile, truncate_layers, remove_modules — see ACCELERATION.md.
  • A dotted key reaches a module inside a component, for a component that holds the model rather than being one.
  • A components block that group offloads anything, or marks anything on_demand, already keeps the pipeline itself off the device - the components are placed individually, so moving the whole pipeline would load it in full before the hooks and wrappers exist. Nothing extra is needed for that.

preserve_device_placement covers the case that is left: a component loaded already placed, which must not be moved afterwards. A device_map load or a quantization that pins its tensors to one device is the usual reason.

"transformer": {
    "configuration": {
        "component_type": "FluxTransformer2DModel",
        "preserve_device_placement": true
    },
    "from_pretrained_arguments": {
        "model_name": "black-forest-labs/FLUX.1-dev",
        "subfolder": "transformer",
        "device_map": "cuda"
    }
}

Renamed: this setting was do_not_send_to_device. The old name is no longer recognized - a workflow still using it will load the component and then move it to the device anyway, since an unknown key is ignored rather than rejected. Rename the key.

On-demand components

"residency": "on_demand" sits between the two placements above. A device component holds VRAM for the whole run, wasted on a component used twice; group offloading streams per submodule forward, so it restreams the model once per call of every leaf - ruinous for a VAE, whose tiled decode calls its blocks once per tile. On-demand moves the model as a whole around each call, so a tiling loop sits inside a single pair of transfers.

"components": {
    "vae": { "device": "cuda", "residency": "on_demand" },
    "audio_vae": { "device": "cuda", "residency": "on_demand" }
}

The component rests on the CPU and is moved to device around whichever of forward, encode and decode it defines, then moved back and the freed VRAM released to the driver. Nested calls are counted, so a decode that calls forward internally is moved once, not twice.

  • Use it for a component that is large but called a handful of times - a VAE that encodes references at the start and decodes the result at the end. Freeing it for the denoise loop is the whole point.
  • Not for a component called every step. A denoising transformer would pay per-call transfers 20-50 times; group offloading is the tool for those.
  • Cannot be combined with group_offload on the same component - a group offloaded module holds one group at a time and ignores whole-model moves, so the two cannot both own its placement. Configuring both is rejected at load.
  • Ignored when the component's device is the CPU, where there is nothing to move it off of.

On a 24GB card, templates/minimax/reference-to-video.json peaks at 18.9GiB of reserved VRAM with on-demand VAEs against 23.2GiB resident, and the tighter resident fit costs 40 allocator retries - cache flushes forced by a failed allocation - where the on-demand run has none. The headroom is also what lets the chained variant run: its later segments carry an extra reference and need ~1.9GiB more than the first.

The same holds for the frame-conditioned workflows. Generating 124 frames at 960x544 from a keyframe, with everything else held equal:

VAE placement peak reserved allocator retries
resident 22.71GiB 22
on-demand 18.03GiB 0

The resident run also logs a memory mapping failed with OOM warning per retry, with as little as 3MB free while it tries to map 20MB. It completes - the allocator flushes its cache and succeeds on the retry - but each one is a synchronising stall, and a run that close to the limit fails outright on any workload that needs slightly more. Every MiniMax H3 example uses on-demand VAEs for this reason.

Example: reference-to-video.json, image-to-video.json

Releasing a pipeline mid-workflow

Pipelines stay loaded for the whole run (and across runs, in the server's persistent worker) so repeated steps reuse them. When a workflow chains two large models that cannot both fit - generate with one, upscale with another - release the first once its step completes instead of configuring offload on everything:

{
    "name": "generate",
    "release_pipeline": true,
    "pipeline": { ... }
}

The step-level release_pipeline flag unloads the step's pipeline after its results are saved. A later pipeline_reference to a released step is an error, and the process-wide step cache will not retain it.

What keeps a pipeline warm: the server's worker keeps a pipeline loaded across runs for as long as what it loads is unchanged. A LoRA's scale and alpha and a scheduler's (or audio_scheduler's) shift are applied in place on the warm pipeline, so iterating on them costs no reload; removing (nulling) an alpha or a shift reloads, so the checkpoint default comes back. Changing a LoRA's model_name or weight_name, the quantization, the placement or the scheduler type reloads it.

Releasing task models mid-workflow

Task models - the checkpoints behind text_generation, segment, depth_estimator and the rest - are cached separately from pipelines, so that a step running its task once per result does not reload the same weights on every iteration. Nothing evicts that cache during a run, which matters when a task loads a large model on the device ahead of a generation step: a prompt-expanding language model would hold its weights for the whole run. release_models clears it once the step completes:

{
    "name": "expand_prompt",
    "release_models": true,
    "workflow": { "path": "builtin:h3_context_ir.json", "arguments": { ... } }
}

The flag applies to any step type, and on a workflow step it fires once the whole sub-workflow has finished. It clears every cached task model, not only this step's, and a later step needing one of them reloads it.

Example: enhance-prompt.json

A step nothing reads does not run

Before the first step executes, the engine drops any step whose result no later step reads and which writes no file, and warns once per dropped step saying which and why. dialogue-short cast from portraits that already exist used to run its two Z-Image steps anyway and throw the pictures away - about a minute of GPU per episode on something nothing looked at (#122).

Four things keep a step:

  • it saves - a result with a content_type, and save not false. A workflow whose whole point is writing three images references nothing, so this is the rule that keeps elision from being destructive. "save": false is how a step says it is scaffolding.
  • it is the last step - it is the run's answer, whatever it declares.
  • something reads it - previous_result:/from_previous_result (including previous_result:step.property), a gather: (which is a list of those by the time this runs), a pipeline_reference naming it, or a reused_components entry naming a component it shares.
  • Elision is transitive, so dropping a step can drop the step it read in turn.

release_pipeline on an elided step moves onto the last surviving step before it when that step loaded the same pipeline, and release_models moves unconditionally - a release that vanished with its step would leak the memory it existed to free. The plan a validate call answers with is computed after elision, so steps, downloads_required and the cost it quotes are the work that will actually happen, and it lists what was dropped under elided_steps; the run manifest records the same list.

If a step you expected to run is named in the warnings, the usual cause is a reference to it spelled wrong somewhere later, or a step that was meant to declare a result.

VAE Options

"configuration": {
    "vae": {
        "enable_slicing": true,
        "enable_tiling": true
    }
}
  • enable_slicing — Process VAE in slices to reduce memory
  • enable_tiling — Tile large images through the VAE

LoRAs

Attach one or more LoRAs to a pipeline with loras, a sibling of configuration:

"loras": [
    { "model_name": "XLabs-AI/flux-RealismLora", "adapter_name": "realism", "scale": 0.8 },
    { "model_name": "user/other-lora", "weight_name": "lora.safetensors", "subfolder": "loras" }
]
  • model_name — the LoRA's hub repo, required. null switches the entry off: it is not loaded and the step runs without it. A template's loras list is fixed JSON, so this is how a caller drops its adapter - pass the variable behind model_name as null (MiniMax H3: lora_model_name). Validation and the run both warn (lora_disabled). A step-distillation (turbo) LoRA sets the step count and sigma shift the step runs at, so raise those to the base model's schedule with it or the step runs undertrained.
  • weight_name / subfolder — pick a specific weights file within the repo.
  • adapter_name — name passed to set_adapters(). Defaults to the LoRA's index in the list.
  • scale — the adapter's weight, passed to set_adapters(). Defaults to 1.0.

See workflows/templates/lora.json for a full example.

IP-Adapter

"ip_adapter": {
    "model_name": "h94/IP-Adapter",
    "weight_name": "ip-adapter_sdxl.bin",
    "scale": 0.6
}

model_name is required; weight_name, subfolder and scale are optional. The adapter image itself is passed as a normal ip_adapter_image pipeline argument. See workflows/templates/ip-adapter.json.

Sharing Components Across Steps

Two pipeline steps that load the same underlying component (a shared text encoder, for instance) can avoid loading it twice:

"configuration": { "component_type": "FluxPipeline", "shared_components": ["text_encoder"] }
"configuration": { "component_type": "FluxPipeline", "reused_components": ["text_encoder"] }

The step naming shared_components stores those components after it loads; a later step naming the same names in reused_components gets them instead of loading its own copy. The names must match exactly between the two steps. Either list can sit in the step's configuration or beside it on the pipeline itself.

How the component reaches the second pipeline depends on what kind it is. A standard pipeline takes it as a from_pretrained argument. A modular pipeline cannot — it is built from the component specs in its own index — so it is registered with update_components() before load_components() runs, which is also what keeps load_components() from pulling a second copy: it only loads what is not already there. That is what lets two MiniMax-H3 steps of different tasks (t2va and ref2va load different transformer partitions) share the 14GB text encoder and the VAEs between them.

A reused component keeps the device placement the step that shared it gave it. Any components entry naming one is skipped with a log line rather than applied a second time — offloading hooks do not survive being installed twice, and the step that loaded the component is the one that decided how it is placed.

Sharing outlives the pipeline that did it: a step can share a component and still set release_pipeline, which frees everything else it loaded while the shared component stays alive for the steps that reuse it.

Attention and Performance

"configuration": {
    "component_type": "FluxPipeline",
    "attention_backend": "flash_hub",
    "enable_attention_slicing": true,
    "no_generator": false
}
  • enable_attention_slicing — process attention in slices to reduce memory at some cost to speed. Automatic on MPS unless disable_attention_slicing is set (UNet/ControlNet models only; worth disabling for SDXL on a Mac).
  • attention_backend — selects a diffusers attention backend (e.g. "flash_hub") for the duration of each pipeline call.
  • prompt_weighting — enables A1111-style prompt weighting ((word:1.5), [word], ((word))) and prompts over 77 tokens. Currently supports Flux pipelines. Mutually exclusive with remote_text_encoder.
  • no_generator — set true to skip creating a torch.Generator for pipelines that don't accept one.
  • inversion — run the pipeline's invert() instead of the pipeline itself; the step returns the inverted/image latents for a later step to consume (see community-pipeline.json).
  • generate — run the pipeline's generate() instead, for components with a generation head (the step returns generated_ids).

Cache Acceleration

A cache block speeds up inference by skipping redundant computation:

"configuration": {
    "cache": { "type": "first_block", "threshold": 0.05 }
}

cache wraps diffusers' own cache hooks - type is one of first_block, faster, mag, taylorseer or text_kv, each with its own tuning fields (threshold, num_inference_steps, max_skip_steps, retention_ratio, cache_interval, max_order, mag_ratios, calibrate — see dw/workflow_schema.json for which fields apply to which type). See workflows/templates/step-caching.json.

Device and Dtype

Device is auto-detected (CUDA > MPS > CPU). Dtype is set per-component:

"from_pretrained_arguments": {
    "model_name": "black-forest-labs/FLUX.1-dev",
    "torch_dtype": "torch.bfloat16"
}

A step can name a device instead, in a pipeline configuration (which becomes the default for that pipeline's components), in a component configuration, or in a task's arguments. A device naming a backend the machine running the workflow does not have is translated to the one it does, with a warning, so a workflow written on a CUDA box runs on a Mac and back again:

"configuration": {
    "component_type": "FluxPipeline",
    "device": "cuda"
}

Only the backend is translated. A device index survives when the backend matches, so cuda:1 on a single-GPU CUDA box remains the error it always was; when the backend does not match, the index is dropped and the warning says so — a workflow that meant to spread work across two accelerators will not on a machine that has one. "device": "cpu" is never translated, since pinning a step to the CPU is how a GPU-specific problem gets ruled out.

Modular Pipelines

Modular pipelines (ModularPipeline and its subclasses) load their configuration and their component weights separately, so from_pretrained_arguments only names the model and load_components pulls the weights:

"configuration": {
    "component_type": "MiniMaxMusic3ModularPipeline",
    "load_components": { "dtype": "torch.bfloat16" },
    "components_manager": { "enable_auto_cpu_offload": true }
}
  • load_components — arguments for load_components(). Use dtype for the component dtype and names to load only some of the components. quantization_config is keyed by component name, since a modular pipeline loads each component itself:

    "load_components": {
        "dtype": "torch.bfloat16",
        "quantization_config": {
            "transformer": {
                "configuration": { "config_type": "TorchAoConfig" },
                "arguments": {
                    "quant_type": "torchao.quantization.Int8WeightOnlyConfig",
                    "modules_to_not_convert": ["proj_in", "proj_out"]
                }
            },
            "language_model": {
                "configuration": { "config_type": "transformers.TorchAoConfig" },
                "arguments": { "quant_type": "torchao.quantization.Int8WeightOnlyConfig" }
            }
        }
    }

    A component the map does not name loads unquantized. Note which TorchAoConfig each component takes: the diffusers one for its own models, the transformers one for a transformers model such as a conditioner.

  • configs — values the pipeline's blocks declare and read while they run. They are neither components nor call arguments, which is why they have a block of their own:

    "configs": {
        "canvas_short_edge": 768,
        "reference_image_short_edge": 1024
    }

    The names are whatever the pipeline itself declares, so they differ per model rather than being a fixed list here — MiniMax-H3 declares canvas_short_edge (768), canvas_max_pixels (1032192) and reference_image_short_edge (2048), the last being the resolution its image references are encoded at. A name the pipeline does not declare raises rather than passing quietly, since a dropped config reads as a setting that did nothing.

  • components_manager — attaches a ComponentsManager, which tracks the pipeline's components. With enable_auto_cpu_offload it keeps only the running components on the device and moves the rest to system memory, reserving memory_reserve_margin (default "3GB") of free device memory. It requires a device that reports free memory (CUDA) and replaces offload, which modular pipelines do not support.

A modular pipeline returns whatever its output argument asks for — one output by name, or several of them together:

"arguments": {
    "prompt": "variable:prompt",
    "output": ["videos", "audio", "sampling_rate"]
}

Asked for several, the outputs come back keyed by name. Video generated with its own soundtrack is muxed into a single video/mp4 file, the same way a video pipeline's own output is, and a later step can still reference any of the outputs by name.

Some repositories hold more than one task's weights. workflow names the task, which prunes the pipeline to the blocks that task runs, so only the components it needs are downloaded and loaded:

"from_pretrained_arguments": {
    "model_name": "MiniMaxAI/MiniMax-H3",
    "workflow": "t2va"
}

A task is chosen by the arguments the step passes, so one workflow name can cover more than one of them: MiniMax-H3's fl2va takes an image, a last_image, or both. Given only a last_image it generates up to that frame, inventing everything that leads to it — see workflows/templates/minimax/last-frame-only.json beside workflows/templates/minimax/first-and-last-frame.json.

See workflows/templates/minimax/music.json and workflows/templates/minimax/video-with-audio.json for full examples.

H3: generating to a held soundtrack

A MiniMax-H3 step (t2va, fl2va or ref2va) takes a pipeline argument hold_audio: an audio value, spelled asset:<file>, output:<file> or previous_result:<step>.audio (an AudioTrack or AudioVideo). The video is generated to that soundtrack. The track goes through H3's Ref2VA reference-audio encode, is written over the target audio rows and held clean as conditioning through denoise, so the model draws picture that fits audio it is not allowed to change:

"pipeline": {
    "arguments": {
        "prompt": "a drummer in a lit rehearsal room",
        "num_frames": 125,
        "hold_audio": "asset:track.wav"
    },
    "output": ["videos", "audio", "sampling_rate"]
}

"audio" must be in output (as it already is in video-with-audio) for the held track to be what the step returns. A track longer than the video is cropped, a shorter one padded with silence. The step's audio is the caller's original waveform fitted to the video's duration (num_frames / 24 s) at its own sample rate, not a VAE round trip. previous_result:base.audio holds a track an earlier step generated, music for example.

run_workflow's arguments can set only a declared variable, so unless the template has one, add "hold_audio": "variable:hold_audio" and a hold_audio variable, or edit the argument. Validation refuses hold_audio on a step that is not H3, on a modular pipeline loaded with a workflow other than t2va, fl2va or ref2va, on a path whose extension is not .wav, .mp3, .flac or .ogg, and on any {"media_type": ...} dict; a run refuses a pipeline with no hold blocks or a value that is not audio. Without hold_audio the output is unchanged.

Hold plus the reference is the music-video default: the template passes each shot's slice as a MiniMaxH3AudioReference and its hold_audio variable defaults to previous_result:slice; set it to null for the reference alone. The match_audio chains (chain-matched-to-audio, chain-matched-and-aligned) keep hold_audio null by default, since a chain has no per-shot slice to name; set it to an audio file to hold. Measured on lem (#808: music-video, 4 shots of 124 frames, two seeds and so two songs, offset of each shot's own audio against its source slice by measure_sync), the reference alone put one of four shots 1.37 s late (seed 42) and another 0.18 s early (seed 7), with correlation 0.35 to 0.96; reference plus hold measured 0.00 s on all eight shots, correlation 0.95 to 1.00. The earlier A/B (#619: chain-matched-to-audio, a 10 s sung track, seed 42, one run per arm) found a held track with image-only references kept the soundtrack exact but the mouth was open at about 3 of 6 sung-word onsets, against 6 of 6 with the reference alone, and read as speech to camera rather than singing. #808 measured the audio's offset, not mouth onsets, so how closely the mouth follows a held-plus-referenced song is still a visual check. Hold is also for picture that must fit audio exactly, such as motion cut to music. refine_strength also uses hold, to keep the base pass's own audio, which was generated jointly with that video - a different case from lip sync to supplied audio (see the refine section below).

H3: holding a clip with guides

A MiniMax-H3 t2va or fl2va step takes a pipeline argument guides: a list of {"video": <reference>, "frame": <pixel frame>}, where the video is an asset:, output: or previous_result: reference. Each guide clip is VAE-encoded on the generation canvas and placed as condition rows, timed from the target frame it lands on, that are never denoised, so the generated video is held to earlier footage. On fl2va the image / last_image keyframes combine with guides. The clip is fitted to the canvas (scaled to cover, centre-cropped). By default a guide holds video only and its own audio is not used; "audio": true holds the guide's audio too (below).

"pipeline": {
    "arguments": {
        "prompt": "variable:prompt",
        "num_frames": "variable:num_frames",
        "width": "variable:width",
        "height": "variable:height",
        "num_inference_steps": "variable:num_inference_steps",
        "guides": [
            {"video": "asset:opening.mp4", "frame": 0},
            {"video": "asset:closing.mp4", "frame": 85}
        ],
        "output": ["videos", "audio", "sampling_rate"]
    }
}

This is the text_to_video_audio step of workflows/templates/minimax/video-with-audio.json with guides added. As with hold_audio, run_workflow's arguments can set only a declared variable, so reference a guides variable or edit the argument.

Rule Value
frame a multiple of 17 (a VAE chunk boundary): 0, 17, 34, ...
clip length 1, 5 or 17m + 5 frames (22, 39, 56, ... 124); any other length is cut down to the longest such length (n < 5 gives 1, 5 <= n < 22 gives 5, else 17 * ((n - 5) // 17) + 5) with a warning
extent frame + length must not run past num_frames (H3 rounds it up to 17n + 5; default 124); ending exactly at the end is fine
count at most 4 guides per step (a VRAM limit: each guide frame adds attention rows, and the template's vram_estimate charges each guide's frames on the canvas); guides: [] is the same as none
where t2va or fl2va on an H3 pipeline; not with references (ref2va), not on a non-H3 pipeline

Validation refuses all of the above before the run (dw/guides.py); a previous_result: guide is checked at run time, when its clip exists.

Holding the guide's audio: "audio": true

A guide with "audio": true holds its soundtrack as well as its frames: {"video": "asset:opening.mp4", "frame": 0, "audio": true}. The guide video's audio over the guide's span is encoded by the audio VAE (posterior mode, as ref2va reference audio is) and placed as condition audio rows timed at the target's own audio time, then held, never denoised. The output's audio under the guide reproduces the guide's audio, and the model generates the rest to run on from it. Audio runs at 40 latents a second against 24 fps video, so a guide of n frames at frame k holds round((k + n) * 40 / 24) - floor(k * 40 / 24) audio latents: 37 for a 22-frame guide at frame 0, 65 (about 1.6 s) for a 39-frame one. A guide that does not start at frame 0 has its audio padded to the audio VAE's hop first, so its latents land at the right time. The clip is read as 24 fps footage, as its frames are.

Rule Value
audio true or false; absent is false. Any other value is refused
soundtrack "audio": true on a guide whose video has no audio stream is refused - before the run for an asset:/output: file, at run time for a previous_result: clip
with hold_audio allowed; a guide's audio rows come first, then the held track's

A guide holds its span's look as well as its composition and motion: the frames under it come out close to a copy of the clip, whatever the prompt says about style. A full-length guide at frame 0 therefore reproduces the take rather than restyling it - measured on lem, a 124-frame guide with the prompt "the same scene as a hand-painted watercolour animation" came back photoreal (#648). The prompt steers the frames the guides leave free. Restyling a whole take needs a LoRA trained on guides (#612), not the guide alone.

With no guides (or guides: []) the step runs the stock layout and draws the same noise from its seed as before. Repeat runs at one seed are bit-identical only on a server with cudnn_deterministic: true (ACCELERATION.md, TF32 and cuDNN); with the default false, two runs at seed 42 match frame for frame by eye, but their audio levels differ by a few dB.

Chaining with a guide: continuity: "guide"

A chain on an H3 t2va or fl2va step can carry its seam with a guide instead of a keyframe: "chain": {"segments": 3, "continuity": "guide", "guide_frames": 22}. Every segment after the first gets the previous segment's last P = guide_frames frames as a guide at frame 0 (appended to the step's own guides; the carried clip counts as a guide, so a step that also writes its own guides may list at most three of them with continuity: "guide" - validation says so before the run rather than after the first segment), with "audio": true when carry_audio is (the default), so motion and voice run on across the seam rather than restarting from a still. On fl2va, image is set to the guide's first frame. The next segment opens with a near-copy of those P frames, so P frames are trimmed from its head: N segments of F frames give F + (N - 1)(F - P) frames - 3 x 124 at P = 22 is 328.

Rule Value
guide_frames 22 (default) or 39 - a whole-latent guide length; anything else is refused
where MiniMax-H3 t2va or fl2va; refused on ref2va, with references, or on a non-H3 step - guides stay off ref2va
carry_frames refused - the carry is always the last guide_frames frames
trim_frames ignored (the trim is P), with a note in the job log
crossfade_ms ignored when carry_audio is true: the guide held the audio, so the seam is joined with only a declick
num_frames must exceed P, or the chain cannot progress

workflows/templates/minimax/chained-segments.json takes continuity and guide_frames as variables. The ref2va chains (chain-video-continuity and the match_audio chains) keep last_segment.

Promoting an H3 take to 768p in latent space: upscale_h3_latents and decode_h3_latents

Once a 960x544 MiniMax-H3 take reads the way it should, upscale_h3_latents and decode_h3_latents promote it to 1344x768 without denoising it again - cheaper than a native 768p render, since only a small 3D-convolution network and a VAE decode run, not the transformer. These two tasks run no refine pass, so the result is sharper than the 544p take but cannot show detail the base pass never generated; the refine section below adds one. Measured on a 124-frame crowd scene, it took 7.8 min against a native 768p render's 12.7 min, and the faces came out soft and waxy where the native render's were distinct. So this is a measurement path, not a catalog template: the catalog's native 768p render is templates/minimax/video-with-audio-768p, and it is the one to use when faces matter.

Target width/height must be multiples of 16, each between 1x and 4x the base latents' own size, and within H3's 1344x768 (or portrait 768x1344) canvas. The latents pass from base to up to decode entirely in memory. A previous_result: property on a modular step names a key of the dict its output list returns, spelled as the pipeline spells it: base.latents, base.audio and base.sampling_rate - not sample_rate, which is what the saved AudioVideo calls it, and which the dict does not carry. A property no result carries is an error rather than an empty list, so a misspelt one fails the step instead of skipping the mux that reads it. The base step carries no result at all, so nothing beyond its return value is written - upscaling only makes sense for a take chosen from something already reviewed, so the 544p pass that produced it is not itself a deliverable here.

The upscaler weights (~691 MB, LBH-123-AI/Minimax_h3_latent_Upscaler, Apache-2.0, read at a pinned revision) and the H3 VAE download on first use, the same as any other model. Neither is counted in plan.downloads_required, since that walk collects from_pretrained_arguments sources on pipeline steps and does not see a task argument naming a Hugging Face repo - a box that has run the base pass before but never this task can still stall mid-run pulling the upscaler.

{
    "id": "H3LatentUpscalePreview",
    "description": "Promote a MiniMax-H3 take from 960x544 to 1344x768 in latent space, decoded without refinement.",
    "variables": {
        "prompt": "prompt:minimax/fox_dawn_context_ir",
        "num_frames": 124,
        "num_inference_steps": 9,
        "video_shift": 12.0,
        "audio_shift": 3.0,
        "weights_dtype": "{int4}",
        "lora_scale": 1.0,
        "lora_alpha": null,
        "lora_model_name": "lightx2v/Minimax-h3-Turbo",
        "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors",
        "lora_adapter_name": "turbo",
        "seed": 42
    },
    "variable_constraints": {
        "num_frames": {
            "modulus": 17,
            "remainder": 5,
            "min_frames": 124,
            "max_frames": 345,
            "snap": "up",
            "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps"
        }
    },
    "seed": "variable:seed",
    "steps": [
        {
            "name": "base",
            "pipeline": {
                "configuration": {
                    "component_type": "ModularPipeline",
                    "pre_load_modules": [
                        "sdnq"
                    ],
                    "load_components": {
                        "dtype": "torch.bfloat16",
                        "quantization_config": {
                            "transformer": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "variable:weights_dtype",
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "use_quantized_matmul": true,
                                    "dequantize_fp32": false,
                                    "modules_to_not_convert": [
                                        "proj_in",
                                        "audio_proj_in",
                                        "context_embedder",
                                        "time_embedder",
                                        "time_proj",
                                        "token_refiner",
                                        "norm_out",
                                        "proj_out",
                                        "audio_proj_out"
                                    ]
                                }
                            },
                            "text_encoder": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "variable:weights_dtype",
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "dequantize_fp32": false,
                                    "modules_to_not_convert": [
                                        ".model.visual",
                                        "lm_head"
                                    ]
                                }
                            },
                            "vae": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "{int8}",
                                    "quant_conv": true,
                                    "use_quantized_matmul_conv": true,
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "dequantize_fp32": false
                                }
                            }
                        }
                    },
                    "components": {
                        "transformer": {
                            "group_offload": {
                                "offload_type": "block_level",
                                "num_blocks_per_group": 2,
                                "use_stream": true,
                                "record_stream": true,
                                "low_cpu_mem_usage": true
                            }
                        },
                        "text_encoder": {
                            "remove_modules": [
                                "lm_head"
                            ]
                        },
                        "text_encoder.model": {
                            "truncate_layers": {
                                "language_model.layers": 51
                            },
                            "group_offload": {
                                "offload_type": "leaf_level"
                            }
                        },
                        "vae": {
                            "device": "cuda",
                            "residency": "on_demand"
                        },
                        "audio_vae": {
                            "device": "cuda",
                            "residency": "on_demand"
                        }
                    }
                },
                "from_pretrained_arguments": {
                    "model_name": "MiniMaxAI/MiniMax-H3",
                    "workflow": "t2va"
                },
                "loras": [
                    {
                        "model_name": "variable:lora_model_name",
                        "weight_name": "variable:lora_weight_name",
                        "adapter_name": "variable:lora_adapter_name",
                        "scale": "variable:lora_scale",
                        "alpha": "variable:lora_alpha"
                    }
                ],
                "scheduler": {
                    "shift": "variable:video_shift"
                },
                "audio_scheduler": {
                    "shift": "variable:audio_shift"
                },
                "arguments": {
                    "prompt": "variable:prompt",
                    "num_frames": "variable:num_frames",
                    "width": 960,
                    "height": 544,
                    "num_inference_steps": "variable:num_inference_steps",
                    "output": [
                        "videos",
                        "audio",
                        "sampling_rate",
                        "latents"
                    ]
                }
            }
        },
        {
            "name": "up",
            "task": {
                "command": "upscale_h3_latents",
                "arguments": {
                    "latents": "previous_result:base.latents",
                    "width": 1344,
                    "height": 768
                }
            }
        },
        {
            "name": "decode",
            "task": {
                "command": "decode_h3_latents",
                "arguments": {
                    "latents": "previous_result:up"
                }
            }
        },
        {
            "name": "mux",
            "task": {
                "command": "pair_audio",
                "arguments": {
                    "video": "previous_result:decode",
                    "audio": "previous_result:base.audio",
                    "sample_rate": "previous_result:base.sampling_rate",
                    "fit": "video"
                }
            },
            "result": {
                "content_type": "video/mp4",
                "subfolder": "final"
            }
        }
    ]
}

To keep the 544p take beside the promotion for comparison, add one more pair_audio step with "video": "previous_result:base.videos" and the same audio/sample_rate, saved with "result": {"content_type": "video/mp4", "fps": 24, "subfolder": "intermediate"}. base.videos is the pipeline's batch - a list holding the one video - and pair_audio unwraps a batch of one (a batch of several is refused, since one track goes under one video). The dict's frames carry no frame rate, so the result's fps says it; without it the file is written at the 8 fps fallback.

Refining the upscaled latents: refine_strength

upscale_h3_latents followed by decode_h3_latents adds no detail. To let the transformer add some, pass the upscaled latents to an H3 step as latents with refine_strength. The step adds noise to the video latents up to sigma = refine_strength and denoises from there down to 0, so the picture is kept and the fine detail is regenerated at 1344x768. The refined video is decoded by the step itself, so there is no decode_h3_latents step, and the step's videos is the deliverable.

Refine re-denoises the video only. The audio rows have to be held, so pass the base pass's track as hold_audio (previous_result:base.audio) and the step's audio is that track. This is a different use of hold from the one above: it keeps the audio the base pass generated jointly with that video, rather than fitting picture to supplied audio. For lip sync to supplied audio see the hold section above (#619, #795).

{
    "id": "H3LatentUpscaleRefine",
    "description": "Promote a MiniMax-H3 take from 960x544 to 1344x768 in latent space, then refine it at low strength.",
    "variables": {
        "prompt": "prompt:minimax/fox_dawn_context_ir",
        "num_frames": 124,
        "num_inference_steps": 9,
        "video_shift": 12.0,
        "audio_shift": 3.0,
        "weights_dtype": "{int4}",
        "lora_scale": 1.0,
        "lora_alpha": null,
        "lora_model_name": "lightx2v/Minimax-h3-Turbo",
        "lora_weight_name": "minimax_h3_fl2v_turbo_8step_v1.0_bf16.safetensors",
        "lora_adapter_name": "turbo",
        "seed": 42
    },
    "variable_constraints": {
        "num_frames": {
            "modulus": 17,
            "remainder": 5,
            "min_frames": 124,
            "max_frames": 345,
            "snap": "up",
            "reason": "the video VAE encodes 17 * n + 5 frames, and MiniMax-H3 generates between 5 and 15 seconds at 24 fps"
        }
    },
    "seed": "variable:seed",
    "steps": [
        {
            "name": "base",
            "pipeline": {
                "configuration": {
                    "component_type": "ModularPipeline",
                    "pre_load_modules": [
                        "sdnq"
                    ],
                    "load_components": {
                        "dtype": "torch.bfloat16",
                        "quantization_config": {
                            "transformer": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "variable:weights_dtype",
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "use_quantized_matmul": true,
                                    "dequantize_fp32": false,
                                    "modules_to_not_convert": [
                                        "proj_in",
                                        "audio_proj_in",
                                        "context_embedder",
                                        "time_embedder",
                                        "time_proj",
                                        "token_refiner",
                                        "norm_out",
                                        "proj_out",
                                        "audio_proj_out"
                                    ]
                                }
                            },
                            "text_encoder": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "variable:weights_dtype",
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "dequantize_fp32": false,
                                    "modules_to_not_convert": [
                                        ".model.visual",
                                        "lm_head"
                                    ]
                                }
                            },
                            "vae": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "{int8}",
                                    "quant_conv": true,
                                    "use_quantized_matmul_conv": true,
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "dequantize_fp32": false
                                }
                            }
                        }
                    },
                    "components": {
                        "transformer": {
                            "group_offload": {
                                "offload_type": "block_level",
                                "num_blocks_per_group": 2,
                                "use_stream": true,
                                "record_stream": true,
                                "low_cpu_mem_usage": true
                            }
                        },
                        "text_encoder": {
                            "remove_modules": [
                                "lm_head"
                            ]
                        },
                        "text_encoder.model": {
                            "truncate_layers": {
                                "language_model.layers": 51
                            },
                            "group_offload": {
                                "offload_type": "leaf_level"
                            }
                        },
                        "vae": {
                            "device": "cuda",
                            "residency": "on_demand"
                        },
                        "audio_vae": {
                            "device": "cuda",
                            "residency": "on_demand"
                        }
                    }
                },
                "from_pretrained_arguments": {
                    "model_name": "MiniMaxAI/MiniMax-H3",
                    "workflow": "t2va"
                },
                "loras": [
                    {
                        "model_name": "variable:lora_model_name",
                        "weight_name": "variable:lora_weight_name",
                        "adapter_name": "variable:lora_adapter_name",
                        "scale": "variable:lora_scale",
                        "alpha": "variable:lora_alpha"
                    }
                ],
                "scheduler": {
                    "shift": "variable:video_shift"
                },
                "audio_scheduler": {
                    "shift": "variable:audio_shift"
                },
                "arguments": {
                    "prompt": "variable:prompt",
                    "num_frames": "variable:num_frames",
                    "width": 960,
                    "height": 544,
                    "num_inference_steps": "variable:num_inference_steps",
                    "output": [
                        "videos",
                        "audio",
                        "sampling_rate",
                        "latents"
                    ]
                }
            }
        },
        {
            "name": "up",
            "task": {
                "command": "upscale_h3_latents",
                "arguments": {
                    "latents": "previous_result:base.latents",
                    "width": 1344,
                    "height": 768
                }
            }
        },
        {
            "name": "refine",
            "pipeline": {
                "configuration": {
                    "component_type": "ModularPipeline",
                    "pre_load_modules": [
                        "sdnq"
                    ],
                    "load_components": {
                        "dtype": "torch.bfloat16",
                        "quantization_config": {
                            "transformer": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "variable:weights_dtype",
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "use_quantized_matmul": true,
                                    "dequantize_fp32": false,
                                    "modules_to_not_convert": [
                                        "proj_in",
                                        "audio_proj_in",
                                        "context_embedder",
                                        "time_embedder",
                                        "time_proj",
                                        "token_refiner",
                                        "norm_out",
                                        "proj_out",
                                        "audio_proj_out"
                                    ]
                                }
                            },
                            "text_encoder": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "variable:weights_dtype",
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "dequantize_fp32": false,
                                    "modules_to_not_convert": [
                                        ".model.visual",
                                        "lm_head"
                                    ]
                                }
                            },
                            "vae": {
                                "configuration": {
                                    "config_type": "sdnq.SDNQConfig"
                                },
                                "arguments": {
                                    "weights_dtype": "{int8}",
                                    "quant_conv": true,
                                    "use_quantized_matmul_conv": true,
                                    "quantization_device": "cuda",
                                    "return_device": "cpu",
                                    "dequantize_fp32": false
                                }
                            }
                        }
                    },
                    "components": {
                        "transformer": {
                            "group_offload": {
                                "offload_type": "block_level",
                                "num_blocks_per_group": 2,
                                "use_stream": true,
                                "record_stream": true,
                                "low_cpu_mem_usage": true
                            }
                        },
                        "text_encoder": {
                            "remove_modules": [
                                "lm_head"
                            ]
                        },
                        "text_encoder.model": {
                            "truncate_layers": {
                                "language_model.layers": 51
                            },
                            "group_offload": {
                                "offload_type": "leaf_level"
                            }
                        },
                        "vae": {
                            "device": "cuda",
                            "residency": "on_demand"
                        },
                        "audio_vae": {
                            "device": "cuda",
                            "residency": "on_demand"
                        }
                    }
                },
                "from_pretrained_arguments": {
                    "model_name": "MiniMaxAI/MiniMax-H3",
                    "workflow": "t2va"
                },
                "loras": [
                    {
                        "model_name": "variable:lora_model_name",
                        "weight_name": "variable:lora_weight_name",
                        "adapter_name": "variable:lora_adapter_name",
                        "scale": "variable:lora_scale",
                        "alpha": "variable:lora_alpha"
                    }
                ],
                "scheduler": {
                    "shift": "variable:video_shift"
                },
                "audio_scheduler": {
                    "shift": "variable:audio_shift"
                },
                "arguments": {
                    "prompt": "variable:prompt",
                    "num_frames": "variable:num_frames",
                    "width": 1344,
                    "height": 768,
                    "num_inference_steps": 5,
                    "output": [
                        "videos",
                        "audio",
                        "sampling_rate"
                    ],
                    "latents": "previous_result:up",
                    "refine_strength": 0.2,
                    "hold_audio": "previous_result:base.audio"
                }
            },
            "result": {
                "content_type": "video/mp4",
                "fps": 24,
                "subfolder": "final"
            }
        }
    ]
}

num_inference_steps is the number of sigma points, ending at 0, spaced by the scheduler's shift, so a refine runs num_inference_steps - 1 denoise evaluations. Here 5 points at shift 12 (the example's video_shift) and strength 0.2 are sigma 0.2, 0.157, 0.110, 0.058, 0: 4 evaluations. The time a refine takes scales with num_inference_steps. refine_strength sets where it starts, not how many steps run: a higher strength moves the picture further from the upscaled take, a lower one stays closer to it.

Validation refuses a refine_strength that is not a number between 0 and 1 (exclusive), a refine with no latents or no hold_audio, a step that is not H3, and a num_inference_steps below 2.

Chained Video Generation

Video pipelines generate short clips - a chain block on a pipeline step runs the pipeline once per segment and stitches the segments into one long video. The model loads once; each segment's last frame is carried into the next segment as its keyframe, the duplicated boundary frames are trimmed, and frames and audio are joined into a single file:

"pipeline": {
    "configuration": { "component_type": "LTX2ImageToVideoPipeline" },
    "from_pretrained_arguments": { "model_name": "Lightricks/LTX-2.5-Diffusers" },
    "chain": {
        "segments": 3,
        "trim_frames": 2,
        "crossfade_ms": 80
    },
    "arguments": { "prompt": "variable:prompt", "image": "variable:image" }
}
  • segments — how many times the pipeline runs. Total length is roughly segments * num_frames, minus trim_frames per seam.
  • match_audio — instead of a count, derive the length from the step's hold_audio track, or else the one audio reference in its references. The audio is sliced into frame-aligned per-segment chunks, each segment is generated against its slice - a held track's slice is that segment's hold_audio, a reference's replaces the reference - and the final video is muxed with the original, unsliced track - so the soundtrack has no seams at all. Requires num_frames (the per-segment length) and a frame rate. Exactly one of segments or match_audio must be given.
  • continuity — how continuity carries across segments. last_frame (the default) extracts each segment's last frame and passes it to the next segment - single-frame conditioning, which carries pose and colour. last_segment carries the previous segment itself (frames and its generated soundtrack) into the next as a video reference, which also carries motion, camera, and voice across the seam; it requires a segment_argument that takes a references list. guide (MiniMax-H3 t2va/fl2va only) lays the previous segment's last guide_frames frames into the next as a held frame-0 guide (see H3: holding a clip with guides, Chaining with a guide).
  • guide_frames — with guide, how many frames the guide carries: 22 (default) or 39.
  • carry_frames — with last_segment, bound the carry to the last N frames of the segment (the audio is cut to the same span). Unset carries the whole segment.
  • carry_audio — with last_segment or guide, whether the carried reference includes its soundtrack (default true).
  • segment_argument — where the carried frame or reference lands: image (default) for image-to-video pipelines, or references for reference-conditioned modular pipelines, where it is appended alongside the workflow's own.
  • trim_frames — image-to-video pipelines reproduce their keyframe as frame 0, so this many frames are dropped from the head of every segment after the first (default 1). The matching audio is used as crossfade material, so video and audio stay exactly in sync. It also bounds the crossfade window: trim_frames / fps seconds (at 24 fps, trim_frames: 2 allows the full default 75 ms fade).
  • crossfade_ms — equal-power crossfade applied to generated audio at each seam (default 75). Not used with match_audio, which keeps the original track.
  • fps — frame rate for the chain's audio math. Defaults to the pipeline's frame_rate argument; pipelines with a fixed rate need it set (MiniMax H3: 24).
  • frame_snap — the constraint the pipeline puts on num_frames, used to snap the final match_audio segment to a valid length. MiniMax H3 accepts 17n+5 frames between 124 and 345: { "modulus": 17, "remainder": 5, "min_frames": 124, "max_frames": 345 }. Where the workflow already declares that rule as a variable_constraints entry, write "frame_snap": "constraint:num_frames" instead, so the numbers live in one place (What a variable is allowed to be).
  • prompts — optional per-segment prompt list for narrative progression; segment i uses prompts[min(i, len - 1)].
  • save_segments — write each completed segment to the output directory as a playable mp4 and free its frames, bounding memory to roughly one segment regardless of chain length. The final video is streamed from the segment files at save time, and they are removed once it is written (keep_segments: true retains them). A crashed chain leaves the finished segments behind - stitch them by hand by listing their paths in a concat_videos step (trim_frames: 0, the trim was already applied). Requires PyAV and a frame rate. The trade-off is one extra encode/decode cycle through h264 for the segment files.

The chain runs inside one iteration of the step, so it composes with previous_result fan-out (three keyframes in, three chained videos out), and a pipeline_reference step can carry its own chain. Seeds behave like a normal run: the step's generator advances across segments, so one seed reproduces the whole chain. Expect some visual drift across many segments with last_frame continuity - it is single-frame conditioning; last_segment continuity exists for exactly that, where the pipeline can take a video reference.

See workflows/templates/ltx2/chained-segments.json, workflows/templates/minimax/chained-segments.json, and workflows/templates/minimax/chain-matched-to-audio.json (audio-matched lip-sync of arbitrary length).

Schedulers

Override the default scheduler:

"scheduler": {
    "configuration": {
        "scheduler_type": "DPMSolverMultistepScheduler"
    },
    "from_config_args": {
        "use_karras_sigmas": true
    }
}

A scheduler block may also carry shift, the exponential sigma shift for schedulers that take one (MiniMax H3's released checkpoint: 12.0 for video, 3.0 for audio). A pipeline that carries a second scheduler takes an audio_scheduler block with the same shape - MiniMax H3 steps video and audio latents down two schedules whose shifts are set independently.

Seeds

Set a seed for reproducibility at workflow, step, or pipeline level - most specific wins: a pipeline's own seed overrides its step's, which overrides the workflow's:

{
    "id": "my_workflow",
    "seed": 42,
    "steps": [
        { "name": "step1", "seed": 123, "pipeline": { "seed": 7, ... } }
    ]
}

Omit seed entirely to let the workflow draw a random one at run time. The seed a run actually used - drawn or named - is recorded in its manifest.json, so a run you liked can be reproduced after the fact.

Beside that manifest the run also writes workflow.json — the realized workflow, meaning the one that actually ran. Every mutable input is pinned into it: the caller's arguments folded into the variables defaults, the seed the run used, each prompt: reference replaced by the stored text, and each output:<identity>/latest/<file> (or /v<N>/) rewritten to the run id it resolved to. asset:, constant:, previous_result: and builtin: are kept as written — each already names something pinned by the asset library or by the manifest's dw_version — and a sub-workflow named by local path is kept with its file's SHA-256 recorded in the manifest. The manifest also lists which stored prompts were inlined, since inlining loses the name.

A step that joins shots (concat_videos, dissolve_videos, or a pipeline step with a chain) also records where each one landed, as shots on its manifest entry (and on its step_end event):

{
  "step": "cut",
  "files": ["final/film.mp4"],
  "subfolder": "final",
  "shots": [
    {"name": "shot@a", "start_frame": 0,   "num_frames": 121, "start_sample": 0,      "num_samples": 242267},
    {"name": "shot@b", "start_frame": 121, "num_frames": 97,  "start_sample": 242267, "num_samples": 194000}
  ]
}

The shots partition the file's frames: the num_frames add up to the frame count. The sample fields are measured off the track the join built, not worked out from the frames. That means a shot whose track ran long shows it here: the first shot above is 267 samples longer than 121 frames at 24 fps. They are null when the video has no track, and for a chain that uses match_audio. A shot is named shot@<key> when the step's videos entry was a previous_result:shot@<key> reference, else by its input's position (video N, a chain's segment N). A dissolve's shots after the first carry overlap_frames, the head they share with the shot before. A step that wrote several joined files marks each shot with its file.

The steps that keep the frames pass shots on. stabilize and the per-frame tasks keep them as they are. interpolate_frames rescales them to the new frame count and clears the samples. pair_audio measures the samples again against the new track - every shot but the last is round(start_frame / fps * sample_rate), and the last one runs to the track's actual end - measured again, once the file is written, against what it decodes to. So its num_samples can be a few dozen samples off round(num_frames * sample_rate / fps): the AAC encode's trim, recorded in the job's event log (not a dropped sample - a real mismatch between the track and the video's length is its own warning, audio_video_length_mismatch or audio_padded_to_video/audio_trimmed_to_video with fit: "video"). Everything else drops them: an audio task's track, say, or a video read back from a file. get_gallery_metadata reports the recorded shots as media.shots, and get_output_frames(seams=true) uses them when you pass no boundaries.

The file is a valid workflow, and running it again is python -m dw.run workflow.json or handing its contents to run_workflow as inline_workflow — but either way the asset: and output: names in it resolve against the server's or CLI's own libraries, not against the run directory, so doing this from inside that directory reproduces the run only when its libraries are the ones the original run used too. Writing the file is best effort, exactly like the manifest — a run that produced its files has succeeded either way — and --output-layout flat writes no run directory, so it writes neither file.

Any of the three levels accepts a variable: reference, which is how a seed becomes settable per run without editing the file:

{
    "variables": { "seed": 42 },
    "seed": "variable:seed",
    "steps": [ ... ]
}
python -m dw.run workflows/models/z-image.json seed=1234

Declare the variable with an integer default, as above: the value from the command line arrives as a string and is converted to the declared type. A string that is not a variable: reference is rejected by the schema.

The seed also reaches sub-workflows: a delegated workflow step runs the child under the parent's seed unless the child names its own. Without that a child draws its own random seed, and a workflow whose real generation happens inside a sub-workflow would not reproduce from the seed it was given.

The step cache that lets a reproduced step skip re-running (see Runs in Workspaces) is scoped to the output directory a run writes into, which on dw.serve is the pinned workspace's own outputs/. Two workspaces holding what looks like the same prior run - same workflow, same seed, same arguments - do not share a cache entry, so validate_workflow's plan.cached_steps answers for the workspace the call is pinned to, not for every workspace that happens to hold a matching run. Deleting a workspace takes its cache entries with it, the same as deleting its outputs/ directory would.

Type System

Dynamic type conversion applies to certain values:

  • Keys ending in _type or _dtype, or named dtype: "torch.bfloat16" becomes torch.bfloat16
  • Dotted names: "sdnq.SDNQConfig" loads the class via importlib
  • Escape with braces to keep as string: "{nf4}" stays as "nf4"
  • content_type and offload_type are exempt even though they end in _type - they name a category, not a Python type, so their value always stays a plain string (the {} escape is accepted but not required for these two keys)
  • Values prefixed with constant: are read from python rather than copied into the workflow: "constant:diffusers.pipelines.ltx2.utils.DISTILLED_SIGMA_VALUES"

Constant References

Some arguments have a value the library already declares: the sigma schedule a distilled model was trained on, the negative prompt a model family ships. Reference it with constant: and its dotted python name instead of copying it into the workflow:

"sigmas": "constant:diffusers.pipelines.ltx2.utils.DISTILLED_SIGMA_VALUES",
"negative_prompt": "constant:diffusers.pipelines.ltx2.utils.DEFAULT_NEGATIVE_PROMPT"

The leading part of the name that imports is the module, and the rest is read from it - so a constant held in a config object is reachable too:

"prompt_max_new_tokens": "constant:diffusers.pipelines.ltx2.utils.GEMMA4_PROMPT_ENHANCEMENT_CONFIG.max_new_tokens"

A reference resolves anywhere in a workflow's arguments, including in a variables default, where it becomes the value a caller overrides - and its type, since a variable is declared by its default:

"variables": { "negative_prompt": "constant:diffusers.pipelines.ltx2.utils.DEFAULT_NEGATIVE_PROMPT" }

A constant is data. The name has to resolve to a value - anything callable is refused, because a type is named with a *_type argument and constructed there, and reaching a function this way would be evaluating python rather than referencing it. Mutable values are copied, so a pipeline that consumes its schedule in place cannot edit the library's constant for the rest of the session.

The value the library declares is the value the workflow gets, which is the point: a constant that changes upstream changes here, and one that is renamed or moved fails loudly rather than leaving a stale copy behind.

Prompt References

A prompt worth keeping is worth keeping once. Stored prompts live as JSON files in a prompt library - the prompts/ folder by default - and a workflow argument written as prompt: plus the file's name (without .json, optionally one folder deep) loads its text at run time:

"prompt": "prompt:scenic_landscape",
"prompt": "prompt:minimax/fox_dawn_t2va"

A prompt file holds the text plus the metadata the server's Prompts page shows:

{
  "text": "A sweeping alpine valley at golden hour...",
  "description": "General-purpose scenic landscape",
  "intended_model": "z-image",
  "negative_prompt": "blurry, low quality",
  "tags": ["landscape", "golden-hour"]
}

Only text is required, and it is what the reference resolves to. intended_model is informational - the engine ignores it, but the library badges and filters by it, and the server's prompt enhancer uses it to preselect a preset. One spelling per family: list_prompts(intended_model=...) matches the whole value exactly, so minimax-music beside minimax-music3 hides half a shelf, and tests/test_prompt_library.py sweeps the repo's library for a variant.

The library's location is resolved in order: the DW_PROMPT_DIR environment variable (which --prompt-dir on dw.serve sets), then ./prompts in the working directory when it exists, then the first prompts/ folder found walking up from the workflow file's own directory - which is how a repo workflow run from any working directory still reaches the library beside it. dw.serve resolves the directory once at startup with this same order (anchored at its workflow directory) and pins it for every job, so the Prompts page and prompt: resolution always agree on one library. References are rooted at that one directory - not at the workflow file - so the same reference means the same text from every workflow. Like constant:, a reference resolves anywhere in a workflow's arguments, including a variables default, and it always resolves to exactly one string: it never multiplies a step's iterations the way previous_result: references do. A prompt's text may not itself begin with a reference prefix such as variable: - the engine refuses it rather than resolving text as syntax.

Asset References

A workflow's plain media paths resolve against the workflow file's own directory, which means a workflow that reads anything has to keep that thing beside it. An asset: reference is rooted at the asset library instead - the workspace's assets/ folder - so a workflow and the media it reads do not have to live in the same place:

"image": "asset:iris.png",
"video": "asset:gyre/frames/web.mp4",
"references": [
    {
        "reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3ImageReference",
        "from_file": "asset:subject.png"
    }
]

A reference names a file with its extension, at most four folders deep, and resolves to that file's path - so it works under any argument that accepts a path: image, video, a from_file, a list of any of them, or a task argument that names a file. What loads the path is unchanged; only where the path comes from is.

The library's location is resolved in order: the DW_ASSET_DIR environment variable (which --asset-dir on dw.serve sets), then the workspace's assets/ when a workspace was named explicitly, then ./assets in the working directory when it exists, then the first assets/ folder found walking up from the workflow file's own directory. See Workspaces.

A reference can only name a file inside the library: .., an absolute path, or a symlink pointing out of it are all refused. Browser uploads land in the library's uploads/ folder and come back as asset:uploads/<name>, so a workflow saved after an upload still resolves on the next run.

Output References

Multi-stage work — generate stills, then animate them; generate a score, then mux it — used to mean copying files out of the output directory and back in beside the next workflow. An output: reference names what an earlier run wrote, directly:

"image": "output:ltx2/Gyre/latest/Gyre-still.0-0.0.png",
"audio": "output:ltx2/GyreScore/20260905-181530-a1b2c3d4/Gyre-score.10-0.0.wav"

The name is a path under the output directory — the workflow's identity, the run, and the file (see Runs). Writing latest where the run id goes resolves to the newest run of that workflow that holds the file, which is what lets a second-stage workflow name the first stage's product without being edited after every run - and keeps working when the newest run failed part way, or reused every step from the cache and so wrote nothing of its own but a manifest. Runs sort by their id, which starts with a UTC timestamp, so "newest" needs no file timestamps and survives a directory being copied. v<N> in the same position names the run whose version is N - the v4 the gallery labels its files with - so the number a person was told is a name a workflow can take. Unlike latest it picks exactly one run: v4 not holding the file is an error, not a reason to try v3. latest and v<N> only select a run where run directories are; a workflow or file that happens to be called either is still named as itself.

Like asset:, a reference resolves to a path and then whatever loads paths loads it, so it works under image, video, a from_file, or a list of them. The audio tasks take a video file's path too and use the soundtrack muxed into it, which is how a finished cut is scored in a later run without re-cutting it. It resolves against the output directory the run was told to write to, and cannot leave it: .., an absolute path, and a symlink pointing out are all refused.

To name an earlier step of the same run, use previous_result: instead — that passes the value in memory rather than through the filesystem.

A generated file worth reusing repeatedly is better kept than referenced by the run that made it: POST /api/assets/keep (the gallery's Keep as asset, or MCP's keep_output) copies it into the workspace's asset library under a name you choose, and from then on it is an asset: reference like any other — stable whatever happens to the run directory it came from.

Objects Built From a File

Some pipelines take arguments that are objects rather than plain media. An argument that names a type and a from_file is constructed by that type's own from_file():

"references": [
    {
        "reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3ImageReference",
        "from_file": "subject.png"
    },
    {
        "reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3AudioReference",
        "from_file": "voice.wav"
    }
]

Loading the media this way rather than as a plain image or video argument is what brings its frame rate or sample rate along with it, which MiniMax-H3 resamples a reference from. The file may be a path — relative to the workflow file, like all media a workflow names — or a URL, and is validated like any other media. variable: references work as the file location; previous_result: does not, since the object is built when the workflow loads — use from_previous_result for that. A dict that merely contains a from_file key without a *_type key is not an object description and is passed through untouched.

An entry in a list whose source is null is left out of that list. That is what makes a reference optional: write it as an ordinary entry whose from_file (or from_previous_result) is a variable, declare the variable null, and a run that is given nothing for it generates exactly as it did before the reference existed — one workflow serving both, instead of two spellings of the same steps. It applies to from_file, from_previous_result and from_arguments alike. On its own rather than in a list there is nothing to leave it out of, so a null source there is an error.

Any other key goes wherever the type can take it: to from_file() where its signature names it, and onto the object it returns where it does not. That is what corrects a decoded file, which is the only thing that knows what the container claimed:

{
    "reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3VideoReference",
    "from_file": "motion.mp4",
    "fps": 30.0,
    "audio": null
}

fps overrides a rate the container got wrong — MiniMax-H3 resamples a reference onto its own 24 fps, so a wrong rate is a request conditioned at the wrong speed — and audio: null drops the decoded soundtrack, leaving a reference that conditions on motion and camera alone. A name that is neither an argument of from_file() nor a field of the object raises, with the fields it does have.

See workflows/templates/minimax/reference-to-video.json for a full example.

Objects Built From an Earlier Step

The same object can be built from what an earlier step generated, by naming the step instead of a file:

"references": [
    {
        "reference_type": "diffusers.modular_pipelines.minimax_h3.MiniMaxH3ImageReference",
        "from_previous_result": "draw_subject"
    }
]

from_file cannot do this — it names a file, and the object is built when the workflow loads, before any step has run. from_previous_result waits: the description is checked at load time and constructed once the step it names has produced its media, which is what lets one workflow generate a subject and then condition on it without writing it out and reading it back.

The media never touches the disk, so it arrives as the step produced it. Which field it lands in comes from the type's own kind:

kind Built from
image The generated image
video The generated frames, and the soundtrack generated with them if there was one
audio The generated soundtrack - or, for a step that produced audio alone (a music pipeline, a slice_audio task), the waveform itself. The rate travels with the waveform when the pipeline or task reports one (an AudioTrack - AudioLDM2, StableAudio, generate_speech); declare sample_rate beside from_previous_result only for a waveform from a task or file that carries none, and a declared rate always wins

Any other key is a field of the object and wins over what the media carried — "fps": 30.0 where the producing pipeline generated at a rate the consuming one does not share, for instance. A step that produced several artifacts fans out the same way every previous_result reference does: four images in, four videos out.

See workflows/templates/minimax/generated-subject-reference.json for a full example.

Objects Built From Named Arguments

Not every type a pipeline takes knows how to open a file. LTX-2's keyframe conditions and IC-LoRA references are plain dataclasses holding frames the caller already loaded, plus the numbers that say what to do with them. Those are written as the arguments to construct the object with:

"conditions": [
    {
        "condition_type": "diffusers.pipelines.ltx2.pipeline_ltx2_condition.LTX2VideoCondition",
        "from_arguments": {
            "frames": { "media_type": "image", "location": "first.png" },
            "index": 0,
            "strength": 1.0
        }
    },
    {
        "condition_type": "diffusers.pipelines.ltx2.pipeline_ltx2_condition.LTX2VideoCondition",
        "from_arguments": {
            "frames": { "media_type": "image", "location": "last.png" },
            "index": -1,
            "strength": 1.0
        }
    }
]

from_arguments holds every argument the type is constructed with - a key beside it raises rather than being silently dropped, and so does an argument the type does not take, naming the ones it does. The arguments inside are ordinary arguments: a media reference loads there, a variable: reference resolves there, and a previous_result: reference waits the way from_previous_result does - the object is constructed once the step it names has run.

Which of the three forms a type wants is decided by the type, not by preference:

Form For a type that
from_file opens the media itself, bringing its frame or sample rate along (MiniMax-H3's references)
from_previous_result declares a media kind, so a step's output lands in the right field on its own
from_arguments is a plain record of fields - no from_file(), no kind (LTX-2's conditions and references)

See workflows/templates/ltx2/keyframes.json for the file form and workflows/templates/ltx2/extend-clip.json for the one built from an earlier step.

Frames Across a Step Boundary

A pipeline that generates video with a soundtrack returns the two paired, and the result muxes them into one file. A step that works on the frames alone - a latent upsampler, an interpolator - returns frames without it. Two tasks carry the pieces across:

  • video_frames takes a generated video and returns its frames as one (frames, height, width, channels) uint8 array - the 0-255 shape LTX-2's conditions want, and one artifact rather than one per frame.
  • pair_audio puts a soundtrack back beside frames that lost it, so the step that saves them writes a single muxed mp4.
{
    "name": "film",
    "task": {
        "command": "pair_audio",
        "arguments": {
            "video": "previous_result:edit",
            "audio": "previous_result:balanced",
            "sample_rate": "variable:sample_rate"
        }
    },
    "result": { "content_type": "video/mp4", "fps": 24 }
}

audio takes either a waveform or the earlier step whose video carried the soundtrack, which brings its sample rate along; here it is an earlier step's waveform, so sample_rate is given explicitly. The frames keep the rate they arrived with - video given a file or an asset: carries that file's fps through to the saved mp4 - so result.fps is only needed for frames that bring no rate of their own. A mono track needs no preparation: an mp4 audio stream takes stereo and nothing else, so saving duplicates the single channel into two and emits a warning saying it did.

The track and the frames are two lengths a workflow used to have to keep equal by hand. "fit": "video" derives one from the other instead: the track is cut to exactly the frames it is laid over, or padded with silence and warned about when it is shorter than they are. That is what a soundtrack over a cut whose length is an argument needs - nothing in a workflow can multiply a list's length by a frame count, so music-video.json sliced a fixed 496 frames of song (it now cuts each shot to its own cut_frames) while its cut followed a shots list, and a two-shot run wrote 10.3 s of picture into a 20.7 s container and reported succeeded with no warnings (#142). Left unset the track is used as it is and a disagreement is warned about rather than passing in silence.

Which shape a pipeline argument wants is the pipeline's business, and the two LTX-2 paths differ: a keyframe condition is mapped from 0-255, so it takes the video_frames array, while an IC-LoRA reference goes through the video processor, which expects the [0, 1] frames the pipeline returned - previous_result:step.frames hands those over untouched.

Example: workflows/templates/assemble-and-score.json