diff --git a/docs/src/content/docs/Users Guide/03.Models/Local Models/anima.mdx b/docs/src/content/docs/Users Guide/03.Models/Local Models/anima.mdx index 3ca2830fbac..83275ec83ff 100644 --- a/docs/src/content/docs/Users Guide/03.Models/Local Models/anima.mdx +++ b/docs/src/content/docs/Users Guide/03.Models/Local Models/anima.mdx @@ -1,7 +1,7 @@ --- title: Anima description: Generate anime-style images with Anima, a 2B Cosmos Predict2 diffusion transformer with an LLM adapter, including its schedulers, regional prompting and ControlNet-LLLite adapters. -lastUpdated: 2026-09-30 +lastUpdated: 2026-10-02 sidebar: order: 1 --- @@ -12,7 +12,8 @@ built-in **LLM adapter** that translates the encoder's output for the transforme 16-channel Wan 2.1 / Qwen Image VAE. InvokeAI supports Anima for text-to-image, and on the Canvas for image-to-image, inpainting, outpainting -and regional prompts. +and regional prompts. Two community finetunes are supported as well: **Anima-2.9B** and **Anima-3.8B** +(see [below](#community-finetunes)). ## License @@ -53,6 +54,38 @@ three are listed. FLUX VAEs are **not** compatible. Anima also needs a T5-XXL tokenizer, which ships with InvokeAI; no T5 model has to be installed. +## Community finetunes + +Both finetunes keep Anima's VAE, schedulers and Canvas support, and install from the Starter Models. + +| Model | What changes | Components | +| ----- | ------------ | ---------- | +| **Anima-2.9B** (Preview v1) | 40 transformer blocks instead of 28, trained on 1.7M more images (cutoff July 2026). An int8 build (~3.1 GB) halves the memory of the bf16 one (~5.8 GB); it is not faster. | Same as Anima | +| **Anima-3.8B** (v1.1) | 52 blocks, plus a **semantic connector** bundled in the checkpoint that adds a second text encoder, **Qwen3.5 4B**, for better prompt adherence, multi-character binding and mixed natural-language/tag prompts. | Qwen3 0.6B **and** Qwen3.5 4B encoder, VAE | + +With an Anima-3.8B model selected, the **Components** section shows a third picker, **Qwen3.5 Encoder**, +and generation needs it. Its starter model (~4.8 GB) installs together with Anima-3.8B. The connector +runs once per denoising step, so Anima-3.8B is slower per step than Anima-2.9B (about 1.3 against 1.7 +iterations per second at 832×1216 on an RTX 4090), and its Qwen3.5 encoder adds memory while the prompt +is encoded. **Reset all to model defaults** sets 40 steps at CFG 6, the settings of the author's reference +workflow. + +[FP8 Storage](/configuration/optimization/fp8-storage/) changes Anima-3.8B's images more than Anima's: at +the same seed the composition can change, where Anima keeps it and only details differ. + +Only the v1.1 checkpoint of Anima-3.8B is supported. The earlier v1.0 transformer needs a separate adapter +file and is refused at install with a message that says so. + +Both finetunes are derivatives of Anima and fall under the same non-commercial license. + +:::note[LoRAs and LLLite adapters] +LoRAs and ControlNet-LLLite adapters address the transformer's blocks by position, and the finetunes +insert their extra blocks between the original ones. Invoke moves an adapter trained on Anima to the +blocks the finetune kept from Anima, and one trained on Anima-2.9B to its blocks in Anima-3.8B. +Anima-2.9B keeps Anima's blocks unchanged, so Anima adapters work there as on Anima itself. +Anima-3.8B changed them slightly and adds the semantic connector, so results there can differ more. +::: + ## Generation settings Selecting an Anima model applies these defaults: @@ -110,8 +143,8 @@ Base 1.0; the pose adapter is notably weak. Only the Inpainting and Sketch adapt | Node | Purpose | | ---------------------------- | --------------------------------------------------------- | -| **Main Model - Anima** | Loads the transformer, Qwen3 encoder and VAE | -| **Prompt - Anima** | Encodes a prompt, with an optional regional mask | +| **Main Model - Anima** | Loads the transformer, Qwen3 encoder and VAE, and for Anima-3.8B the Qwen3.5 encoder | +| **Prompt - Anima** | Encodes a prompt, with an optional regional mask; connect the Qwen3.5 encoder for Anima-3.8B | | **Denoise - Anima** | Runs sampling; accepts img2img latents, masks and LLLite adapters | | **Image to Latents - Anima** | VAE encode | | **Latents to Image - Anima** | VAE decode | diff --git a/docs/src/content/docs/contributing/new-model-integration.mdx b/docs/src/content/docs/contributing/new-model-integration.mdx index 44811bfc1b2..f53f0e34d76 100644 --- a/docs/src/content/docs/contributing/new-model-integration.mdx +++ b/docs/src/content/docs/contributing/new-model-integration.mdx @@ -1,7 +1,7 @@ --- title: New Model Type Integration Checklist description: A step-by-step checklist for integrating a new model architecture into InvokeAI, from the model manager to the webv2 frontend and the tests that enforce completeness. -lastUpdated: 2026-09-25 +lastUpdated: 2026-10-02 --- import { Steps, FileTree } from '@astrojs/starlight/components'; @@ -20,6 +20,14 @@ The examples use a hypothetical architecture, `NewModel`, with the base value `n - **Krea-2**: a single-stream DiT with a Qwen3-VL encoder, installable as Diffusers, single-file or GGUF. - **Z-Image**: variants with different defaults, LoRA and control. - **Qwen-Image**: reference images and an edit variant. +- **Anima**: depth-expanded finetunes (Anima-2.9B, Anima-3.8B) and a variant with a second text encoder. +::: + +:::tip[Extending an architecture that already exists?] +A finetune with a different depth, a new variant or an extra encoder for an architecture InvokeAI already +supports is a different job: most of this checklist does not apply, and some traps apply only there -- +installed models whose stored records must keep validating, and adapters trained on the original layout. +Start with [Extending an existing architecture](#14-extending-an-existing-architecture). ::: :::caution[Frontend scope] @@ -109,13 +117,21 @@ File: `invokeai/backend/model_manager/taxonomy.py` - `variant_type_adapter` in `taxonomy.py` - `ModelRecordChanges.variant` in `invokeai/app/services/model_records/model_records_base.py` - `tests/backend/architectures/test_variants.py` checks that the three agree and that the values are unique. + `tests/backend/architectures/test_variants.py` checks that the three agree and that the values are unique. A variant enum on a `base=Any` config (an encoder's) belongs to no architecture, so it also goes into `BASE_AGNOSTIC_VARIANT_ENUMS` in that test. + + :::caution[Adding a variant to an architecture with installed models] + The model config's `variant` field must be **required** (`Field()` with no default), and existing records need a [migration](#persisted-records-and-migrations). Both shortcuts break every installed model of that class: + - Without a migration, a stored record without `variant` fails validation and is skipped on read; the models vanish from the model list. + - With a default, `Config_Base.get_tag` adds the variant to the class's discriminator tag, which a stored record's dict does not carry, so no record of that class deserializes at all. + ::: 3. **Add encoder types (only for a new text encoder).** - There is no generic text-encoder type. Each encoder family has its own `ModelType` member, and a `ModelFormat` member for its folder layout. Existing examples are `Qwen3Encoder`, `Qwen3VLEncoder`, `QwenVLEncoder`, `MistralEncoder` and `T5Encoder`. + There is no generic text-encoder type. Each encoder family has its own `ModelType` member. A matching `ModelFormat` member is needed only if the encoder ships as a folder; a single-file encoder uses `ModelFormat.Checkpoint`. Existing examples are `Qwen3Encoder`, `Qwen3VLEncoder`, `Qwen35Encoder`, `QwenVLEncoder`, `MistralEncoder` and `T5Encoder`. + + Reuse an existing type whenever the encoder is a model InvokeAI already supports. Krea-2, Ideogram 4 and MiniMax H3 all use `ModelType.Qwen3VLEncoder`, told apart by `Qwen3VLVariantType`. Check whether the encoder weights are a stock checkpoint before adding a type: Anima-3.8B's `qwen35_4b.safetensors` ships its last layer without the MLP and replaces the LM head with a projection, and the conditioning was trained on exactly that. - Reuse an existing type whenever the encoder is a model InvokeAI already supports. Krea-2, Ideogram 4 and MiniMax H3 all use `ModelType.Qwen3VLEncoder`, told apart by `Qwen3VLVariantType`. Check whether the encoder weights are a stock checkpoint before adding a type. + Same width is not same family. A new family needs its own type when its architecture differs, even where an existing variant matches its hidden size: the Qwen3.5 4B and the Qwen3 4B are both 2560 wide. :::tip[Checklist: Taxonomy]{icon="approve-check"} @@ -296,6 +312,8 @@ Configs identify a model on disk. Every non-abstract subclass of `Config_Base` r ... ``` + Widths and key prefixes are not enough where two families share them. A `model.`-prefixed Qwen3.5 export satisfies every heuristic of the Qwen3 encoder configs; only its `linear_attn` layers tell it apart, so `Qwen35Encoder_Checkpoint_Config` requires them and the Qwen3 configs reject them. + **Configs must exclude each other.** Identification iterates `Config_Base.CONFIG_CLASSES`, which is a set, so the order in the `AnyModelConfig` union decides nothing. When two configs match the same file, `matches_sort_key` breaks the tie, which amounts to chance. For every existing config whose heuristic the new files could satisfy, add a negative check on one side or both. The same applies to LoRA and VAE configs, whose heuristics are often loose, for example "any key starting with `transformer_blocks.`". Cover each exclusion with a detection test. 3. **VAE config (only for a new VAE)** (`configs/vae.py`) @@ -335,8 +353,23 @@ Configs identify a model on disk. Every non-abstract subclass of `Config_Base` r ```python title="invokeai/backend/model_manager/configs/factory.py" Annotated[Main_Diffusers_NewModel_Config, Main_Diffusers_NewModel_Config.get_tag()], ``` + +6. **Refuse what is recognized but not supported** + + Raise `InvalidMatchError` for a file the config recognizes and cannot run, with a message that says what to install instead. `NotAMatchError` would let it fall through to another config or to `unknown`, and installing it as a working model would be worse. Anima refuses the Anima-3.8B v1.0 transformer this way: it loads cleanly as a 52-block Anima, but needs an adapter file InvokeAI does not load; its header (`qwen35_joint_dit_blocks`) is what gives it away. +### Persisted records and migrations + +Every installed model is stored as its config's JSON in the `models` table. `ModelRecordServiceSQL` *skips* a stored config that no longer validates, with a log line and nothing in the UI: the model just disappears from the model list. So any change that makes stored records invalid -- a new required field, a renamed enum value, a narrowed `Literal` -- needs a migration in `invokeai/app/services/shared/sqlite_migrator/migrations/` that rewrites them: + +- Name it `migration__.py` with a `build_migration(...)` that returns a `Migration`. Its `depends_on` names the newest migration on the main line. Discovery is automatic. +- Request `app_config` in `build_migration` if the callback has to open model files; stored paths may be relative to `app_config.models_path`. Fall back to a safe value when a file is gone. +- Touch only the records the change concerns (type, base and format), and leave records that already carry the field alone. +- Test it against an in-memory `models` table, as `test_migration_2026_10_01_add_anima_variant.py` does. + +`migration_2026_09_16_add_qwen3_vl_encoder_variant` (a new required encoder variant) and `migration_2026_10_01_add_anima_variant` (a new required main-model variant that reads each checkpoint's header) are complete examples. + :::tip[Checklist: Model configs]{icon="approve-check"} - [ ] Main configs per format (Diffusers, checkpoint, GGUF) - [ ] Detection helper and variant detection @@ -344,6 +377,8 @@ Configs identify a model on disk. Every non-abstract subclass of `Config_Base` r - [ ] VAE config and exclusions in the existing VAE detectors (if the VAE is new) - [ ] Encoder configs (if the encoder type is new) - [ ] Every new config in the `AnyModelConfig` union +- [ ] `InvalidMatchError` for recognized files that cannot run +- [ ] A migration for every change that would invalidate stored records ::: --- @@ -387,18 +422,21 @@ Loaders turn a config into an in-memory model. Every module in this folder is im Community single-file checkpoints often use key names that differ from diffusers, for example fused projections. Verify the conversion numerically against the diffusers module on a small config. + **Read structural hyperparameters from the checkpoint.** A loader that hard-codes the depth or a width and loads with `strict=False` builds the wrong model without an error when a finetune changes them: it fills the first N blocks of a deeper checkpoint and drops the rest as unexpected keys, which `log_unexpected_keys` reports at DEBUG only. `reject_incomplete_load` does not catch this either -- nothing is *missing*. Anima's loader took Anima-2.9B's first 28 of 40 blocks this way and generated degraded images. Count blocks from the state dict (`count_anima_dit_blocks`), refuse gaps, and read further hyperparameters from the safetensors header where the checkpoint records them (`read_safetensors_metadata`). + 3. **VAE loader (only for a new VAE)** Build the diffusers VAE class and load the converted state dict. `vae.py` and `flux.py` handle existing VAEs from native and ComfyUI layouts. 4. **Text encoder loader (only for a new encoder type)** - Register with `base=BaseModelType.Any` and the encoder's `ModelType`, and serve both `SubModelType.TextEncoder` and `SubModelType.Tokenizer`. Vendor tokenizers and configs the loader needs, as `invokeai/backend/qwen2_5_vl/` does, so a single-file install loads without network access. Add the vendored files to `[tool.setuptools.package-data]` in `pyproject.toml`, or a wheel install fails with `FileNotFoundError`. + Register with `base=BaseModelType.Any` and the encoder's `ModelType`, and serve both `SubModelType.TextEncoder` and `SubModelType.Tokenizer`. Vendor tokenizers and configs the loader needs, as `invokeai/backend/qwen2_5_vl/` does, so a single-file install loads without network access. Add the vendored files to `[tool.setuptools.package-data]` in `pyproject.toml`, or a wheel install fails with `FileNotFoundError`; `tests/test_package_data.py` checks every `*.json` and `*.json.gz` under `invokeai/backend`. :::tip[Checklist: Model loaders]{icon="approve-check"} - [ ] Loader per registered `(base, type, format)` - [ ] Key conversion for single-file and GGUF layouts, verified numerically against diffusers +- [ ] Depth and other structural hyperparameters read from the checkpoint, not hard-coded - [ ] Quantized formats via the existing `invokeai/backend/quantization/` helpers - [ ] FP8 Storage: the loader casts, or declares `Unimplemented` / `NotApplicable` at registration - [ ] VAE and encoder loaders (if new) @@ -588,6 +626,7 @@ Packages are discovered automatically. A directory of node modules without an `_ - **Backend package (optional).** Put reusable math in `invokeai/backend/new_model/`: noise, packing, position IDs, schedules and text encoding. FLUX, FLUX.2 and Krea-2 have one; Qwen-Image and Z-Image keep the loop in the denoise node. Use a package once code is shared between nodes or needs its own tests. - **Schedulers.** Flow-matching architectures use diffusers' `FlowMatchEulerDiscreteScheduler`, or the shared maps in `invokeai/backend/flux/schedulers.py`: `FLUX_SCHEDULER_MAP`, `ZIMAGE_SCHEDULER_MAP` and `ERNIE_IMAGE_SCHEDULER_MAP`, each with name values and labels. Declare which set the UI offers in `FeaturesFacet.scheduler_set`. Set `scheduler_applies_to_graph=True` only if the denoise node really takes a scheduler field. Reproduce the reference sigma schedule and shift (`mu`, time-shift type, terminal shift) exactly, and test it against the reference scheduler. - **External noise (optional).** Only if the denoise node accepts a noise tensor: extend `LatentNoiseType` and its shape and grid logic in `invokeai/app/invocations/latent_noise.py`, and extend `noise_type` in `invocations/noise.py`. Validate with `validate_noise_tensor_shape`. +- **Timestep-dependent conditioning.** Most architectures turn the text embeddings into the transformer's context once, before the loop. If a module between encoder and transformer reads the timestep -- Anima-3.8B's semantic connector does -- the context has to be recomputed at every step, for the positive, the negative and every regional conditioning. Pass the timestep exactly as the reference does: the connector scales sigma by 1000 before embedding it, so it gets a float32 sigma, where bf16's rounding would move its high frequencies by radians. A mask that leaves a query no key (an empty prompt encoded as one masked token) needs a defined result; `masked_sdpa` makes it zero instead of leaving it to the SDPA backend. - **Inpainting.** Rectified-flow models use `RectifiedFlowInpaintExtension` (`invokeai/backend/rectified_flow/rectified_flow_inpaint_extension.py`). - **Previews.** `LatentSpaceFacet` provides them; there is nothing to add in the step callback. @@ -767,10 +806,14 @@ Paths below are relative to `invokeai/frontend/webv2/src/features/generation/cor - Add a `case` to `getBaseComponentSectionPolicy` in `baseGenerationPolicies.ts`. Its `default` returns an empty policy silently. - Add encoder filters in `componentCompatibility.ts`. If a main model can bring its own components, update `isBundledMainForBase`. VAE acceptance comes from the backend. + - A slot that only one variant needs goes into the policy conditionally: `getBaseComponentSectionPolicy` receives the model, so branch on `model.variant`, as FLUX.2 dev and Anima-3.8B do. - A **new** settings key must be added in all of these places: - `GenerateSettings` (`types.ts`) - `GenerateComponentValueKey`, `COMPONENT_SETTING_LABELS`, `getComponentPolicyContext` and `getDefaultGenerateSettings` (`baseGenerationPolicies.ts`) - - `normalizeGenerateSettings` and `cloneGenerateWidgetValues` (`settings.ts`) + - the second `getComponentPolicyContext` in `src/features/generation/ui/GenerateComponentsSection.tsx` + - `normalizeGenerateSettings`, `cloneGenerateWidgetValues` and `syncGenerateWidgetValuesWithModels` (`settings.ts`) + - `RecalledComponentSetting` and `RECALLED_COMPONENTS` (`src/workbench/image-actions/imageRecall.ts`), under the metadata key the builder writes + - the test fixtures that build a complete `GenerateWidgetValues` literal (`imageRecall.test.ts`, `workbenchState.test.ts`, `importGalleryImages.test.ts`). Vitest does not type-check them; `lint:tsc` does. 5. **Canvas** @@ -789,11 +832,17 @@ Paths below are relative to `invokeai/frontend/webv2/src/features/generation/cor - `src/features/models/core/baseIdentity.ts`: `MODEL_BASES`, plus its pinned test - `src/features/models/core/types.ts`: `ModelBase` - - `src/features/models/core/taxonomy.ts`: variant values + - `src/features/models/core/taxonomy.ts`: variant values -- labels in `MODEL_VARIANT_LABELS`, main-model variants in `MAIN_VARIANTS_BY_BASE`, the variants of other types in `VARIANTS_BY_TYPE` - `src/features/models/core/relationships.ts`: which encoder and VAE types may link to the base - `src/features/workflow/core/modelRequirements.ts`: `BASE_LABELS` - New user-facing strings go in `invokeai/frontend/webv2/public/locales/en.json`. + A new **model type** (an encoder family) has its own set: + - `ModelTaxonomyType` in `types.ts`, and its label in the type list at the top of `taxonomy.ts` + - `NULL_BASE_ALLOWANCES` and `SINGLETON_LINK_TYPES` in `relationships.ts` + - the type label in `modelRequirements.ts` + - `CPU_ONLY_TYPES` in `src/features/models/ui/detail/CpuOnlySetting.tsx` if the config has `cpu_only` + 8. **Per-family UI (only for controls unique to the architecture)** `src/features/generation/ui/GenerateRenderSection.tsx` branches on the family for controls such as Krea-2 seed variance and Wan settings. Controls every architecture has come from capabilities; they need no code. @@ -804,17 +853,18 @@ Paths below are relative to `invokeai/frontend/webv2/src/features/generation/cor pnpm -C invokeai/frontend/webv2 test src/features/generation/core/graphCoverage.test.ts -u ``` - This rewrites `__snapshots__/generateGraphNodeTypes.json`. Then raise `SUPPORTED_BASE_COUNTS["generate"]` in `tests/app/invocations/test_frontend_graph_node_types.py`, and `LITERAL_FLOORS` if needed. That Python test checks every node type, edge field and literal value the builders emit against the backend's invocation registry. + This rewrites `__snapshots__/generateGraphNodeTypes.json`. Graph coverage and the remix round trip compile each base in the shapes listed in `graphCoverage.testing.ts`; a variant that changes the graph needs its own entry in `SHAPE_OVERRIDES`, and a component variant a slot filters for needs to be in `CANDIDATE_VARIANTS`, or neither test ever builds that graph. Then raise `SUPPORTED_BASE_COUNTS["generate"]` in `tests/app/invocations/test_frontend_graph_node_types.py`, and `LITERAL_FLOORS` if needed. That Python test checks every node type, edge field and literal value the builders emit against the backend's invocation registry. :::tip[Checklist: webv2]{icon="approve-check"} - [ ] Capabilities fixture regenerated - [ ] `contracts.ts`, `supportedBases.ts` and its test - [ ] Builder in `graph.ts` and registered in `GRAPH_BUILDERS` -- [ ] Component slot policy, encoder filters, new settings keys everywhere +- [ ] Component slot policy (per variant where only one needs it), encoder filters, new settings keys everywhere - [ ] `CANVAS_I2L_NODE_TYPES` and `BASE_CASES` - [ ] Recall - [ ] Model registry, labels, `en.json` +- [ ] `SHAPE_OVERRIDES` / `CANDIDATE_VARIANTS` for variant-dependent graphs - [ ] Graph snapshot regenerated, `SUPPORTED_BASE_COUNTS` raised - [ ] Changed interactions verified in the browser (`--webv2`) ::: @@ -857,19 +907,23 @@ Several tests pin hand-maintained tables, so a new architecture has to extend th | `tests/backend/architectures/test_latent_space.py` | `DECLARED_LATENT_SPACES` | | `tests/backend/architectures/test_default_settings.py` | Rows in `DEFAULT_SETTINGS_MATRIX` | | `tests/backend/architectures/test_conditioning.py` | The pinned numbers of conditioning types and architectures | -| `tests/backend/architectures/test_variants.py` | Nothing if the variant lists and unique values are right; it checks them | +| `tests/backend/architectures/test_variants.py` | A `base=Any` variant enum in `BASE_AGNOSTIC_VARIANT_ENUMS`; an architecture that gains variants out of the pinned list of variant-less ones. Otherwise it checks the variant lists and unique values itself | | `tests/backend/architectures/test_modality.py` | Nothing if `GENERATION_MODES` matches `ModalityFacet` | | `tests/backend/architectures/test_capabilities_fixture.py` | Regenerate the webv2 fixture (see section 11) | | `tests/app/invocations/test_frontend_graph_node_types.py` | `SUPPORTED_BASE_COUNTS`, possibly `LITERAL_FLOORS` | | `tests/backend/model_manager/load/test_diffusers_0XX_compatibility.py` | New classes, when the dependency is bumped | +| `tests/test_package_data.py` | Nothing if `package-data` covers the vendored files; it checks them | +| `src/features/generation/core/graphCoverage.testing.ts` (webv2) | `SHAPE_OVERRIDES` and `CANDIDATE_VARIANTS` for variant-dependent graphs | +| `src/workbench/image-actions/imageRecall.test.ts` (webv2) | The recorded component setting in the remix round-trip list | Beyond those tables, test what can silently go wrong: - detection of every format, including negative cases against similar architectures - key conversion against the diffusers module - the text-encoder template - the sampling schedule against the reference scheduler +- every migration, against an in-memory `models` table -Use real lightweight modules with small configs rather than mocks. +Use real lightweight modules with small configs rather than mocks. For layout tests, capture the key names and shapes of the real checkpoint into a fixture under `tests/backend/model_manager/load/state_dicts/` and build the model on the `meta` device: real extents are billions of elements. Where a module is ported rather than taken from `diffusers`/`transformers`, compare it once against the reference implementation on the real weights before writing the small-config tests. **CI gates for generated artifacts.** New or changed invocations change the OpenAPI schema. CI (`openapi-checks.yml`, `typegen-checks.yml`) requires the shared package's `invokeai/frontend/api/openapi.json` and `invokeai/frontend/api/schema.ts` to be current. Regenerate both from `invokeai/frontend/api` in the repository's Python environment after installing that package's locked dependencies: @@ -889,6 +943,49 @@ python ../../../scripts/generate_openapi_schema.py | pnpm typegen --- +## 14. Extending an existing architecture + +Finetunes and variants of a supported architecture reuse almost everything above. Decide first what the new checkpoints actually change: + +| What changes | Treat it as | Work | +| --- | --- | --- | +| Only the weights (finetune, merge) | nothing | at most a starter model | +| A structural hyperparameter nothing outside the loader depends on, such as depth | loader detection, **not** a variant | read it from the checkpoint ([5.2](#5-model-loaders)), test against the real key layout | +| What the graph or UI needs: an extra encoder, a different conditioning path, different defaults | a **variant** | the checklist below | +| The latent space, the VAE family or the transformer family | a new base | the rest of this guide | + +The scaffolder does not apply; it creates files for a new base. + +Anima-2.9B (40 blocks instead of 28) needed only the second row. Anima-3.8B (52 blocks, plus a connector that reads a second encoder) needed the third: + + +1. **Taxonomy.** The variant enum, with values unique across all enums, listed in the three places of [2.2](#2-taxonomy). Values describe what differs (`anima_qwen35`: conditioned on Qwen3.5 as well), not a release name, so the next finetune with the same requirements fits. + +2. **Config and migration.** A required `variant` field, detection from the checkpoint, `override_fields.pop("variant", None)` before detection, and a [migration](#persisted-records-and-migrations) for installed records. A model installed before the variant existed may already be the new kind; read its header in the migration rather than assume. + +3. **Declaration.** `VariantFacet`, per-variant `DefaultSettingsFacet` entries, the regenerated capabilities fixture. The model's defaults reach the UI through **Reset all to model defaults**; switching models keeps the user's settings. + +4. **Invocations.** New inputs optional, so existing workflows keep validating; node version bumps. The model loader checks the main model's variant and fails with a message naming what is missing, rather than letting the denoise node meet a `None`. + +5. **Frontend.** The per-variant slot and builder branch ([11.4](#11-frontend-webv2)), the variant in `MAIN_VARIANTS_BY_BASE` and `MODEL_VARIANT_LABELS`, `SHAPE_OVERRIDES` for both variants. + +6. **Unsupported releases.** Refuse a recognized checkpoint you do not support with `InvalidMatchError` and a pointer to the one you do. + +7. **Existing adapters.** LoRAs and control adapters address blocks by name, and names carry the block index. A depth expansion that inserts blocks between the original ones shifts every index after the first insertion, so an adapter trained on the original loads and runs against the wrong blocks. A remap needs the insertion positions, and model cards rarely list them, but the weights do. Compare each block of the expanded checkpoint with the original's: a block that was not trained further is bit-identical to its source, and one trained further still has a cosine similarity near 1 to it, while an inserted copy stays far lower. Record the measured table next to the architecture (Anima: `invokeai/backend/anima/block_layout.py`), apply it where the adapter binds, and say in the user guide which results to expect. + + +:::tip[Checklist: Extending an architecture]{icon="approve-check"} +- [ ] Decided: weights only, loader detection, variant or new base +- [ ] Structural hyperparameters read from the checkpoint, tested against a real-captured key fixture +- [ ] Variant: required field, detection, migration with a test, `VariantFacet`, defaults per variant +- [ ] Optional node inputs, variant check with a clear error in the model loader +- [ ] Per-variant slot and graph branch, variant labels, `SHAPE_OVERRIDES` +- [ ] `InvalidMatchError` for recognized but unsupported releases +- [ ] Adapter compatibility documented +::: + +--- + ## Summary: Minimal integration A minimal txt2img integration with a Diffusers main model, reusing an existing VAE and encoder: diff --git a/invokeai/app/invocations/anima/anima_denoise.py b/invokeai/app/invocations/anima/anima_denoise.py index 5e3c40736d4..adc67402b0a 100644 --- a/invokeai/app/invocations/anima/anima_denoise.py +++ b/invokeai/app/invocations/anima/anima_denoise.py @@ -60,6 +60,7 @@ from invokeai.backend.model_manager.taxonomy import BaseModelType from invokeai.backend.patches.layer_patcher import LayerPatcher, PatchSpec from invokeai.backend.patches.lora_conversions.anima_lora_constants import ANIMA_LORA_TRANSFORMER_PREFIX +from invokeai.backend.patches.lora_conversions.anima_lora_conversion_utils import anima_lora_for_depth from invokeai.backend.patches.model_patch_raw import ModelPatchRaw from invokeai.backend.rectified_flow.rectified_flow_inpaint_extension import ( RectifiedFlowInpaintExtension, @@ -482,16 +483,58 @@ def _load_text_conditionings( t5xxl_ids=cond_info.t5xxl_ids, t5xxl_weights=cond_info.t5xxl_weights, mask=mask, + qwen35_states=cond_info.qwen35_states, + qwen35_mask=cond_info.qwen35_mask, ) ) return text_conditionings + @staticmethod + def _run_llm_adapter( + transformer, + tc: AnimaTextConditioning, + dtype: torch.dtype, + timesteps: torch.Tensor | None, + ) -> torch.Tensor: + """Run the LLM Adapter -- or, on Anima-3.8B, the semantic connector -- for one conditioning. + + Args: + transformer: The AnimaTransformer instance (must be on device). + tc: The conditioning. + dtype: Inference dtype. + timesteps: The flow timestep, float32, shape (1,). Only the semantic connector reads it; it + is None for every other model, whose context does not change between steps. + + Returns: + Context of shape (1, 512, 1024). + """ + qwen3_embeds = tc.qwen3_embeds.unsqueeze(0) # (1, seq_len, 1024) + t5xxl_ids = tc.t5xxl_ids.unsqueeze(0) # (1, seq_len) + t5xxl_weights = None + if tc.t5xxl_weights is not None: + t5xxl_weights = tc.t5xxl_weights.unsqueeze(0).unsqueeze(-1).to(dtype=dtype) # (1, seq_len, 1) + semantic_states = None + semantic_mask = None + if timesteps is not None and tc.qwen35_states is not None: + # (num_layers, seq_len, 2560) -> one (1, seq_len, 2560) tensor per layer. + semantic_states = [state.unsqueeze(0).to(dtype=dtype) for state in tc.qwen35_states] + semantic_mask = tc.qwen35_mask.unsqueeze(0) if tc.qwen35_mask is not None else None + return transformer.preprocess_text_embeds( + qwen3_embeds.to(dtype=dtype), + t5xxl_ids, + t5xxl_weights=t5xxl_weights, + semantic_states=semantic_states, + semantic_mask=semantic_mask, + timesteps=timesteps, + ) + def _run_llm_adapter_for_regions( self, transformer, text_conditionings: list[AnimaTextConditioning], dtype: torch.dtype, + timesteps: torch.Tensor | None = None, ) -> AnimaRegionalTextConditioning: """Run the LLM Adapter separately for each regional conditioning and concatenate. @@ -499,6 +542,7 @@ def _run_llm_adapter_for_regions( transformer: The AnimaTransformer instance (must be on device). text_conditionings: List of per-region conditioning data. dtype: Inference dtype. + timesteps: See `_run_llm_adapter`. Returns: AnimaRegionalTextConditioning with concatenated context and masks. @@ -509,20 +553,8 @@ def _run_llm_adapter_for_regions( cur_len = 0 for tc in text_conditionings: - qwen3_embeds = tc.qwen3_embeds.unsqueeze(0) # (1, seq_len, 1024) - t5xxl_ids = tc.t5xxl_ids.unsqueeze(0) # (1, seq_len) - t5xxl_weights = None - if tc.t5xxl_weights is not None: - t5xxl_weights = tc.t5xxl_weights.unsqueeze(0).unsqueeze(-1) # (1, seq_len, 1) - - # Run the LLM Adapter to produce context for this region - context = transformer.preprocess_text_embeds( - qwen3_embeds.to(dtype=dtype), - t5xxl_ids, - t5xxl_weights=t5xxl_weights.to(dtype=dtype) if t5xxl_weights is not None else None, - ) # context shape: (1, 512, 1024) — squeeze batch dim - context_2d = context.squeeze(0) # (512, 1024) + context_2d = self._run_llm_adapter(transformer, tc, dtype, timesteps).squeeze(0) # (512, 1024) context_embeds_list.append(context_2d) context_ranges.append(Range(start=cur_len, end=cur_len + context_2d.shape[0])) @@ -537,6 +569,26 @@ def _run_llm_adapter_for_regions( context_ranges=context_ranges, ) + @staticmethod + def _check_qwen3_5_conditioning( + context: InvocationContext, + uses_connector: bool, + positive: list[AnimaTextConditioning], + negative: list[AnimaTextConditioning], + ) -> None: + """Refuse Anima-3.8B conditioning without Qwen3.5 states; note Qwen3.5 states nothing will read.""" + conditionings = [*positive, *negative] + if uses_connector: + if any(tc.qwen35_states is None for tc in conditionings): + raise ValueError( + "This Anima model bundles a Qwen3.5 semantic connector, but a prompt was encoded without " + "Qwen3.5. Connect the model loader's Qwen3.5 Encoder output to every Anima prompt node." + ) + elif any(tc.qwen35_states is not None for tc in conditionings): + context.logger.warning( + "This Anima model has no semantic connector; the prompt's Qwen3.5 encoding is unused." + ) + def _run_diffusion(self, context: InvocationContext) -> torch.Tensor: device = TorchDevice.choose_torch_device() inference_dtype = TorchDevice.choose_anima_inference_dtype(device) @@ -710,64 +762,68 @@ def _run_diffusion(self, context: InvocationContext) -> torch.Tensor: exit_stack.enter_context( LayerPatcher.apply_smart_model_patches( model=transformer, - patches=self._lora_iterator(context), + patches=self._lora_iterator(context, len(transformer.blocks)), prefix=ANIMA_LORA_TRANSFORMER_PREFIX, dtype=inference_dtype, cached_weights=cached_weights, ) ) - # Run LLM Adapter for each regional conditioning to produce context vectors. - # This must happen with the transformer on device since it uses the adapter weights. - if has_regional: - pos_regional = self._run_llm_adapter_for_regions(transformer, pos_text_conditionings, inference_dtype) - pos_context = pos_regional.context_embeds.unsqueeze(0) # (1, total_ctx_len, 1024) + # Anima-3.8B's semantic connector conditions the context on the timestep, so its context + # is recomputed at every step; every other Anima model's context is computed once here. + uses_connector = bool(getattr(transformer, "has_semantic_connector", False)) + self._check_qwen3_5_conditioning( + context, uses_connector, pos_text_conditionings, neg_text_conditionings or [] + ) - # Build regional prompting extension with cross-attention mask - regional_extension = AnimaRegionalPromptingExtension.from_regional_conditioning( - pos_regional, img_seq_len + def build_contexts( + sigma: float | None, + ) -> tuple[torch.Tensor, torch.Tensor | None, AnimaRegionalTextConditioning | None]: + # float32, as the reference passes it: the connector's timestep embedding scales sigma by + # 1000, where bf16's rounding of sigma would move its high frequencies by radians. + timesteps = ( + torch.tensor([sigma * ANIMA_MULTIPLIER], device=device, dtype=torch.float32) + if sigma is not None + else None ) - - # For negative, concatenate all regions without masking (matches Z-Image behavior) - neg_context = None - if do_cfg and neg_text_conditionings is not None: - neg_regional = self._run_llm_adapter_for_regions( - transformer, neg_text_conditionings, inference_dtype + # Must run with the transformer on device since it uses the adapter weights. + if has_regional: + pos_regional = self._run_llm_adapter_for_regions( + transformer, pos_text_conditionings, inference_dtype, timesteps ) - neg_context = neg_regional.context_embeds.unsqueeze(0) - else: - # Single conditioning — run LLM Adapter via normal forward path - tc = pos_text_conditionings[0] - pos_qwen3_embeds = tc.qwen3_embeds.unsqueeze(0) - pos_t5xxl_ids = tc.t5xxl_ids.unsqueeze(0) - pos_t5xxl_weights = None - if tc.t5xxl_weights is not None: - pos_t5xxl_weights = tc.t5xxl_weights.unsqueeze(0).unsqueeze(-1) - - # Pre-compute context via LLM Adapter - pos_context = transformer.preprocess_text_embeds( - pos_qwen3_embeds.to(dtype=inference_dtype), - pos_t5xxl_ids, - t5xxl_weights=pos_t5xxl_weights.to(dtype=inference_dtype) - if pos_t5xxl_weights is not None - else None, - ) - - neg_context = None + pos = pos_regional.context_embeds.unsqueeze(0) # (1, total_ctx_len, 1024) + # For negative, concatenate all regions without masking (matches Z-Image behavior) + neg = None + if do_cfg and neg_text_conditionings is not None: + neg_regional = self._run_llm_adapter_for_regions( + transformer, neg_text_conditionings, inference_dtype, timesteps + ) + neg = neg_regional.context_embeds.unsqueeze(0) + return pos, neg, pos_regional + pos = self._run_llm_adapter(transformer, pos_text_conditionings[0], inference_dtype, timesteps) + neg = None if do_cfg and neg_text_conditionings is not None: - ntc = neg_text_conditionings[0] - neg_qwen3 = ntc.qwen3_embeds.unsqueeze(0) - neg_ids = ntc.t5xxl_ids.unsqueeze(0) - neg_weights = None - if ntc.t5xxl_weights is not None: - neg_weights = ntc.t5xxl_weights.unsqueeze(0).unsqueeze(-1) - neg_context = transformer.preprocess_text_embeds( - neg_qwen3.to(dtype=inference_dtype), - neg_ids, - t5xxl_weights=neg_weights.to(dtype=inference_dtype) if neg_weights is not None else None, - ) - - regional_extension = None + neg = self._run_llm_adapter(transformer, neg_text_conditionings[0], inference_dtype, timesteps) + return pos, neg, None + + first_sigma = sigmas[0] if uses_connector else None + pos_context, neg_context, pos_regional = build_contexts(first_sigma) + context_sigma = first_sigma + + def contexts_at(sigma: float) -> tuple[torch.Tensor, torch.Tensor | None]: + nonlocal pos_context, neg_context, context_sigma + if uses_connector and sigma != context_sigma: + pos_context, neg_context, _ = build_contexts(sigma) + context_sigma = sigma + return pos_context, neg_context + + # The regional cross-attention mask depends only on how many context tokens each region + # contributes (512 each), not on their values, so it is built once. + regional_extension = ( + AnimaRegionalPromptingExtension.from_regional_conditioning(pos_regional, img_seq_len) + if pos_regional is not None + else None + ) # Apply regional prompting patch if we have regional masks exit_stack.enter_context(patch_anima_for_regional_prompting(transformer, regional_extension)) @@ -802,11 +858,12 @@ def _run_transformer(ctx: torch.Tensor, x: torch.Tensor, t: torch.Tensor) -> tor timestep = torch.tensor( [it.sigma_curr * ANIMA_MULTIPLIER], device=device, dtype=inference_dtype ).expand(latents.shape[0]) + step_pos_context, step_neg_context = contexts_at(it.sigma_curr) - noise_pred_cond = _run_transformer(pos_context, latents, timestep).float() + noise_pred_cond = _run_transformer(step_pos_context, latents, timestep).float() - if do_cfg and neg_context is not None: - noise_pred_uncond = _run_transformer(neg_context, latents, timestep).float() + if do_cfg and step_neg_context is not None: + noise_pred_uncond = _run_transformer(step_neg_context, latents, timestep).float() noise_pred = noise_pred_uncond + self.guidance_scale * (noise_pred_cond - noise_pred_uncond) else: noise_pred = noise_pred_cond @@ -858,11 +915,12 @@ def _run_transformer(ctx: torch.Tensor, x: torch.Tensor, t: torch.Tensor) -> tor timestep = torch.tensor( [sigma_curr * ANIMA_MULTIPLIER], device=device, dtype=inference_dtype ).expand(latents.shape[0]) + step_pos_context, step_neg_context = contexts_at(sigma_curr) - noise_pred_cond = _run_transformer(pos_context, latents, timestep).float() + noise_pred_cond = _run_transformer(step_pos_context, latents, timestep).float() - if do_cfg and neg_context is not None: - noise_pred_uncond = _run_transformer(neg_context, latents, timestep).float() + if do_cfg and step_neg_context is not None: + noise_pred_uncond = _run_transformer(step_neg_context, latents, timestep).float() noise_pred = noise_pred_uncond + self.guidance_scale * (noise_pred_cond - noise_pred_uncond) else: noise_pred = noise_pred_cond @@ -933,8 +991,12 @@ def step_callback(state: PipelineIntermediateState) -> None: return step_callback - def _lora_iterator(self, context: InvocationContext) -> Iterator[PatchSpec]: - """Iterate over LoRA models to apply to the transformer.""" + def _lora_iterator(self, context: InvocationContext, transformer_depth: int) -> Iterator[PatchSpec]: + """Iterate over LoRA models to apply to the transformer. + + A LoRA trained on a shallower Anima has its blocks moved to where this depth-expanded model keeps + them (see `invokeai.backend.anima.block_layout`). + """ for lora in self.transformer.loras: lora_info = context.models.load(lora.lora) if not isinstance(lora_info.model, ModelPatchRaw): @@ -942,4 +1004,11 @@ def _lora_iterator(self, context: InvocationContext) -> Iterator[PatchSpec]: f"Expected ModelPatchRaw for LoRA '{lora.lora.key}', got {type(lora_info.model).__name__}. " "The LoRA model may be corrupted or incompatible." ) - yield (lora_info.model, lora.weight, lora_info.model_in_ram()) + patch, moved_from = anima_lora_for_depth(lora_info.model, transformer_depth) + if moved_from is not None: + name = lora_info.config.name if lora_info.config is not None else lora.lora.key + context.logger.info( + f"LoRA '{name}' was trained on a {moved_from}-block Anima; applying it to the " + f"matching blocks of this {transformer_depth}-block model." + ) + yield (patch, lora.weight, lora_info.model_in_ram()) diff --git a/invokeai/app/invocations/anima/anima_model_loader.py b/invokeai/app/invocations/anima/anima_model_loader.py index 11ac96e7897..dd98ddced44 100644 --- a/invokeai/app/invocations/anima/anima_model_loader.py +++ b/invokeai/app/invocations/anima/anima_model_loader.py @@ -1,3 +1,5 @@ +from typing import Optional + from invokeai.app.invocations.baseinvocation import ( BaseInvocation, BaseInvocationOutput, @@ -9,12 +11,13 @@ from invokeai.app.invocations.model import ( ModelIdentifierField, Qwen3EncoderField, + Qwen35EncoderField, TransformerField, VAEField, ) from invokeai.app.services.shared.invocation_context import InvocationContext from invokeai.backend.architectures import accepted_vae_bases -from invokeai.backend.model_manager.taxonomy import BaseModelType, ModelType, SubModelType +from invokeai.backend.model_manager.taxonomy import AnimaVariantType, BaseModelType, ModelType, SubModelType @invocation_output("anima_model_loader_output") @@ -23,6 +26,11 @@ class AnimaModelLoaderOutput(BaseInvocationOutput): transformer: TransformerField = OutputField(description=FieldDescriptions.transformer, title="Transformer") qwen3_encoder: Qwen3EncoderField = OutputField(description=FieldDescriptions.qwen3_encoder, title="Qwen3 Encoder") + qwen3_5_encoder: Optional[Qwen35EncoderField] = OutputField( + default=None, + description=f"{FieldDescriptions.qwen3_5_encoder}. Set only for an Anima-3.8B model.", + title="Qwen3.5 Encoder", + ) vae: VAEField = OutputField(description=FieldDescriptions.vae, title="VAE") @@ -31,7 +39,7 @@ class AnimaModelLoaderOutput(BaseInvocationOutput): title="Main Model - Anima", tags=["model", "anima"], category="model", - version="1.4.0", + version="1.5.0", classification=Classification.Prototype, ) class AnimaModelLoaderInvocation(BaseInvocation): @@ -41,6 +49,7 @@ class AnimaModelLoaderInvocation(BaseInvocation): - Transformer: Cosmos Predict2 DiT + LLM Adapter (from single-file checkpoint) - Qwen3 Encoder: Qwen3 0.6B (standalone single-file) - VAE: AutoencoderKLQwenImage / Wan 2.1 VAE (standalone single-file) + - Qwen3.5 Encoder: Qwen3.5 4B, for Anima-3.8B only, whose bundled semantic connector reads it The T5-XXL tokenizer needed for LLM Adapter token IDs is bundled in the package, so no T5-XXL encoder model needs to be installed. @@ -69,6 +78,14 @@ class AnimaModelLoaderInvocation(BaseInvocation): title="Qwen3 Encoder", ) + qwen3_5_encoder_model: Optional[ModelIdentifierField] = InputField( + default=None, + description="Standalone Qwen3.5 4B Encoder model. Required by Anima-3.8B, ignored by every other Anima model.", + input=Input.Direct, + ui_model_type=ModelType.Qwen35Encoder, + title="Qwen3.5 Encoder", + ) + def invoke(self, context: InvocationContext) -> AnimaModelLoaderOutput: # Transformer always comes from the main model transformer = self.model.model_copy(update={"submodel_type": SubModelType.Transformer}) @@ -83,5 +100,25 @@ def invoke(self, context: InvocationContext) -> AnimaModelLoaderOutput: return AnimaModelLoaderOutput( transformer=TransformerField(transformer=transformer, loras=[]), qwen3_encoder=Qwen3EncoderField(tokenizer=qwen3_tokenizer, text_encoder=qwen3_encoder), + qwen3_5_encoder=self._qwen3_5_encoder(context), vae=VAEField(vae=vae), ) + + def _qwen3_5_encoder(self, context: InvocationContext) -> Optional[Qwen35EncoderField]: + """The Qwen3.5 encoder if the main model reads one, else None. Refuses a model that needs one and has none.""" + variant = getattr(context.models.get_config(self.model), "variant", None) + if variant != AnimaVariantType.Qwen35: + if self.qwen3_5_encoder_model is not None: + context.logger.warning( + f"{self.model.name} is not conditioned on Qwen3.5; ignoring the selected Qwen3.5 encoder." + ) + return None + if self.qwen3_5_encoder_model is None: + raise ValueError( + f"{self.model.name} bundles a Qwen3.5 semantic connector and needs a Qwen3.5 4B encoder. " + "Install one (qwen35_4b.safetensors) and select it as the Qwen3.5 Encoder." + ) + return Qwen35EncoderField( + tokenizer=self.qwen3_5_encoder_model.model_copy(update={"submodel_type": SubModelType.Tokenizer}), + text_encoder=self.qwen3_5_encoder_model.model_copy(update={"submodel_type": SubModelType.TextEncoder}), + ) diff --git a/invokeai/app/invocations/fields.py b/invokeai/app/invocations/fields.py index 4225422f645..e85bb75bb74 100644 --- a/invokeai/app/invocations/fields.py +++ b/invokeai/app/invocations/fields.py @@ -163,6 +163,7 @@ class FieldDescriptions: glm_encoder = "GLM (THUDM) tokenizer and text encoder" qwen3_encoder = "Qwen3 tokenizer and text encoder" qwen3_vl_encoder = "Qwen3-VL tokenizer and text encoder" + qwen3_5_encoder = "Qwen3.5 tokenizer and text encoder" mistral_encoder = "Mistral tokenizer/processor and text encoder" clip_embed_model = "CLIP Embed loader" clip_g_model = "CLIP-G Embed loader" diff --git a/invokeai/app/invocations/model.py b/invokeai/app/invocations/model.py index 8528aa29161..737cab9e2f5 100644 --- a/invokeai/app/invocations/model.py +++ b/invokeai/app/invocations/model.py @@ -162,6 +162,13 @@ class Qwen3VLEncoderField(BaseModel): loras: List[LoRAField] = Field(default_factory=list, description="LoRAs to apply on model loading") +class Qwen35EncoderField(BaseModel): + """Field for the Qwen3.5 text encoder Anima-3.8B's semantic connector reads.""" + + tokenizer: ModelIdentifierField = Field(description="Info to load tokenizer submodel") + text_encoder: ModelIdentifierField = Field(description="Info to load text_encoder submodel") + + class WanT5EncoderField(BaseModel): """Field for the UMT5-XXL text encoder used by Wan 2.2 models.""" diff --git a/invokeai/app/invocations/text_encoder/anima_text_encoder.py b/invokeai/app/invocations/text_encoder/anima_text_encoder.py index 7c446e17f48..3647a766c7d 100644 --- a/invokeai/app/invocations/text_encoder/anima_text_encoder.py +++ b/invokeai/app/invocations/text_encoder/anima_text_encoder.py @@ -7,6 +7,10 @@ Both outputs are stored together in AnimaConditioningInfo and used by the LLM Adapter inside the transformer during denoising. +Anima-3.8B additionally reads Qwen3.5 4B hidden states (layers 7, 15, 23, 31) through the semantic +connector bundled in its transformer. They are encoded here when a Qwen3.5 encoder is connected and +stored beside the Qwen3 ones; the connector itself runs inside the denoising loop, per step. + Key differences from Z-Image text encoder: - Anima uses Qwen3 0.6B (base model, NOT instruct) — no chat template - Anima additionally tokenizes with T5-XXL tokenizer to get token IDs @@ -28,9 +32,10 @@ TensorField, UIComponent, ) -from invokeai.app.invocations.model import Qwen3EncoderField +from invokeai.app.invocations.model import Qwen3EncoderField, Qwen35EncoderField from invokeai.app.invocations.primitives import AnimaConditioningOutput from invokeai.app.services.shared.invocation_context import InvocationContext +from invokeai.backend.anima.semantic_connector import QWEN35_LAYER_INDICES from invokeai.backend.patches.layer_patcher import LayerPatcher, PatchSpec from invokeai.backend.patches.lora_conversions.anima_lora_constants import ANIMA_LORA_QWEN3_PREFIX from invokeai.backend.patches.model_patch_raw import ModelPatchRaw @@ -52,13 +57,16 @@ # Qwen3 0.6B supports 32K context but the LLM Adapter doesn't need that much. QWEN3_MAX_SEQ_LEN = 8192 +# The reference encoder for Anima-3.8B's semantic connector truncates the Qwen3.5 prompt here. +QWEN3_5_MAX_SEQ_LEN = 1024 + @invocation( "anima_text_encoder", title="Prompt - Anima", tags=["prompt", "conditioning", "anima"], category="conditioning", - version="1.4.0", + version="1.5.0", classification=Classification.Prototype, idle_gpu_offloadable=True, ) @@ -80,10 +88,17 @@ class AnimaTextEncoderInvocation(BaseInvocation): default=None, description="A mask defining the region that this conditioning prompt applies to.", ) + qwen3_5_encoder: Qwen35EncoderField | None = InputField( + default=None, + title="Qwen3.5 Encoder", + description=f"{FieldDescriptions.qwen3_5_encoder}. Connect it for Anima-3.8B, which reads both encoders.", + input=Input.Connection, + ) @torch.no_grad() def invoke(self, context: InvocationContext) -> AnimaConditioningOutput: qwen3_embeds, t5xxl_ids, t5xxl_weights = self._encode_prompt(context) + qwen35_states, qwen35_mask = self._encode_qwen3_5(context) if self.qwen3_5_encoder else (None, None) # Move to CPU for storage qwen3_embeds = qwen3_embeds.detach().to("cpu") @@ -96,6 +111,8 @@ def invoke(self, context: InvocationContext) -> AnimaConditioningOutput: qwen3_embeds=qwen3_embeds, t5xxl_ids=t5xxl_ids, t5xxl_weights=t5xxl_weights, + qwen35_states=qwen35_states, + qwen35_mask=qwen35_mask, ) ] ) @@ -219,6 +236,52 @@ def _encode_prompt( return qwen3_embeds, t5xxl_ids, None + def _encode_qwen3_5(self, context: InvocationContext) -> tuple[torch.Tensor, torch.Tensor]: + """Encode the prompt with Qwen3.5 the way Anima-3.8B's reference encoder does. + + Returns: + Tuple of (states, mask), on the CPU. + - states: Shape (num_layers, seq_len, 2560), one row per layer in `QWEN35_LAYER_INDICES`. + - mask: Shape (seq_len,), True for real tokens. + """ + assert self.qwen3_5_encoder is not None + context.util.signal_progress("Running Qwen3.5 text encoder") + tokenizer_info = context.models.load(self.qwen3_5_encoder.tokenizer) + with tokenizer_info.model_on_device() as (_, tokenizer): + if not isinstance(tokenizer, PreTrainedTokenizerBase): + raise TypeError(f"Expected PreTrainedTokenizerBase for tokenizer, got {type(tokenizer).__name__}.") + # No template, no special tokens -- the prompt's tokens as they are. + token_ids = tokenizer.encode(self.prompt, add_special_tokens=False) + pad_token_id = tokenizer.pad_token_id + if len(token_ids) > QWEN3_5_MAX_SEQ_LEN: + logger.warning(f"Prompt was truncated to {QWEN3_5_MAX_SEQ_LEN} tokens for the Qwen3.5 encoder.") + token_ids = token_ids[:QWEN3_5_MAX_SEQ_LEN] + # An empty prompt becomes one padding token that nothing attends to, as in the reference, so the + # connector adds no Qwen3.5 signal for it. (The reference also masks every token from the first + # occurrence of id 151643 on -- the Qwen3 padding id, which in Qwen3.5's vocabulary is the + # ordinary token " 내용". That is not reproduced.) + is_empty = not token_ids + if is_empty: + if pad_token_id is None: + raise ValueError("The Qwen3.5 tokenizer has no padding token to encode an empty prompt with.") + token_ids = [pad_token_id] + mask = torch.full((len(token_ids),), not is_empty, dtype=torch.bool) + + # Imported here: transformers' qwen3_5 modeling module costs ~0.25 s, which every app start would + # otherwise pay for a node that only Anima-3.8B uses. + from invokeai.backend.qwen3_5.qwen3_5_encoder import Qwen35Encoder + + encoder_info = context.models.load(self.qwen3_5_encoder.text_encoder) + with encoder_info.model_on_device() as (_, encoder): + if not isinstance(encoder, Qwen35Encoder): + raise TypeError(f"Expected a Qwen3.5 encoder, got {type(encoder).__name__}.") + input_ids = torch.tensor([token_ids], device=encoder_info.compute_device) + # The connector was trained on what Anima-3.8B's encoder checkpoint computes, and that file + # ships its last tapped layer without the MLP. + states = encoder(input_ids, QWEN35_LAYER_INDICES, last_layer_attention_only=True) + stacked = torch.cat(states, dim=0).detach().to("cpu") + return stacked, mask + def _lora_iterator(self, context: InvocationContext) -> Iterator[PatchSpec]: """Iterate over LoRA models to apply to the Qwen3 text encoder.""" for lora in self.qwen3_encoder.loras: diff --git a/invokeai/app/services/model_records/model_records_base.py b/invokeai/app/services/model_records/model_records_base.py index 582cc1ef01a..4fa27d97d2a 100644 --- a/invokeai/app/services/model_records/model_records_base.py +++ b/invokeai/app/services/model_records/model_records_base.py @@ -21,6 +21,7 @@ from invokeai.backend.model_manager.configs.lora import LoraModelDefaultSettings from invokeai.backend.model_manager.configs.main import MainModelDefaultSettings from invokeai.backend.model_manager.taxonomy import ( + AnimaVariantType, BaseModelType, ClipVariantType, Flux2VariantType, @@ -36,6 +37,7 @@ PiDDecoderVariantType, Qwen3VariantType, Qwen3VLVariantType, + Qwen35VariantType, QwenImageVariantType, SchedulerPredictionType, WanLoRAVariantType, @@ -152,6 +154,8 @@ def validate_source_url(cls, v: Any) -> Optional[str]: | WanLoRAVariantType | Qwen3VariantType | Qwen3VLVariantType + | Qwen35VariantType + | AnimaVariantType | Krea2VariantType | MiniMaxH3VariantType | LTX2VariantType diff --git a/invokeai/app/services/shared/sqlite_migrator/migrations/migration_2026_10_01_add_anima_variant.py b/invokeai/app/services/shared/sqlite_migrator/migrations/migration_2026_10_01_add_anima_variant.py new file mode 100644 index 00000000000..66cc4975688 --- /dev/null +++ b/invokeai/app/services/shared/sqlite_migrator/migrations/migration_2026_10_01_add_anima_variant.py @@ -0,0 +1,99 @@ +"""Record the variant on Anima main models installed before Anima had more than one. + +`Main_Checkpoint_Anima_Config.variant` became a required field when Anima-3.8B joined: its bundled +semantic connector needs a Qwen3.5 encoder beside the Qwen3 one, and the frontend reads the +variant to ask for it. Records written before that carry no `variant`, and a stored config that +fails to validate is *skipped* on read (`ModelRecordServiceSQL._select_models`) -- every installed +Anima model would silently vanish from the model list. + +Nearly every such record is the Qwen3-only variant, because Anima-3.8B could not load before. It +could be *installed*, though: its checkpoint identified as an ordinary Anima. So the checkpoint's +header is read where the file is still there, and a bundled connector decides; a record whose file +cannot be read gets the Qwen3-only variant, which is what it was being loaded as. + +Mirrors `2026_09_16_add_qwen3_vl_encoder_variant`. +""" + +import json +import sqlite3 +from logging import Logger +from pathlib import Path +from typing import Any + +from safetensors import safe_open + +from invokeai.app.services.config.config_default import InvokeAIAppConfig +from invokeai.app.services.shared.sqlite_migrator.sqlite_migrator_common import Migration +from invokeai.backend.model_manager.taxonomy import AnimaVariantType, BaseModelType, ModelFormat, ModelType + +# Kept literal rather than imported from the config module: a migration must keep meaning what it +# meant when it was written, whatever the config later renames. +_CONNECTOR_SEGMENT = "anima_v2_connector." + + +def _bundles_semantic_connector(path: Path) -> bool: + with safe_open(path, framework="pt", device="cpu") as checkpoint: + return any(key.startswith(_CONNECTOR_SEGMENT) or f".{_CONNECTOR_SEGMENT}" in key for key in checkpoint.keys()) + + +class AddAnimaVariantCallback: + def __init__(self, app_config: InvokeAIAppConfig, logger: Logger) -> None: + self._models_path = app_config.models_path + self._logger = logger + + def __call__(self, cursor: sqlite3.Cursor) -> None: + cursor.execute("SELECT id, config FROM models;") + rows = cursor.fetchall() + + migrated: dict[str, int] = {} + for model_id, config_json in rows: + try: + config: dict[str, Any] = json.loads(config_json) + except json.JSONDecodeError as e: + self._logger.error("Invalid config JSON for model %s: %s", model_id, e) + raise + + # Only the single-file main model gained the field. Anima's LLLite control adapters share + # the base and have no variant at all. + if ( + config.get("type") != ModelType.Main.value + or config.get("base") != BaseModelType.Anima.value + or config.get("format") != ModelFormat.Checkpoint.value + or "variant" in config + ): + continue + + variant = self._variant_of(config) + config["variant"] = variant.value + cursor.execute("UPDATE models SET config = ? WHERE id = ?;", (json.dumps(config), model_id)) + migrated[variant.value] = migrated.get(variant.value, 0) + 1 + + if migrated: + self._logger.info(f"Recorded the variant on Anima model config(s): {migrated}") + + def _variant_of(self, config: dict[str, Any]) -> AnimaVariantType: + raw_path = config.get("path") + if not isinstance(raw_path, str) or not raw_path: + return AnimaVariantType.Qwen3 + path = Path(raw_path) + if not path.is_absolute(): + path = self._models_path / path + try: + if _bundles_semantic_connector(path): + return AnimaVariantType.Qwen35 + except Exception as e: + self._logger.warning( + "Could not read the header of Anima model %s (%s); recording it as %s.", + config.get("name", raw_path), + e, + AnimaVariantType.Qwen3.value, + ) + return AnimaVariantType.Qwen3 + + +def build_migration(app_config: InvokeAIAppConfig, logger: Logger) -> Migration: + return Migration( + id="2026_10_01_add_anima_variant", + depends_on="2026_09_26_add_workflow_revision", + callback=AddAnimaVariantCallback(app_config=app_config, logger=logger), + ) diff --git a/invokeai/backend/anima/anima_transformer.py b/invokeai/backend/anima/anima_transformer.py index 8f42d30b8fe..f7cc928306a 100644 --- a/invokeai/backend/anima/anima_transformer.py +++ b/invokeai/backend/anima/anima_transformer.py @@ -16,7 +16,7 @@ import logging import math -from typing import Optional, Tuple +from typing import TYPE_CHECKING, Optional, Tuple import torch import torch.nn.functional as F @@ -24,6 +24,9 @@ from einops.layers.torch import Rearrange from torch import nn +if TYPE_CHECKING: + from invokeai.backend.anima.semantic_connector import AnimaSemanticConnectorConfig + logger = logging.getLogger(__name__) @@ -785,6 +788,27 @@ def _apply_rope(x: torch.Tensor, cos: torch.Tensor, sin: torch.Tensor) -> torch. return (x * cos.unsqueeze(1)) + (_rotate_half(x) * sin.unsqueeze(1)) +def masked_sdpa( + q: torch.Tensor, k: torch.Tensor, v: torch.Tensor, attn_mask: Optional[torch.Tensor] = None +) -> torch.Tensor: + """`scaled_dot_product_attention` where a query row with no key to attend to yields zeros. + + Only the math backend defines that case (as zeros); the fused kernels may return NaN for it. + It is reachable: an empty prompt is encoded as one masked padding token, so every key of that + context is masked. Such rows are computed against an unmasked key set and then zeroed. + + Args: + attn_mask: Boolean, True where attention is allowed, broadcastable to (B, H, Lq, Lk). + """ + if attn_mask is None or attn_mask.dtype != torch.bool: + return F.scaled_dot_product_attention(q, k, v, attn_mask=attn_mask) + has_key = attn_mask.any(dim=-1, keepdim=True) + if bool(has_key.all()): + return F.scaled_dot_product_attention(q, k, v, attn_mask=attn_mask) + out = F.scaled_dot_product_attention(q, k, v, attn_mask=attn_mask | ~has_key) + return out * has_key.to(out.dtype) + + class LLMAdapterRotaryEmbedding(nn.Module): """Rotary position embedding for the LLM Adapter's attention layers.""" @@ -842,7 +866,7 @@ def forward( q = _apply_rope(q, *pos_q) k = _apply_rope(k, *pos_k) - y = F.scaled_dot_product_attention(q, k, v, attn_mask=attn_mask) + y = masked_sdpa(q, k, v, attn_mask) y = y.transpose(1, 2).reshape(x.shape[0], x.shape[1], -1).contiguous() return self.o_proj(y) @@ -1018,29 +1042,63 @@ class AnimaTransformer(MiniTrainDIT): "final_layer", ] - def __init__(self, *args, **kwargs): + def __init__(self, *args, semantic_connector: Optional["AnimaSemanticConnectorConfig"] = None, **kwargs): super().__init__(*args, **kwargs) self.llm_adapter = LLMAdapter() + if semantic_connector is not None: + from invokeai.backend.anima.semantic_connector import AnimaSemanticConnector + + # Named as in the Anima-3.8B bundle (`net.anima_v2_connector.*`), so its keys load as-is. + # FP8 Storage casts all of it. Keeping its timestep path (`time_mlp`, `time_modulation`, 159M + # params) in bf16 by analogy with `t_embedder` was measured and changes nothing: 30 steps, three + # seeds, PSNR against bf16 8.4/10.3/9.6 dB with that exception and 8.4/10.2/9.6 dB without, and + # 8.3/10.1/9.6 dB with the whole connector in bf16. FP8 Storage moves Anima-3.8B's composition + # through the DiT blocks, not the connector; Anima base stays at 14-16 dB. + self.anima_v2_connector = AnimaSemanticConnector(semantic_connector) + + @property + def has_semantic_connector(self) -> bool: + """True for Anima-3.8B: the context depends on Qwen3.5 hidden states and on the timestep.""" + return hasattr(self, "anima_v2_connector") def preprocess_text_embeds( self, text_embeds: torch.Tensor, text_ids: Optional[torch.Tensor], t5xxl_weights: Optional[torch.Tensor] = None, + *, + semantic_states: Optional[list[torch.Tensor]] = None, + semantic_mask: Optional[torch.Tensor] = None, + timesteps: Optional[torch.Tensor] = None, ) -> torch.Tensor: - """Run the LLM Adapter to produce conditioning for the DiT. + """Run the LLM Adapter (or, on Anima-3.8B, the semantic connector) to produce conditioning for the DiT. Args: text_embeds: Qwen3 hidden states. Shape: (batch, seq_len, 1024). text_ids: T5-XXL token IDs. Shape: (batch, seq_len). If None, returns text_embeds directly. t5xxl_weights: Optional per-token weights. Shape: (batch, seq_len, 1). + semantic_states: Qwen3.5 hidden states, one per connector layer. Each (batch, seq_len, 2560). + Required by, and only read by, a transformer with the semantic connector. + semantic_mask: True for valid Qwen3.5 tokens. Shape: (batch, seq_len). + timesteps: The flow timestep (sigma). Shape: (batch,). Required with the semantic connector, + whose output changes with it. Returns: Conditioning tensor. Shape: (batch, 512, 1024), zero-padded if needed. """ if text_ids is None: return text_embeds - out = self.llm_adapter(text_embeds, text_ids) + if self.has_semantic_connector: + if semantic_states is None or timesteps is None: + raise ValueError( + "This Anima model bundles a Qwen3.5 semantic connector and needs Qwen3.5 conditioning " + "and the timestep. Encode the prompt with a Qwen3.5 encoder connected." + ) + out = self.anima_v2_connector( + self.llm_adapter, text_embeds, text_ids, semantic_states, semantic_mask, timesteps + ) + else: + out = self.llm_adapter(text_embeds, text_ids) if t5xxl_weights is not None: out = out * t5xxl_weights if out.shape[1] < 512: diff --git a/invokeai/backend/anima/block_layout.py b/invokeai/backend/anima/block_layout.py new file mode 100644 index 00000000000..2319f09b355 --- /dev/null +++ b/invokeai/backend/anima/block_layout.py @@ -0,0 +1,47 @@ +"""Where the depth-expanded Anima finetunes keep the blocks of the model they grew from. + +Anima-2.9B and Anima-3.8B insert new DiT blocks between the existing ones (LLaMA-Pro style interleaved expansion), +so an adapter that addresses blocks by index -- a LoRA, a ControlNet-LLLite -- trained on the shallower model would, +from the third block on, patch different blocks than the ones it was trained on. These tables send it to the right +ones. Measured from the released checkpoints: + +- Anima-2.9B-preview-v1 (40 blocks) holds all 28 blocks of Anima base-v1.0 bit for bit: its card says only the new + layers were trained. The 12 new blocks sit at 2, 5, 8, 11, 14, 17, 21, 24, 27, 30, 33 and 36. +- Anima-3.8B v1.1 (52 blocks) holds the 40 blocks of Anima-2.9B nearly unchanged: each has a mean cosine similarity of + at least 0.9997 to its 2.9B block, and 11 are bit-identical. The 12 new blocks, at 3, 7, 11, ..., 47, are at about + 0.6 to the neighbor they were copied from. Ten blocks of Anima base-v1.0 are still bit-identical in it, all at the + positions below. + +Keyed by depth: these are the only depth-expanded Anima models. A future finetune with one of these depths but +another layout would need its own entry. +""" + +from typing import Optional + +ANIMA_BASE_DEPTH = 28 + +# (source depth, target depth) -> the target position of each source block. +_POSITIONS: dict[tuple[int, int], tuple[int, ...]] = { + (28, 40): (0, 1, 3, 4, 6, 7, 9, 10, 12, 13, 15, 16, 18, 19, 20, 22, 23, 25, 26, 28, 29, 31, 32, 34, 35, 37, 38, 39), + (40, 52): ( + 0, 1, 2, 4, 5, 6, 8, 9, 10, 12, 13, 14, 16, 17, 18, 20, 21, 22, 24, 25, + 26, 28, 29, 30, 32, 33, 34, 36, 37, 38, 40, 41, 42, 44, 45, 46, 48, 49, 50, 51, + ), +} # fmt: skip +_POSITIONS[(28, 52)] = tuple(_POSITIONS[(40, 52)][i] for i in _POSITIONS[(28, 40)]) + +KNOWN_DEPTHS = (28, 40, 52) + + +def adapter_block_positions(max_block_index: int, target_depth: int) -> Optional[tuple[int, ...]]: + """Where, in a `target_depth`-block model, the blocks of the model an adapter was trained on sit. + + The adapter's own depth is read off the highest block it addresses: the smallest known depth that has that block. + An adapter trained on Anima base touches block 27; one trained on Anima-2.9B, a block past 27. Returns None when + the adapter addresses blocks by index as before -- same depth, a deeper adapter than the model, or no layout for + the pair. + """ + source_depth = next((depth for depth in KNOWN_DEPTHS if depth > max_block_index), None) + if source_depth is None or source_depth >= target_depth: + return None + return _POSITIONS.get((source_depth, target_depth)) diff --git a/invokeai/backend/anima/conditioning_data.py b/invokeai/backend/anima/conditioning_data.py index b96c807835d..8d2186a1e54 100644 --- a/invokeai/backend/anima/conditioning_data.py +++ b/invokeai/backend/anima/conditioning_data.py @@ -39,6 +39,10 @@ class AnimaTextConditioning: t5xxl_ids: torch.Tensor t5xxl_weights: torch.Tensor | None = None mask: torch.Tensor | None = None + qwen35_states: torch.Tensor | None = None + """Qwen3.5 hidden states for Anima-3.8B's semantic connector. Shape: (num_layers, seq_len, 2560).""" + qwen35_mask: torch.Tensor | None = None + """True for valid Qwen3.5 tokens. Shape: (seq_len,).""" @dataclass diff --git a/invokeai/backend/anima/control_net_lllite.py b/invokeai/backend/anima/control_net_lllite.py index dd36f2e5b2f..df818222745 100644 --- a/invokeai/backend/anima/control_net_lllite.py +++ b/invokeai/backend/anima/control_net_lllite.py @@ -32,6 +32,9 @@ import torch.nn.functional as F from torch import nn +from invokeai.backend.anima.block_layout import adapter_block_positions +from invokeai.backend.util.logging import InvokeAILogger + ASPP_DEFAULT_DILATIONS: tuple[int, ...] = (1, 2, 4, 8) _SAVED_COND_PREFIX = "lllite_conditioning1." @@ -511,10 +514,21 @@ def set_multiplier(self, multiplier: float) -> None: m.multiplier = multiplier def apply_to(self, transformer: nn.Module) -> None: - """Swap the forward of each target Linear in ``transformer``. Idempotent.""" + """Swap the forward of each target Linear in ``transformer``. Idempotent. + + An adapter trained on a shallower Anima binds to the blocks where this depth-expanded model keeps + them (see ``invokeai.backend.anima.block_layout``). + """ self.restore() + depth = len(transformer.blocks) + positions = adapter_block_positions(self._max_block_index(), depth) + if positions is not None: + InvokeAILogger.get_logger(__name__).info( + f"ControlNet-LLLite addresses {self._max_block_index() + 1} blocks; binding it to the matching " + f"blocks of this {depth}-block model." + ) for m in self.lllite_modules: - target = self._resolve_target(transformer, m.lllite_name) + target = self._resolve_target(transformer, m.lllite_name, positions) if not isinstance(target, nn.Linear): raise TypeError(f"LLLite target for '{m.lllite_name}' is {type(target).__name__}, expected nn.Linear") if target.in_features != m.in_dim: @@ -536,12 +550,19 @@ def restore(self) -> None: for m in self.lllite_modules: m.unbind() + def _max_block_index(self) -> int: + return max( + int(match.group(1)) for m in self.lllite_modules if (match := MODULE_NAME_PATTERN.match(m.lllite_name)) + ) + @staticmethod - def _resolve_target(transformer: nn.Module, name: str) -> nn.Module: + def _resolve_target(transformer: nn.Module, name: str, positions: Sequence[int] | None = None) -> nn.Module: match = MODULE_NAME_PATTERN.match(name) if match is None: raise ValueError(f"Unrecognized LLLite module name: '{name}'") block_idx = int(match.group(1)) + if positions is not None: + block_idx = positions[block_idx] blocks = transformer.blocks if block_idx >= len(blocks): raise ValueError( diff --git a/invokeai/backend/anima/semantic_connector.py b/invokeai/backend/anima/semantic_connector.py new file mode 100644 index 00000000000..35e889cc689 --- /dev/null +++ b/invokeai/backend/anima/semantic_connector.py @@ -0,0 +1,347 @@ +"""Anima-3.8B's semantic connector: Qwen3.5 conditioning on top of Anima's native LLM adapter. + +Anima-3.8B v1.1 bundles this connector into its checkpoint under `anima_v2_connector.`. It replaces +the plain `LLMAdapter` call: the native adapter's six blocks still run, unchanged and in order, and +after each one two residual cross-attentions add Qwen3.5 information -- + +- the *quality anchor* attends to a per-block mix of four Qwen3.5 hidden layers (7, 15, 23, 31); +- the *v2 injection* attends to a 64-token bank that a timestep-aware perceiver resampler distills + from the same four layers. + +The resampler is conditioned on the diffusion timestep, so unlike the native adapter's output the +connector's output changes from step to step and has to be recomputed inside the denoising loop. + +Ported from the reference ComfyUI extension (https://github.com/GumGum10/comfyui-anima-3-8B, MIT), +`semantic_connector_v2.py` and `progressive_cross_adapter.py`. Module and parameter names follow +the checkpoint so the bundled weights load without a key map. Training-only machinery +(initialization, trainability, gradient checkpointing) is left out. +""" + +import math +from collections.abc import Sequence +from dataclasses import dataclass +from typing import Optional, Tuple + +import torch +import torch.nn.functional as F +from torch import nn + +from invokeai.backend.anima.anima_transformer import LLMAdapter, LLMAdapterAttention, masked_sdpa + +#: The Qwen3.5 4B hidden layers the connector was trained on. +QWEN35_LAYER_INDICES: tuple[int, ...] = (7, 15, 23, 31) +#: Qwen3.5 4B hidden size. +QWEN35_HIDDEN_SIZE = 2560 + + +@dataclass(frozen=True) +class AnimaSemanticConnectorConfig: + """The connector's hyperparameters. The bundle records them in its header (see `from_metadata`).""" + + num_queries: int = 64 + resampler_blocks: int = 6 + resampler_dim: int = 2048 + resampler_heads: int = 16 + mlp_hidden_dim: int = 5632 + semantic_source_dim: int = QWEN35_HIDDEN_SIZE + layer_indices: tuple[int, ...] = QWEN35_LAYER_INDICES + + #: The architecture name the bundle declares for the connector it carries. + ARCHITECTURE = "anima_qwen35_quality_anchored_semantic_connector_v2" + + @classmethod + def from_metadata(cls, metadata: dict[str, str]) -> "AnimaSemanticConnectorConfig": + """Read the hyperparameters from a v1.1 bundle's safetensors header. + + Raises: + ValueError: if the header declares a connector architecture other than the one ported here. + """ + architecture = metadata.get("anima_v2_adapter_architecture") + if architecture is not None and architecture != cls.ARCHITECTURE: + raise ValueError(f"Unsupported Anima semantic connector architecture {architecture!r}.") + + def number(name: str, default: int) -> int: + return int(metadata.get(f"anima_v2_adapter_{name}", default)) + + layer_indices = QWEN35_LAYER_INDICES + if (raw := metadata.get("anima_v2_adapter_layer_indices")) is not None: + layer_indices = tuple(int(v) for v in raw.strip("[]").split(",") if v.strip()) + return cls( + num_queries=number("semantic_query_tokens", cls.num_queries), + resampler_blocks=number("semantic_resampler_blocks", cls.resampler_blocks), + resampler_dim=number("semantic_resampler_dim", cls.resampler_dim), + resampler_heads=number("semantic_resampler_heads", cls.resampler_heads), + mlp_hidden_dim=number("semantic_resampler_mlp_hidden_dim", cls.mlp_hidden_dim), + layer_indices=layer_indices, + ) + + +def sinusoidal_timestep_embedding(timesteps: torch.Tensor, dim: int, max_period: int = 10_000) -> torch.Tensor: + """Embed Anima's continuous [0, 1] flow timestep. Scaled by 1000 like the reference.""" + half = dim // 2 + frequencies = torch.exp( + -math.log(max_period) * torch.arange(half, device=timesteps.device, dtype=torch.float32) / max(half, 1) + ) + angles = timesteps.float().reshape(-1, 1) * 1_000.0 * frequencies.reshape(1, -1) + embedding = torch.cat((angles.cos(), angles.sin()), dim=-1) + if dim % 2: + embedding = F.pad(embedding, (0, 1)) + return embedding + + +class ResamplerAttention(nn.Module): + """Plain multi-head attention with independently sized query and context streams. No norms, no RoPE.""" + + def __init__(self, query_dim: int, context_dim: int, num_heads: int): + super().__init__() + if query_dim % num_heads: + raise ValueError("query_dim must be divisible by num_heads") + self.num_heads = num_heads + self.head_dim = query_dim // num_heads + self.q_proj = nn.Linear(query_dim, query_dim, bias=False) + self.k_proj = nn.Linear(context_dim, query_dim, bias=False) + self.v_proj = nn.Linear(context_dim, query_dim, bias=False) + self.o_proj = nn.Linear(query_dim, query_dim, bias=False) + + def forward( + self, query: torch.Tensor, context: torch.Tensor, context_mask: Optional[torch.Tensor] = None + ) -> torch.Tensor: + batch, query_tokens, query_dim = query.shape + context_tokens = context.shape[1] + + def heads(value: torch.Tensor, tokens: int) -> torch.Tensor: + return value.reshape(batch, tokens, self.num_heads, self.head_dim).transpose(1, 2) + + q = heads(self.q_proj(query), query_tokens) + k = heads(self.k_proj(context), context_tokens) + v = heads(self.v_proj(context), context_tokens) + mask = None + if context_mask is not None: + mask = context_mask.to(torch.bool).reshape(batch, 1, 1, context_tokens) + attended = masked_sdpa(q, k, v, mask) + return self.o_proj(attended.transpose(1, 2).reshape(batch, query_tokens, query_dim)) + + +class TimestepModulatedNorm(nn.Module): + """Parameter-free LayerNorm whose scale and shift come from the timestep.""" + + def __init__(self, dim: int): + super().__init__() + self.norm = nn.LayerNorm(dim, elementwise_affine=False) + + def forward(self, value: torch.Tensor, scale: torch.Tensor, shift: torch.Tensor) -> torch.Tensor: + return self.norm(value) * (1.0 + scale.unsqueeze(1)) + shift.unsqueeze(1) + + +class SemanticResamplerBlock(nn.Module): + """One timestep-aware block: cross-attention to Qwen3.5, self-attention among queries, SwiGLU.""" + + def __init__(self, dim: int, qwen_dim: int, num_heads: int, mlp_hidden_dim: int): + super().__init__() + self.cross_norm = TimestepModulatedNorm(dim) + self.self_norm = TimestepModulatedNorm(dim) + self.mlp_norm = TimestepModulatedNorm(dim) + self.source_norm = nn.LayerNorm(qwen_dim) + self.cross_attention = ResamplerAttention(dim, qwen_dim, num_heads) + self.self_attention = ResamplerAttention(dim, dim, num_heads) + self.mlp_in = nn.Linear(dim, 2 * mlp_hidden_dim, bias=False) + self.mlp_out = nn.Linear(mlp_hidden_dim, dim, bias=False) + # Scale and shift for each of the three sub-blocks. + self.time_modulation = nn.Linear(dim, 6 * dim, bias=True) + + def forward( + self, + queries: torch.Tensor, + qwen_features: torch.Tensor, + timestep_embedding: torch.Tensor, + qwen_mask: Optional[torch.Tensor], + ) -> torch.Tensor: + modulation = self.time_modulation(F.silu(timestep_embedding)) + cross_scale, cross_shift, self_scale, self_shift, mlp_scale, mlp_shift = modulation.chunk(6, dim=-1) + queries = queries + self.cross_attention( + self.cross_norm(queries, cross_scale, cross_shift), self.source_norm(qwen_features), qwen_mask + ) + normalized = self.self_norm(queries, self_scale, self_shift) + queries = queries + self.self_attention(normalized, normalized) + normalized = self.mlp_norm(queries, mlp_scale, mlp_shift) + gate, value = self.mlp_in(normalized).chunk(2, dim=-1) + return queries + self.mlp_out(F.silu(gate) * value) + + +class TimestepAwareSemanticResampler(nn.Module): + """Compresses four Qwen3.5 layer streams into a bank of semantic query tokens, per timestep.""" + + def __init__( + self, + qwen_dim: int, + output_dim: int, + num_layers: int, + num_queries: int, + num_blocks: int, + model_dim: int, + num_heads: int, + mlp_hidden_dim: int, + ): + super().__init__() + self.num_layers = num_layers + self.model_dim = model_dim + self.query_tokens = nn.Parameter(torch.empty(1, num_queries, model_dim)) + self.layer_embeddings = nn.Parameter(torch.empty(num_layers, 1, qwen_dim)) + self.time_mlp = nn.Sequential(nn.Linear(model_dim, model_dim), nn.SiLU(), nn.Linear(model_dim, model_dim)) + self.blocks = nn.ModuleList( + [SemanticResamplerBlock(model_dim, qwen_dim, num_heads, mlp_hidden_dim) for _ in range(num_blocks)] + ) + self.output_norm = nn.LayerNorm(model_dim) + self.output_projection = nn.Linear(model_dim, output_dim, bias=False) + + def forward( + self, + hidden_states: Sequence[torch.Tensor], + timesteps: torch.Tensor, + source_mask: Optional[torch.Tensor], + ) -> torch.Tensor: + if len(hidden_states) != self.num_layers: + raise ValueError(f"Expected {self.num_layers} Qwen3.5 layers, got {len(hidden_states)}") + batch = hidden_states[0].shape[0] + qwen_features = torch.cat( + [hidden + self.layer_embeddings[index].unsqueeze(0) for index, hidden in enumerate(hidden_states)], dim=1 + ) + qwen_mask = torch.cat([source_mask] * self.num_layers, dim=1) if source_mask is not None else None + + time = sinusoidal_timestep_embedding(timesteps, self.model_dim).to(dtype=qwen_features.dtype) + time = self.time_mlp(time) + queries = self.query_tokens.expand(batch, -1, -1) + for block in self.blocks: + queries = block(queries, qwen_features, time, qwen_mask) + return self.output_projection(self.output_norm(queries)) + + +class QualityAnchor(nn.Module): + """The frozen v1 cross-attentions to Qwen3.5 that the v2 connector was trained on top of. + + One residual cross-attention per native adapter block, each reading its own learned softmax mix + of the four Qwen3.5 layers (`layer_mix_logits`). + """ + + def __init__(self, model_dim: int, num_heads: int, num_blocks: int, semantic_source_dim: int, num_layers: int): + super().__init__() + head_dim = model_dim // num_heads + self.query_norms = nn.ModuleList([nn.RMSNorm(model_dim, eps=1e-6) for _ in range(num_blocks)]) + self.source_norms = nn.ModuleList([nn.RMSNorm(semantic_source_dim, eps=1e-6) for _ in range(num_blocks)]) + self.semantic_attentions = nn.ModuleList( + [LLMAdapterAttention(model_dim, semantic_source_dim, num_heads, head_dim) for _ in range(num_blocks)] + ) + self.layer_mix_logits = nn.Parameter(torch.zeros(num_blocks, num_layers)) + + @staticmethod + def mixed_source(hidden_states: Sequence[torch.Tensor], mix: torch.Tensor, block_index: int) -> torch.Tensor: + return sum(hidden * mix[block_index, layer_index] for layer_index, hidden in enumerate(hidden_states)) # type: ignore[return-value] + + +class AnimaSemanticConnector(nn.Module): + """Quality-anchored semantic connector v2. Wraps a native `LLMAdapter` it does not own. + + The adapter is passed to `forward` instead of being held: it belongs to the transformer, which + patches it with LoRAs and moves it between devices, and a second registration would serialize and + move it twice. + """ + + def __init__( + self, + config: AnimaSemanticConnectorConfig, + model_dim: int = 1024, + num_heads: int = 16, + num_adapter_blocks: int = 6, + ): + super().__init__() + self.config = config + num_layers = len(config.layer_indices) + head_dim = model_dim // num_heads + self.quality_anchor = QualityAnchor( + model_dim, num_heads, num_adapter_blocks, config.semantic_source_dim, num_layers + ) + self.semantic_resampler = TimestepAwareSemanticResampler( + qwen_dim=config.semantic_source_dim, + output_dim=model_dim, + num_layers=num_layers, + num_queries=config.num_queries, + num_blocks=config.resampler_blocks, + model_dim=config.resampler_dim, + num_heads=config.resampler_heads, + mlp_hidden_dim=config.mlp_hidden_dim, + ) + self.v2_query_norms = nn.ModuleList([nn.RMSNorm(model_dim, eps=1e-6) for _ in range(num_adapter_blocks)]) + self.v2_semantic_norms = nn.ModuleList([nn.RMSNorm(model_dim, eps=1e-6) for _ in range(num_adapter_blocks)]) + self.v2_attentions = nn.ModuleList( + [LLMAdapterAttention(model_dim, model_dim, num_heads, head_dim) for _ in range(num_adapter_blocks)] + ) + + @staticmethod + def _attention_mask(mask: Optional[torch.Tensor]) -> Optional[torch.Tensor]: + if mask is None: + return None + mask = mask.to(torch.bool) + return mask[:, None, None, :] if mask.ndim == 2 else mask + + def forward( + self, + native_adapter: LLMAdapter, + native_source: torch.Tensor, + target_input_ids: torch.Tensor, + semantic_hidden_states: Sequence[torch.Tensor], + semantic_mask: Optional[torch.Tensor], + timesteps: torch.Tensor, + ) -> torch.Tensor: + """Produce the DiT's cross-attention context for one denoising step. + + Args: + native_adapter: The transformer's own `LLMAdapter`. + native_source: Qwen3 0.6B hidden states. Shape: (B, L_qwen3, 1024). + target_input_ids: T5-XXL token IDs. Shape: (B, L_t5). + semantic_hidden_states: Qwen3.5 hidden states, one per entry of `config.layer_indices`. + Each of shape (B, L_qwen35, 2560). + semantic_mask: True for valid Qwen3.5 tokens. Shape: (B, L_qwen35). A fully masked row + (the reference's encoding of an empty prompt) contributes nothing. + timesteps: The flow timestep (sigma). Shape: (B,). + + Returns: + Context of shape (B, L_t5, 1024), before Anima's padding to 512 tokens. + """ + if len(semantic_hidden_states) != len(self.config.layer_indices): + raise ValueError( + f"Expected {len(self.config.layer_indices)} Qwen3.5 layers, got {len(semantic_hidden_states)}" + ) + semantic_attention_mask = self._attention_mask(semantic_mask) + semantic_bank = self.semantic_resampler(semantic_hidden_states, timesteps, semantic_mask) + + x = native_adapter.embed(target_input_ids).to(dtype=native_source.dtype) + rotary = native_adapter.rotary_emb + + def positions(length: int) -> Tuple[torch.Tensor, torch.Tensor]: + return rotary(x, torch.arange(length, device=x.device, dtype=torch.long).unsqueeze(0)) + + query_rope = positions(x.shape[1]) + native_rope = positions(native_source.shape[1]) + anchor_rope = positions(semantic_hidden_states[0].shape[1]) + bank_rope = positions(semantic_bank.shape[1]) + anchor_mix = self.quality_anchor.layer_mix_logits.float().softmax(dim=-1).to(x.dtype) + + anchor = self.quality_anchor + for index, native_block in enumerate(native_adapter.blocks): + x = native_block(x, context=native_source, pos_target=query_rope, pos_source=native_rope) + anchor_source = anchor.source_norms[index](anchor.mixed_source(semantic_hidden_states, anchor_mix, index)) + x = x + anchor.semantic_attentions[index]( + anchor.query_norms[index](x), + context=anchor_source, + attn_mask=semantic_attention_mask, + pos_q=query_rope, + pos_k=anchor_rope, + ) + x = x + self.v2_attentions[index]( + self.v2_query_norms[index](x), + context=self.v2_semantic_norms[index](semantic_bank), + pos_q=query_rope, + pos_k=bank_rope, + ) + + return native_adapter.norm(native_adapter.out_proj(x)) diff --git a/invokeai/backend/architectures/defs/anima.py b/invokeai/backend/architectures/defs/anima.py index fec89d0d6e4..d19c83f1fa5 100644 --- a/invokeai/backend/architectures/defs/anima.py +++ b/invokeai/backend/architectures/defs/anima.py @@ -6,9 +6,10 @@ from invokeai.backend.architectures.facets.latent_space import WAN21_16, LatentSpaceFacet from invokeai.backend.architectures.facets.modality import ModalityFacet from invokeai.backend.architectures.facets.vae import VaeCompatibility, VaeFacet +from invokeai.backend.architectures.facets.variant import VariantFacet from invokeai.backend.architectures.registry import register from invokeai.backend.model_manager.configs.default_settings import MainModelDefaultSettings -from invokeai.backend.model_manager.taxonomy import BaseModelType +from invokeai.backend.model_manager.taxonomy import AnimaVariantType, BaseModelType, ModelType from invokeai.backend.stable_diffusion.diffusion.conditioning_data import AnimaConditioningInfo # Anima uses the Wan 2.1 VAE. @@ -17,7 +18,14 @@ LatentSpaceFacet(WAN21_16), ConditioningFacet(AnimaConditioningInfo), DefaultSettingsFacet( - {None: MainModelDefaultSettings(scheduler="euler", steps=35, cfg_scale=4.5, width=1024, height=1024)} + { + # The author's reference workflow for Anima-3.8B v1.1 samples at CFG 6 for 40 steps; the + # model card recommends CFG 4-7 and 28-50 steps. + AnimaVariantType.Qwen35: MainModelDefaultSettings( + scheduler="euler", steps=40, cfg_scale=6.0, width=1024, height=1024 + ), + None: MainModelDefaultSettings(scheduler="euler", steps=35, cfg_scale=4.5, width=1024, height=1024), + } ), ModalityFacet(frozenset({"txt2img", "img2img", "inpaint", "outpaint"}), metadata_slug="anima"), FeaturesFacet( @@ -48,4 +56,6 @@ } ) ), + # Variant values must be globally unique; see taxonomy.py. + VariantFacet({ModelType.Main: AnimaVariantType}), ) diff --git a/invokeai/backend/architectures/facets/variant.py b/invokeai/backend/architectures/facets/variant.py index 0976571f8db..98f6840499d 100644 --- a/invokeai/backend/architectures/facets/variant.py +++ b/invokeai/backend/architectures/facets/variant.py @@ -8,10 +8,11 @@ layer patcher. - The PiD decoder's resolution presets are one enum shared across five bases. -Three variant enums cannot live here at all: `ClipVariantType`, `Qwen3VariantType` and -`MistralVariantType` sit on `base=Any` configs, and `Any` is a sentinel the registry refuses to -register. They are named explicitly in `tests/backend/architectures/test_variants.py` so the -completeness check against `AnyVariant` stays total rather than quietly partial. +The text-encoder variant enums (`ClipVariantType`, `Qwen3VariantType`, `Qwen3VLVariantType`, +`Qwen35VariantType`, `MistralVariantType`) cannot live here at all: they sit on `base=Any` configs, +and `Any` is a sentinel the registry refuses to register. They are named explicitly in +`tests/backend/architectures/test_variants.py` so the completeness check against `AnyVariant` stays +total rather than quietly partial. This facet is deliberately *not* wired into `configs/factory.py`. That module validates a bare variant string against `variant_type_adapter` without passing the base (`build_common_fields`), so @@ -34,7 +35,7 @@ class VariantFacet(Facet): """The variant enums an architecture's models are labelled with, by model type. - Optional: four registered architectures (CogView4, ERNIE-Image, Ideogram 4, Anima) model no + Optional: three registered architectures (CogView4, ERNIE-Image, Ideogram 4) model no variants at all, and declaring an empty facet would be indistinguishable from declaring nothing. """ diff --git a/invokeai/backend/model_manager/configs/factory.py b/invokeai/backend/model_manager/configs/factory.py index a3260ecc65c..6aa10f0bd97 100644 --- a/invokeai/backend/model_manager/configs/factory.py +++ b/invokeai/backend/model_manager/configs/factory.py @@ -132,6 +132,7 @@ PiDDecoder_Checkpoint_SD3_Config, PiDDecoder_Checkpoint_SDXL_Config, ) +from invokeai.backend.model_manager.configs.qwen3_5_encoder import Qwen35Encoder_Checkpoint_Config from invokeai.backend.model_manager.configs.qwen3_encoder import ( Qwen3Encoder_Checkpoint_Config, Qwen3Encoder_GGUF_Config, @@ -537,6 +538,7 @@ def has_model_export(module: Any, name: Any, expected_bases: tuple[type, ...]) - # satisfies the text-only Qwen3 GGUF heuristic in full. Annotated[Qwen3VLEncoder_GGUF_Config, Qwen3VLEncoder_GGUF_Config.get_tag()], Annotated[Qwen3VLEncoder_Qwen3VLEncoder_Config, Qwen3VLEncoder_Qwen3VLEncoder_Config.get_tag()], + Annotated[Qwen35Encoder_Checkpoint_Config, Qwen35Encoder_Checkpoint_Config.get_tag()], # Qwen3 Encoder Annotated[Qwen3Encoder_Qwen3Encoder_Config, Qwen3Encoder_Qwen3Encoder_Config.get_tag()], Annotated[Qwen3Encoder_Checkpoint_Config, Qwen3Encoder_Checkpoint_Config.get_tag()], diff --git a/invokeai/backend/model_manager/configs/main.py b/invokeai/backend/model_manager/configs/main.py index 4e446818809..bb1c7153c76 100644 --- a/invokeai/backend/model_manager/configs/main.py +++ b/invokeai/backend/model_manager/configs/main.py @@ -43,6 +43,7 @@ from invokeai.backend.model_manager.configs.qwen3_encoder import _SDNQ_LOADABLE_QWEN_ARCHITECTURES from invokeai.backend.model_manager.model_on_disk import ModelOnDisk from invokeai.backend.model_manager.taxonomy import ( + AnimaVariantType, BaseModelType, Flux2VariantType, FluxVariantType, @@ -1286,6 +1287,22 @@ def _has_anima_keys(state_dict: dict[str | int, Any]) -> bool: return False +#: Where Anima-3.8B v1.1 bundles its semantic connector, relative to the transformer root. The +#: bundle records it in its header as `anima_v2_connector_prefix` (`net.anima_v2_connector.`). +ANIMA_V2_CONNECTOR_KEY_PREFIX = "anima_v2_connector." + +#: Anima's own wrapper namespaces, as `_has_anima_keys` accepts them. +_ANIMA_KEY_PREFIXES = ("", "net.", "model.diffusion_model.") + + +def _get_anima_variant(state_dict: dict[str | int, Any]) -> AnimaVariantType: + """An Anima checkpoint that bundles the semantic connector needs Qwen3.5 as well as Qwen3.""" + connector_prefixes = tuple(f"{prefix}{ANIMA_V2_CONNECTOR_KEY_PREFIX}" for prefix in _ANIMA_KEY_PREFIXES) + if any(isinstance(key, str) and key.startswith(connector_prefixes) for key in state_dict): + return AnimaVariantType.Qwen35 + return AnimaVariantType.Qwen3 + + class Main_Diffusers_ZImage_Config(Diffusers_Config_Base, Main_Config_Base, Config_Base): """Model config for Z-Image diffusers models (Z-Image-Turbo, Z-Image-Base).""" @@ -2884,6 +2901,11 @@ class Main_Checkpoint_Anima_Config(Checkpoint_Config_Base, Main_Config_Base, Con base: Literal[BaseModelType.Anima] = Field(default=BaseModelType.Anima) format: Literal[ModelFormat.Checkpoint] = Field(default=ModelFormat.Checkpoint) + # Required, and therefore not part of the discriminator tag. A default would put it into the tag + # (`Config_Base.get_tag`), which a stored record's dict does not carry, and every Anima record + # would stop deserializing. Records written before the field existed get it from + # `migration_2026_10_01_add_anima_variant`. + variant: AnimaVariantType = Field(description="Which text encoders the model is conditioned on.") @classmethod def from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self: @@ -2892,8 +2914,10 @@ def from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) - raise_for_override_fields(cls, override_fields) cls._validate_looks_like_anima_model(mod) + cls._reject_unbundled_qwen35_dit(mod) - return cls(**override_fields) + variant = override_fields.pop("variant", None) or _get_anima_variant(mod.load_state_dict()) + return cls(**override_fields, variant=variant) @classmethod def _validate_looks_like_anima_model(cls, mod: ModelOnDisk) -> None: @@ -2901,6 +2925,27 @@ def _validate_looks_like_anima_model(cls, mod: ModelOnDisk) -> None: if not has_anima_keys: raise NotAMatchError("state dict does not look like an Anima model") + @classmethod + def _reject_unbundled_qwen35_dit(cls, mod: ModelOnDisk) -> None: + """Refuse the Anima-3.8B v1.0 DiT, which needs an adapter file this release does not load. + + v1.0 trained 12 inserted blocks jointly with a separate Qwen3.5 cross-attention adapter; on its + own it loads cleanly as a 52-block Anima and generates with a conditioning signal its new + blocks were never trained without. v1.1 bundles a newer connector into the checkpoint and is + the release path, so the v1.0 file is refused here rather than installed as something it is + not. Its header names the joint training; a plain depth-expanded finetune carries no such key. + """ + metadata = mod.metadata() + if "qwen35_joint_dit_blocks" not in metadata: + return + if _get_anima_variant(mod.load_state_dict()) is AnimaVariantType.Qwen35: + return + raise InvalidMatchError( + "This is the Anima-3.8B v1.0 transformer, which only works together with its separate " + "Qwen3.5 adapter file. Install the Anima-3.8B v1.1 checkpoint instead " + "(Anima-3.8B-v1.1.safetensors), which bundles its connector." + ) + class Main_SDNQ_FLUX_Config(Checkpoint_Config_Base, Main_Config_Base, Config_Base): """Model config for SDNQ-quantized FLUX transformer models.""" diff --git a/invokeai/backend/model_manager/configs/qwen3_5_encoder.py b/invokeai/backend/model_manager/configs/qwen3_5_encoder.py new file mode 100644 index 00000000000..d06a32716b0 --- /dev/null +++ b/invokeai/backend/model_manager/configs/qwen3_5_encoder.py @@ -0,0 +1,82 @@ +"""Identification of Qwen3.5 text encoders (single-file). + +Qwen3.5 is a family of its own -- Gated DeltaNet linear attention interleaved with gated full +attention -- so it cannot share `ModelType.Qwen3Encoder` even where widths coincide: the Qwen3.5 4B +and the Qwen3 4B are both 2560 wide. `linear_attn` keys are what tell them apart, and the Qwen3 +configs reject them for the same reason this one requires them. +""" + +from typing import Any, Literal, Optional, Self + +from pydantic import Field + +from invokeai.backend.model_manager.configs.base import Checkpoint_Config_Base, Config_Base +from invokeai.backend.model_manager.configs.identification_utils import ( + NotAMatchError, + raise_for_override_fields, + raise_if_not_file, +) +from invokeai.backend.model_manager.model_on_disk import ModelOnDisk +from invokeai.backend.model_manager.taxonomy import BaseModelType, ModelFormat, ModelType, Qwen35VariantType +from invokeai.backend.quantization.gguf.ggml_tensor import GGMLTensor + +#: Where a Qwen3.5 language model's `layers.` and `embed_tokens.` sit: bare (ComfyUI text-encoder +#: exports, Anima-3.8B's `qwen35_4b.safetensors`), under a causal LM's `model.`, or under the +#: multimodal checkpoint's `model.language_model.`. +QWEN3_5_KEY_PREFIXES = ("", "model.", "model.language_model.") + +_QWEN3_5_HIDDEN_SIZES = {2560: Qwen35VariantType.Qwen35_4B} + + +def has_qwen3_5_linear_attention(state_dict: dict[str | int, Any]) -> bool: + """True if the state dict holds Gated DeltaNet (`linear_attn`) layers -- Qwen3.5, never Qwen3.""" + return any(isinstance(key, str) and ".linear_attn." in key for key in state_dict) + + +def qwen3_5_key_prefix(state_dict: dict[str | int, Any]) -> Optional[str]: + """The prefix the language model's keys live under, or None if this is not a Qwen3.5 model. + + Layer 0 is a linear-attention layer in every Qwen3.5 size (full attention is every fourth). + """ + for prefix in QWEN3_5_KEY_PREFIXES: + if f"{prefix}layers.0.linear_attn.in_proj_qkv.weight" in state_dict and f"{prefix}embed_tokens.weight" in ( + state_dict + ): + return prefix + return None + + +def get_qwen3_5_variant(state_dict: dict[str | int, Any], prefix: str) -> Optional[Qwen35VariantType]: + embed = state_dict.get(f"{prefix}embed_tokens.weight") + shape = getattr(embed, "shape", None) + if shape is None or len(shape) != 2: + return None + return _QWEN3_5_HIDDEN_SIZES.get(int(shape[1])) + + +class Qwen35Encoder_Checkpoint_Config(Checkpoint_Config_Base, Config_Base): + """Configuration for single-file Qwen3.5 text encoders (safetensors).""" + + base: Literal[BaseModelType.Any] = Field(default=BaseModelType.Any) + type: Literal[ModelType.Qwen35Encoder] = Field(default=ModelType.Qwen35Encoder) + format: Literal[ModelFormat.Checkpoint] = Field(default=ModelFormat.Checkpoint) + cpu_only: bool | None = Field(default=None, description="Whether this model should run on CPU only") + variant: Qwen35VariantType = Field(description="Qwen3.5 model size variant") + + @classmethod + def from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self: + raise_if_not_file(mod) + + raise_for_override_fields(cls, override_fields) + + state_dict = mod.load_state_dict() + prefix = qwen3_5_key_prefix(state_dict) + if prefix is None: + raise NotAMatchError("state dict does not look like a Qwen3.5 language model (no linear_attn layer 0)") + if any(isinstance(v, GGMLTensor) for v in state_dict.values()): + raise NotAMatchError("state dict looks like GGUF quantized") + + variant = override_fields.pop("variant", None) or get_qwen3_5_variant(state_dict, prefix) + if variant is None: + raise NotAMatchError("hidden size does not match a known Qwen3.5 variant") + return cls(**override_fields, variant=variant) diff --git a/invokeai/backend/model_manager/configs/qwen3_encoder.py b/invokeai/backend/model_manager/configs/qwen3_encoder.py index ea465061919..33de104f8b0 100644 --- a/invokeai/backend/model_manager/configs/qwen3_encoder.py +++ b/invokeai/backend/model_manager/configs/qwen3_encoder.py @@ -197,6 +197,17 @@ def _has_qwen3_specific_keys(tensor_names: Iterable[str | int]) -> bool: return False +def _has_qwen3_5_layers(tensor_names: Iterable[str | int]) -> bool: + """Check for Qwen3.5's Gated DeltaNet layers, which a Qwen3 model never has. + + Qwen3.5 is as wide as Qwen3 at 4B (2560) and keeps the `model.layers.*` / `q_norm` naming on its + full-attention layers, so it satisfies every heuristic above and would install as a Qwen3 4B + encoder the Qwen3 loader cannot build. Its linear-attention layers are named `linear_attn` in + PyTorch exports and `ssm_*` in llama.cpp ones. + """ + return any(isinstance(key, str) and (".linear_attn." in key or ".ssm_" in key) for key in tensor_names) + + def _get_qwen3_variant_from_state_dict(state_dict: dict[str | int, Any]) -> Optional[Qwen3VariantType]: """Determine Qwen3 variant (0.6B, 4B, or 8B) from state dict based on hidden_size. @@ -315,6 +326,9 @@ def _validate_looks_like_qwen3_model(cls, mod: ModelOnDisk) -> None: raise NotAMatchError( "state dict bundles a Qwen-VL visual tower; this is a Qwen-VL encoder, not a text-only Qwen3 encoder" ) + # Reject Qwen3.5: same width as Qwen3 at 4B, a different architecture (`Qwen35Encoder_*`). + if _has_qwen3_5_layers(state_dict): + raise NotAMatchError("state dict has Qwen3.5 linear-attention layers; this is a Qwen3.5 encoder, not Qwen3") @classmethod def _validate_does_not_look_like_gguf_quantized(cls, mod: ModelOnDisk) -> None: @@ -496,6 +510,9 @@ def _validate_looks_like_qwen3_model(cls, mod: ModelOnDisk) -> None: raise NotAMatchError( "state dict bundles a Qwen-VL visual tower; this is a Qwen-VL encoder, not a text-only Qwen3 encoder" ) + # Reject Qwen3.5: same width as Qwen3 at 4B, a different architecture (`Qwen35Encoder_*`). + if _has_qwen3_5_layers(state_dict): + raise NotAMatchError("state dict has Qwen3.5 linear-attention layers; this is a Qwen3.5 encoder, not Qwen3") # Reject Qwen3-VL language towers. The visual-tower check above cannot see them: llama.cpp # keeps the visual tower in a separate ``mmproj-*.gguf``, so a Qwen3-VL GGUF is structurally # identical to a text-only Qwen3 of the same width (both 36 layers at hidden 2560 for the @@ -568,6 +585,9 @@ def _validate_looks_like_qwen3_model(cls, mod: ModelOnDisk) -> None: raise NotAMatchError( "state dict bundles a Qwen-VL visual tower; this is a Qwen-VL encoder, not a text-only Qwen3 encoder" ) + # Reject Qwen3.5: same width as Qwen3 at 4B, a different architecture (`Qwen35Encoder_*`). + if _has_qwen3_5_layers(state_dict): + raise NotAMatchError("state dict has Qwen3.5 linear-attention layers; this is a Qwen3.5 encoder, not Qwen3") if not _has_qwen3_specific_keys(state_dict): raise NotAMatchError( "state dict lacks Qwen3 QK-normalization (q_norm/k_norm) weights; looks like a Qwen2 model, " diff --git a/invokeai/backend/model_manager/load/model_loaders/anima.py b/invokeai/backend/model_manager/load/model_loaders/anima.py index 4c77dd2ad5f..dd1ab6a7a3f 100644 --- a/invokeai/backend/model_manager/load/model_loaders/anima.py +++ b/invokeai/backend/model_manager/load/model_loaders/anima.py @@ -1,7 +1,8 @@ """Class for Anima model loading in InvokeAI.""" +from collections.abc import Mapping from pathlib import Path -from typing import Optional +from typing import Any, Optional import accelerate @@ -9,7 +10,7 @@ from invokeai.backend.model_manager.configs.base import Checkpoint_Config_Base from invokeai.backend.model_manager.configs.controlnet import ControlNet_Checkpoint_Anima_Config from invokeai.backend.model_manager.configs.factory import AnyModelConfig -from invokeai.backend.model_manager.configs.main import Main_Checkpoint_Anima_Config +from invokeai.backend.model_manager.configs.main import ANIMA_V2_CONNECTOR_KEY_PREFIX, Main_Checkpoint_Anima_Config from invokeai.backend.model_manager.load.fp8_capability import NotApplicable from invokeai.backend.model_manager.load.load_default import ModelLoader, _model_declared_skip_patterns from invokeai.backend.model_manager.load.model_loader_registry import ModelLoaderRegistry @@ -21,6 +22,7 @@ SubModelType, ) from invokeai.backend.quantization.fp8_scaled import ( + Fp8ScaledLayer, attach_fp8_scales, cast_state_dict, dequantize_fp8_scaled, @@ -34,6 +36,13 @@ strip_layer_path_prefix, warn_on_unattached_scales, ) +from invokeai.backend.quantization.int8_convrot import ( + drop_unconsumed_quantization_sidecars, + extract_int8_convrot_markers, + install_int8_convrot_layers, + reject_unmarked_int8_weights, + resolve_quantized_module_paths, +) from invokeai.backend.quantization.load_plan import reserve_for_load from invokeai.backend.util.devices import TorchDevice from invokeai.backend.util.logging import InvokeAILogger @@ -86,9 +95,10 @@ def _filter_non_model_keys(sd: dict) -> dict: } -# Anima's fixed transformer architecture. Kept at module level so tests can instantiate the real -# module graph (e.g. to pin `_skip_layerwise_casting_patterns` to actual dotted module paths) -# without duplicating these values. +# Anima's transformer architecture at the depth of the official release. Kept at module level so tests +# can instantiate the real module graph (e.g. to pin `_skip_layerwise_casting_patterns` to actual +# dotted module paths) without duplicating these values. `num_blocks` is the only value that differs +# between redistributions -- see `anima_transformer_config`. ANIMA_TRANSFORMER_CONFIG = { "max_img_h": 240, "max_img_w": 240, @@ -118,6 +128,39 @@ def _filter_non_model_keys(sd: dict) -> dict: } +def count_anima_dit_blocks(sd: Mapping[str, Any]) -> int: + """Number of DiT blocks in a prefix-stripped Anima state dict. + + Depth-expanded finetunes keep every other dimension of the official 28-block model: Anima-2.9B + has 40 blocks and Anima-3.8B 52, both grown by LLaMA-Pro style interleaved insertion. A model + built at the official depth would load the first 28 of them and drop the rest as unexpected + keys -- reported only at DEBUG, with the surviving blocks in the wrong order, so the result is + a degraded image rather than an error. + + Raises: + ValueError: if the state dict has no DiT blocks or the block indices have gaps. + """ + indices = { + int(parts[1]) + for key in sd + if isinstance(key, str) and key.startswith("blocks.") + for parts in [key.split(".", 2)] + if parts[1].isdigit() + } + if not indices: + raise ValueError("Anima checkpoint has no DiT blocks (no `blocks..` keys after the prefix strip).") + count = max(indices) + 1 + if len(indices) != count: + missing = sorted(set(range(count)) - indices) + raise ValueError(f"Anima checkpoint has gaps in its DiT blocks: {missing[:10]} missing of 0-{count - 1}.") + return count + + +def anima_transformer_config(sd: Mapping[str, Any]) -> dict[str, Any]: + """`ANIMA_TRANSFORMER_CONFIG` at the depth of the checkpoint in `sd` (prefix-stripped).""" + return {**ANIMA_TRANSFORMER_CONFIG, "num_blocks": count_anima_dit_blocks(sd)} + + @ModelLoaderRegistry.register(base=BaseModelType.Anima, type=ModelType.Main, format=ModelFormat.Checkpoint) class AnimaCheckpointModel(ModelLoader): """Class to load Anima transformer models from single-file checkpoints. @@ -150,6 +193,7 @@ def _load_from_singlefile( from safetensors.torch import load_file from invokeai.backend.anima.anima_transformer import AnimaTransformer + from invokeai.backend.anima.semantic_connector import AnimaSemanticConnectorConfig if not isinstance(config, Main_Checkpoint_Anima_Config): raise TypeError( @@ -170,59 +214,101 @@ def _load_from_singlefile( target_device = TorchDevice.choose_torch_device() model_dtype = TorchDevice.choose_anima_inference_dtype(target_device) - # ComfyUI 'scaled fp8': an fp8 weight plus a `weight_scale`. `_filter_non_model_keys` above - # keeps those keys, and `load_state_dict` below rejects the checkpoint outright over them -- - # 500 unexpected keys on a plain scaled export, 749 on one that also ships `comfy_quant` - # markers. Such a checkpoint therefore does not load at all today. - # - # Anima keeps `q_proj`/`k_proj`/`v_proj` separate and the only key rewrite is a prefix strip, - # so a sibling scale travels with its weight and nothing has to be split. - keep_fp8 = self._keep_fp8_weights(config, SubModelType.Transformer) - header_hints = parse_quantization_metadata(read_safetensors_metadata(model_path, logger)) - # The header names layers in the checkpoint's own scheme -- `net.`-prefixed on every Anima - # redistribution measured -- while the scales are read after `_strip_anima_bundle_prefix` - # has run. Without this the per-layer flags, `full_precision_matrix_mult` above all, match - # nothing and are silently ignored. - layer_hints = { - **extract_comfy_quant_hints(sd), - **strip_layer_path_prefix(header_hints), - } - fp8_layers = extract_fp8_scaled_layers(sd, layer_hints=layer_hints) - - # Create an empty AnimaTransformer with Anima's default architecture parameters + # Two ComfyUI side-channel formats share the `weight` + `weight_scale` layout, and a checkpoint + # carries one or the other (Anima-2.9B ships an `int8_tensorwise` build beside its bf16 one). + # The int8 markers come out first: the fp8 path below reads `.comfy_quant` as layer hints and + # would take an int8 weight for an unscaled one. The rejection sits outside the branch so an + # int8 weight whose marker is missing or unreadable is refused instead of cast as raw codes. + int8_markers = extract_int8_convrot_markers(sd) + reject_unmarked_int8_weights(sd, int8_markers, "Anima") + + metadata = read_safetensors_metadata(model_path, logger) + # Anima-3.8B bundles a semantic connector whose hyperparameters live in the header. Without + # it in the model the connector's tensors would fall to `strict=False` and the 52-block DiT + # would generate from the native adapter alone, which its new blocks were not trained on. + semantic_connector = None + if any(key.startswith(ANIMA_V2_CONNECTOR_KEY_PREFIX) for key in sd): + semantic_connector = AnimaSemanticConnectorConfig.from_metadata(metadata) + + keep_fp8 = False + fp8_layers: dict[str, Fp8ScaledLayer] = {} + if int8_markers: + sd = drop_unconsumed_quantization_sidecars(sd) + # The prefix strip is the only key rewrite, so the markers already name the model's modules. + quantized = resolve_quantized_module_paths(int8_markers, key_map={}) + else: + # ComfyUI 'scaled fp8': an fp8 weight plus a `weight_scale`. `_filter_non_model_keys` above + # keeps those keys, and `load_state_dict` below rejects the checkpoint outright over them -- + # 500 unexpected keys on a plain scaled export, 749 on one that also ships `comfy_quant` + # markers. Such a checkpoint therefore does not load at all today. + # + # Anima keeps `q_proj`/`k_proj`/`v_proj` separate and the only key rewrite is a prefix strip, + # so a sibling scale travels with its weight and nothing has to be split. + keep_fp8 = self._keep_fp8_weights(config, SubModelType.Transformer) + header_hints = parse_quantization_metadata(metadata) + # The header names layers in the checkpoint's own scheme -- `net.`-prefixed on every Anima + # redistribution measured -- while the scales are read after `_strip_anima_bundle_prefix` + # has run. Without this the per-layer flags, `full_precision_matrix_mult` above all, match + # nothing and are silently ignored. + layer_hints = { + **extract_comfy_quant_hints(sd), + **strip_layer_path_prefix(header_hints), + } + fp8_layers = extract_fp8_scaled_layers(sd, layer_hints=layer_hints) + + # Build at the checkpoint's own depth: 28 blocks for the official release, more for the + # depth-expanded finetunes (Anima-2.9B: 40, Anima-3.8B: 52). with accelerate.init_empty_weights(): - model = AnimaTransformer(**ANIMA_TRANSFORMER_CONFIG) + model = AnimaTransformer(**anima_transformer_config(sd), semantic_connector=semantic_connector) + if semantic_connector is not None: + logger.info(f"Anima: {len(model.blocks)}-block transformer with a bundled Qwen3.5 semantic connector") skip_patterns = _model_declared_skip_patterns(model) - # Reserve before anything below widens a weight -- the fold and the split both do, and - # reserving afterwards lets either peak land on a cache that was only ever sized for the - # file. `scaled_layers` is what keeps the prediction honest where the weights are kept: the - # split also widens layers whose scale layout `scaled_mm` cannot apply, and without the - # mapping the prediction would charge those 1 byte/element and arrive at 2. Where they are - # not kept the prediction charges every float at `model_dtype`, folded yet or not, so the - # number is the same on either side of the fold -- what changes is when the room exists. - # Building the model first costs nothing: `init_empty_weights` leaves every param on meta. - reserve_for_load( - self._ram_cache.make_room, - sd, - model_dtype, - keep_fp8=keep_fp8, - model=model, - skip_patterns=skip_patterns, - fp8_layers=fp8_layers, - nvfp4_payloads={}, - ) + if int8_markers: + quantized = install_int8_convrot_layers( + model, + sd, + quantized, + model_dtype, + architecture="Anima", + reserve=self._ram_cache.make_room, + skip_patterns=skip_patterns, + ) + logger.info( + f"Anima: kept {len(quantized)} of {len(int8_markers)} layer(s) in int8 " + "(int8_tensorwise checkpoint, dequantized per forward)" + ) + kept = 0 + else: + # Reserve before anything below widens a weight -- the fold and the split both do, and + # reserving afterwards lets either peak land on a cache that was only ever sized for the + # file. `scaled_layers` is what keeps the prediction honest where the weights are kept: the + # split also widens layers whose scale layout `scaled_mm` cannot apply, and without the + # mapping the prediction would charge those 1 byte/element and arrive at 2. Where they are + # not kept the prediction charges every float at `model_dtype`, folded yet or not, so the + # number is the same on either side of the fold -- what changes is when the room exists. + # Building the model first costs nothing: `init_empty_weights` leaves every param on meta. + reserve_for_load( + self._ram_cache.make_room, + sd, + model_dtype, + keep_fp8=keep_fp8, + model=model, + skip_patterns=skip_patterns, + fp8_layers=fp8_layers, + nvfp4_payloads={}, + ) - if fp8_layers and not keep_fp8: - # Neither the matmul nor FP8 Storage asked for them, so keeping them quantized would - # dequantize on every forward to save memory nobody wanted saved. Fold the scale in. - dequantize_fp8_scaled(sd, fp8_layers, model_dtype) - fp8_layers = {} + if fp8_layers and not keep_fp8: + # Neither the matmul nor FP8 Storage asked for them, so keeping them quantized would + # dequantize on every forward to save memory nobody wanted saved. Fold the scale in. + dequantize_fp8_scaled(sd, fp8_layers, model_dtype) + fp8_layers = {} - # Layers the cast would dequantize anyway are folded here too, scale applied, so the cast - # never strips a scale that can no longer be put back. - fp8_layers = split_fp8_scaled_layers(sd, fp8_layers, model_dtype, model=model, skip_patterns=skip_patterns) - kept = cast_state_dict(sd, model_dtype, keep_fp8=keep_fp8, model=model, skip_patterns=skip_patterns) + # Layers the cast would dequantize anyway are folded here too, scale applied, so the cast + # never strips a scale that can no longer be put back. + fp8_layers = split_fp8_scaled_layers(sd, fp8_layers, model_dtype, model=model, skip_patterns=skip_patterns) + kept = cast_state_dict(sd, model_dtype, keep_fp8=keep_fp8, model=model, skip_patterns=skip_patterns) load_result = model.load_state_dict(sd, assign=True, strict=False) log_unexpected_keys("Anima transformer checkpoint", load_result.unexpected_keys) diff --git a/invokeai/backend/model_manager/load/model_loaders/qwen3_5_encoder.py b/invokeai/backend/model_manager/load/model_loaders/qwen3_5_encoder.py new file mode 100644 index 00000000000..9e14c312611 --- /dev/null +++ b/invokeai/backend/model_manager/load/model_loaders/qwen3_5_encoder.py @@ -0,0 +1,117 @@ +"""Loader for single-file Qwen3.5 text encoders.""" + +from pathlib import Path +from typing import Any, Optional + +import accelerate +import torch +from torch import nn + +from invokeai.backend.model_manager.configs.factory import AnyModelConfig +from invokeai.backend.model_manager.configs.qwen3_5_encoder import Qwen35Encoder_Checkpoint_Config, qwen3_5_key_prefix +from invokeai.backend.model_manager.load.load_default import ModelLoader, _model_declared_skip_patterns +from invokeai.backend.model_manager.load.model_loader_registry import ModelLoaderRegistry +from invokeai.backend.model_manager.taxonomy import AnyModel, BaseModelType, ModelFormat, ModelType, SubModelType +from invokeai.backend.quantization.fp8_scaled import cast_state_dict, reject_quantized_side_channel +from invokeai.backend.quantization.load_plan import reserve_for_load +from invokeai.backend.util.devices import TorchDevice +from invokeai.backend.util.state_dict_loading import log_unexpected_keys, reject_incomplete_load + + +class _ZeroMLP(nn.Module): + """Stands in for an MLP sub-block the checkpoint does not ship, so the layer adds nothing there.""" + + def forward(self, x: torch.Tensor) -> torch.Tensor: + return torch.zeros_like(x) + + +def _language_model_state_dict(sd: dict[str, Any], prefix: str) -> dict[str, Any]: + """The decoder's own tensors, keyed as `Qwen35Encoder` names them. + + Everything else is dropped: a multimodal export's vision tower, the LM head, the final norm and, + in Anima-3.8B's `qwen35_4b.safetensors`, the 2560->1024 projection head stored under `norm.*` -- + none of which an encoder that returns intermediate hidden states runs. + """ + return { + key[len(prefix) :]: tensor + for key, tensor in sd.items() + if key.startswith(prefix) and key[len(prefix) :].startswith(("embed_tokens.", "layers.")) + } + + +@ModelLoaderRegistry.register(base=BaseModelType.Any, type=ModelType.Qwen35Encoder, format=ModelFormat.Checkpoint) +class Qwen35EncoderCheckpointLoader(ModelLoader): + """Loads a single-file Qwen3.5 text encoder and the vendored Qwen3.5 tokenizer.""" + + def _load_model(self, config: AnyModelConfig, submodel_type: Optional[SubModelType] = None) -> AnyModel: + if not isinstance(config, Qwen35Encoder_Checkpoint_Config): + raise ValueError("Only Qwen35Encoder_Checkpoint_Config models are supported here.") + + match submodel_type: + case SubModelType.TextEncoder: + return self._load_text_encoder(config) + case SubModelType.Tokenizer: + from invokeai.backend.qwen3_5.qwen3_5_encoder import load_bundled_qwen3_5_tokenizer + + # Single-file checkpoints ship no tokenizer. + return load_bundled_qwen3_5_tokenizer() + + raise ValueError( + f"Only TextEncoder and Tokenizer submodels are supported. Received: {submodel_type.value if submodel_type else 'None'}" + ) + + def _load_text_encoder(self, config: Qwen35Encoder_Checkpoint_Config) -> AnyModel: + from safetensors.torch import load_file + + from invokeai.backend.qwen3_5.qwen3_5_encoder import Qwen35Encoder, qwen3_5_4b_text_config + + model_path = Path(config.path) + target_device = TorchDevice.choose_torch_device() + model_dtype = TorchDevice.choose_bfloat16_safe_dtype(target_device) + + sd = load_file(model_path) + # Anima-3.8B's encoder ships raw fp8 projections (no scale), which the cast below handles + # exactly. A scaled, int8 or nvfp4 build would load *wrong* here rather than fail, so refuse it + # until one exists to support. + reject_quantized_side_channel(sd, f"Qwen3.5 encoder checkpoint {model_path.name}") + prefix = qwen3_5_key_prefix(sd) + if prefix is None: + raise ValueError(f"{model_path.name} is not a Qwen3.5 language model checkpoint.") + sd = _language_model_state_dict(sd, prefix) + + text_config = qwen3_5_4b_text_config() + with accelerate.init_empty_weights(): + model = Qwen35Encoder(text_config) + + # Anima-3.8B's encoder ships its last layer without the MLP sub-block (and its norm), because + # nothing it was used for reads past that layer's attention. Give such a layer an MLP that adds + # nothing, so the layer computes exactly what the file encodes instead of loading half-empty. + for index, layer in enumerate(model.layers): + if not any(key.startswith(f"layers.{index}.mlp.") for key in sd): + layer.mlp = _ZeroMLP() + layer.post_attention_layernorm = nn.Identity() + self._logger.info(f"Qwen3.5 encoder: layer {index} ships without an MLP; running it attention-only.") + + keep_fp8 = self._keep_fp8_weights(config, SubModelType.TextEncoder) + skip_patterns = _model_declared_skip_patterns(model) + reserve_for_load( + self._ram_cache.make_room, + sd, + model_dtype, + keep_fp8=keep_fp8, + model=model, + skip_patterns=skip_patterns, + fp8_layers={}, + nvfp4_payloads={}, + ) + kept = cast_state_dict(sd, model_dtype, keep_fp8=keep_fp8, model=model, skip_patterns=skip_patterns) + + result = model.load_state_dict(sd, strict=False, assign=True) + log_unexpected_keys("Qwen3.5 encoder checkpoint", result.unexpected_keys) + reject_incomplete_load(model, what=f"Qwen3.5 encoder checkpoint {model_path.name}") + sd.clear() + + if kept: + self._logger.info(f"Qwen3.5 encoder: kept {kept} raw fp8 weight(s) quantized ({self._fp8_kept_reason()}).") + model.eval() + return self._apply_fp8_layerwise_casting(model, config, SubModelType.TextEncoder) diff --git a/invokeai/backend/model_manager/starter_models/__init__.py b/invokeai/backend/model_manager/starter_models/__init__.py index b8cc4198fea..25900a1c62b 100644 --- a/invokeai/backend/model_manager/starter_models/__init__.py +++ b/invokeai/backend/model_manager/starter_models/__init__.py @@ -10,6 +10,9 @@ """ from invokeai.backend.model_manager.starter_models.anima import ( + anima_2_9b, + anima_2_9b_int8, + anima_3_8b, anima_base, anima_lllite_depth_preview3, anima_lllite_inpainting, @@ -17,6 +20,7 @@ anima_lllite_pose_preview3, anima_lllite_scribble_preview3, anima_lllite_sketch, + anima_qwen3_5_encoder, anima_vae, ) from invokeai.backend.model_manager.starter_models.cogview4 import cogview4 @@ -496,7 +500,11 @@ alibabacloud_wan26_t2i, alibabacloud_qwen_image_edit_max, anima_base, + anima_2_9b, + anima_2_9b_int8, + anima_3_8b, anima_qwen3_encoder, + anima_qwen3_5_encoder, anima_vae, anima_lllite_inpainting, anima_lllite_sketch, diff --git a/invokeai/backend/model_manager/starter_models/anima.py b/invokeai/backend/model_manager/starter_models/anima.py index 417a7c0c177..768b1daf796 100644 --- a/invokeai/backend/model_manager/starter_models/anima.py +++ b/invokeai/backend/model_manager/starter_models/anima.py @@ -3,9 +3,11 @@ from invokeai.backend.model_manager.starter_models.common import anima_qwen3_encoder from invokeai.backend.model_manager.starter_models.types import StarterModel from invokeai.backend.model_manager.taxonomy import ( + AnimaVariantType, BaseModelType, ModelFormat, ModelType, + Qwen35VariantType, ) anima_vae = StarterModel( @@ -24,9 +26,63 @@ description="Anima Base 1.0 - 2B parameter anime-focused text-to-image model built on Cosmos Predict2 DiT. ~4.5GB", type=ModelType.Main, format=ModelFormat.Checkpoint, + variant=AnimaVariantType.Qwen3, dependencies=[anima_qwen3_encoder, anima_vae], ) +_DERIVATIVE_LICENSE = ( + "A derivative of Anima under the CircleStone Labs Non-Commercial License: the model may not be used " + "commercially (generated images may)." +) + +anima_2_9b = StarterModel( + name="Anima-2.9B (Preview v1)", + base=BaseModelType.Anima, + source="https://huggingface.co/Gazingstars123/Anima-2.9B/resolve/main/Anima-2.9B-preview-v1.safetensors", + description="Community finetune of Anima Base, depth-expanded from 28 to 40 blocks and trained on 1.7M more " + f"anime/illustration images (knowledge cutoff July 2026). Recommended: 28-50 steps, CFG 3.5-5. " + f"{_DERIVATIVE_LICENSE} ~5.8GB", + type=ModelType.Main, + format=ModelFormat.Checkpoint, + variant=AnimaVariantType.Qwen3, + dependencies=[anima_qwen3_encoder, anima_vae], +) + +anima_2_9b_int8 = StarterModel( + name="Anima-2.9B (Preview v1, int8)", + base=BaseModelType.Anima, + source="https://huggingface.co/Gazingstars123/Anima-2.9B/resolve/main/Anima-2.9B-preview-v1_int8_convrot.safetensors", + description="Anima-2.9B in ComfyUI's int8_tensorwise format: half the memory of the bf16 build, dequantized " + f"per forward (no speedup). {_DERIVATIVE_LICENSE} ~3.1GB", + type=ModelType.Main, + format=ModelFormat.Checkpoint, + variant=AnimaVariantType.Qwen3, + dependencies=[anima_qwen3_encoder, anima_vae], +) + +anima_qwen3_5_encoder = StarterModel( + name="Anima-3.8B Qwen3.5 4B Encoder", + base=BaseModelType.Any, + source="https://huggingface.co/lylogummy/Anima-3.8B/resolve/main/text_encoders/qwen35_4b.safetensors", + description="Qwen3.5 4B text encoder (raw fp8) that Anima-3.8B's semantic connector reads. Apache-2.0. ~4.8GB", + type=ModelType.Qwen35Encoder, + format=ModelFormat.Checkpoint, + variant=Qwen35VariantType.Qwen35_4B, +) + +anima_3_8b = StarterModel( + name="Anima-3.8B (v1.1)", + base=BaseModelType.Anima, + source="https://huggingface.co/lylogummy/Anima-3.8B/resolve/main/difussion_models/Anima-3.8B-v1.1.safetensors", + description="Community expansion of Anima-2.9B to 52 blocks with a bundled Qwen3.5 semantic connector, for " + "prompt adherence, multi-character binding and mixed natural-language/tag prompts. Needs both the Qwen3 0.6B " + f"and the Qwen3.5 4B encoder. Recommended: 28-50 steps, CFG 4-7. {_DERIVATIVE_LICENSE} ~8.8GB", + type=ModelType.Main, + format=ModelFormat.Checkpoint, + variant=AnimaVariantType.Qwen35, + dependencies=[anima_qwen3_encoder, anima_qwen3_5_encoder, anima_vae], +) + anima_lllite_inpainting = StarterModel( name="Anima LLLite Inpainting", base=BaseModelType.Anima, diff --git a/invokeai/backend/model_manager/taxonomy.py b/invokeai/backend/model_manager/taxonomy.py index 54821ab52fa..e4cfdb3da10 100644 --- a/invokeai/backend/model_manager/taxonomy.py +++ b/invokeai/backend/model_manager/taxonomy.py @@ -94,6 +94,7 @@ class ModelType(str, Enum): Qwen3Encoder = "qwen3_encoder" QwenVLEncoder = "qwen_vl_encoder" Qwen3VLEncoder = "qwen3_vl_encoder" + Qwen35Encoder = "qwen3_5_encoder" MistralEncoder = "mistral_encoder" WanT5Encoder = "wan_t5_encoder" Gemma2Encoder = "gemma2_encoder" @@ -282,6 +283,34 @@ class Qwen3VLVariantType(str, Enum): of its layers for a 53248-wide feature vector.""" +class Qwen35VariantType(str, Enum): + """Qwen3.5 text encoder variants, by language-model width. + + A family of its own: Qwen3.5 interleaves Gated DeltaNet linear-attention layers with gated full + attention, so its checkpoints share neither architecture nor key layout with `Qwen3VariantType`, + even at the same width. + """ + + Qwen35_4B = "qwen3_5_4b" + """Qwen3.5 4B (hidden_size=2560, 32 layers). Anima-3.8B's semantic encoder, read at layers 7/15/23/31.""" + + +class AnimaVariantType(str, Enum): + """Anima model variants, by how the transformer is conditioned. + + The depth of the DiT (28 blocks for the official release, 40 for Anima-2.9B, 52 for Anima-3.8B) + is not a variant: the loader reads it from the checkpoint and nothing else depends on it. + """ + + Qwen3 = "anima_qwen3" + """Conditioned on Qwen3 0.6B through the native LLM adapter. Every release by CircleStone Labs + and Anima-2.9B.""" + + Qwen35 = "anima_qwen35" + """Additionally conditioned on Qwen3.5 4B through a bundled, timestep-aware semantic connector + (Anima-3.8B v1.1). Needs both encoders, and the connector runs at every denoising step.""" + + class MiniMaxH3VariantType(str, Enum): """MiniMax H3 model variants (task-specific transformer checkpoints sharing every other component).""" @@ -446,6 +475,8 @@ class FluxLoRAFormat(str, Enum): WanLoRAVariantType, Qwen3VariantType, Qwen3VLVariantType, + Qwen35VariantType, + AnimaVariantType, Krea2VariantType, MiniMaxH3VariantType, LTX2VariantType, @@ -463,6 +494,8 @@ class FluxLoRAFormat(str, Enum): | WanLoRAVariantType | Qwen3VariantType | Qwen3VLVariantType + | Qwen35VariantType + | AnimaVariantType | Krea2VariantType | MiniMaxH3VariantType | LTX2VariantType @@ -479,6 +512,8 @@ class FluxLoRAFormat(str, Enum): | WanLoRAVariantType | Qwen3VariantType | Qwen3VLVariantType + | Qwen35VariantType + | AnimaVariantType | Krea2VariantType | MiniMaxH3VariantType | LTX2VariantType diff --git a/invokeai/backend/patches/lora_conversions/anima_lora_conversion_utils.py b/invokeai/backend/patches/lora_conversions/anima_lora_conversion_utils.py index b55a96dca75..1707a2a2122 100644 --- a/invokeai/backend/patches/lora_conversions/anima_lora_conversion_utils.py +++ b/invokeai/backend/patches/lora_conversions/anima_lora_conversion_utils.py @@ -10,10 +10,11 @@ """ import re -from typing import Dict +from typing import Dict, Optional import torch +from invokeai.backend.anima.block_layout import KNOWN_DEPTHS, adapter_block_positions from invokeai.backend.patches.layers.base_layer_patch import BaseLayerPatch from invokeai.backend.patches.layers.utils import any_lora_layer_from_state_dict from invokeai.backend.patches.lora_conversions.anima_lora_constants import ( @@ -298,3 +299,31 @@ def lora_model_from_anima_state_dict(state_dict: Dict[str, torch.Tensor], alpha: layers[final_key] = layer return ModelPatchRaw(layers=layers) + + +_TRANSFORMER_BLOCK_KEY_RE = re.compile(rf"^{re.escape(ANIMA_LORA_TRANSFORMER_PREFIX)}blocks\.(\d+)\.") + + +def anima_lora_for_depth(patch: ModelPatchRaw, target_depth: int) -> tuple[ModelPatchRaw, Optional[int]]: + """The LoRA with its DiT block keys moved to where `target_depth`'s model keeps the blocks it was trained on. + + A LoRA trained on Anima base, applied to a depth-expanded finetune (see `invokeai.backend.anima.block_layout`), + would otherwise patch the wrong blocks from the third one on. Returns the patch -- a new one sharing the layers, + so the cached LoRA stays as it is -- and the depth it was moved from, or the patch itself and None when nothing + moves. Only `blocks.N.` keys of the transformer move; the LLM adapter's own `llm_adapter.blocks.N.` and the text + encoder's keys do not. + """ + indices = [int(match.group(1)) for key in patch.layers if (match := _TRANSFORMER_BLOCK_KEY_RE.match(key))] + if not indices: + return patch, None + positions = adapter_block_positions(max(indices), target_depth) + if positions is None: + return patch, None + + def moved(key: str) -> str: + return _TRANSFORMER_BLOCK_KEY_RE.sub( + lambda m: f"{ANIMA_LORA_TRANSFORMER_PREFIX}blocks.{positions[int(m.group(1))]}.", key, count=1 + ) + + source_depth = next(depth for depth in KNOWN_DEPTHS if depth > max(indices)) + return ModelPatchRaw(layers={moved(key): layer for key, layer in patch.layers.items()}), source_depth diff --git a/invokeai/backend/qwen3_5/__init__.py b/invokeai/backend/qwen3_5/__init__.py new file mode 100644 index 00000000000..64a0c29154b --- /dev/null +++ b/invokeai/backend/qwen3_5/__init__.py @@ -0,0 +1 @@ +"""Qwen3.5 text backbone used as a hidden-state encoder (Anima-3.8B's semantic encoder).""" diff --git a/invokeai/backend/qwen3_5/qwen3_5_4b_text_config.json b/invokeai/backend/qwen3_5/qwen3_5_4b_text_config.json new file mode 100644 index 00000000000..2160becee10 --- /dev/null +++ b/invokeai/backend/qwen3_5/qwen3_5_4b_text_config.json @@ -0,0 +1,76 @@ +{ + "attention_bias": false, + "attention_dropout": 0.0, + "attn_output_gate": true, + "dtype": "bfloat16", + "eos_token_id": 248044, + "full_attention_interval": 4, + "head_dim": 256, + "hidden_act": "silu", + "hidden_size": 2560, + "initializer_range": 0.02, + "intermediate_size": 9216, + "layer_types": [ + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention", + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention", + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention", + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention", + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention", + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention", + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention", + "linear_attention", + "linear_attention", + "linear_attention", + "full_attention" + ], + "linear_conv_kernel_dim": 4, + "linear_key_head_dim": 128, + "linear_num_key_heads": 16, + "linear_num_value_heads": 32, + "linear_value_head_dim": 128, + "max_position_embeddings": 262144, + "mlp_only_layers": [], + "model_type": "qwen3_5_text", + "mtp_num_hidden_layers": 1, + "mtp_use_dedicated_embeddings": false, + "num_attention_heads": 16, + "num_hidden_layers": 32, + "num_key_value_heads": 4, + "rms_norm_eps": 1e-06, + "tie_word_embeddings": true, + "use_cache": true, + "vocab_size": 248320, + "mamba_ssm_dtype": "float32", + "rope_parameters": { + "mrope_interleaved": true, + "mrope_section": [ + 11, + 11, + 10 + ], + "rope_type": "default", + "rope_theta": 10000000, + "partial_rotary_factor": 0.25 + } +} diff --git a/invokeai/backend/qwen3_5/qwen3_5_encoder.py b/invokeai/backend/qwen3_5/qwen3_5_encoder.py new file mode 100644 index 00000000000..580d8f7b205 --- /dev/null +++ b/invokeai/backend/qwen3_5/qwen3_5_encoder.py @@ -0,0 +1,122 @@ +"""Qwen3.5's text backbone, run as an encoder that returns intermediate hidden states. + +Qwen3.5 interleaves Gated DeltaNet linear-attention layers with gated full attention (every fourth +layer). The decoder layers come from `transformers` (`qwen3_5`); this module only owns the loop, so +it can stop at the deepest layer a caller reads and hand back that layer's output without the final +norm `Qwen3_5TextModel.forward` applies. + +Anima-3.8B reads layers 7, 15, 23 and 31 of the 4B model. Its encoder checkpoint +(`qwen35_4b.safetensors`) is not a stock export: it ships layer 31 without its MLP and replaces the +LM head with a 2560->1024 projection Anima never reads. The connector was trained on what that file +computes, so `forward` can run the deepest requested layer attention-only -- matching the file +exactly, and making a stock checkpoint's layer-31 MLP irrelevant rather than wrong. + +The tokenizer is vendored from `Qwen/Qwen3.5-4B` (Apache-2.0) at revision +851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, the revision the reference ComfyUI extension ships, so a +single-file install encodes offline. +""" + +import json +from functools import lru_cache +from pathlib import Path +from typing import Any + +import torch +from torch import nn +from transformers import PreTrainedTokenizerBase +from transformers.models.qwen3_5.configuration_qwen3_5 import Qwen3_5TextConfig +from transformers.models.qwen3_5.modeling_qwen3_5 import Qwen3_5DecoderLayer, Qwen3_5TextRotaryEmbedding + +from invokeai.backend.util.bundled_tokenizer import load_gzipped_tokenizer_dir + +_PACKAGE_DIR = Path(__file__).parent +_TOKENIZER_DIR = _PACKAGE_DIR / "tokenizer" +QWEN3_5_4B_TEXT_CONFIG_PATH = _PACKAGE_DIR / "qwen3_5_4b_text_config.json" + + +@lru_cache(maxsize=1) +def load_bundled_qwen3_5_tokenizer() -> PreTrainedTokenizerBase: + """Load the vendored Qwen3.5 fast tokenizer. Result is cached for the process.""" + return load_gzipped_tokenizer_dir(_TOKENIZER_DIR) + + +def qwen3_5_4b_text_config() -> Qwen3_5TextConfig: + """The Qwen3.5 4B text config, as published in `Qwen/Qwen3.5-4B`'s `config.json`.""" + with open(QWEN3_5_4B_TEXT_CONFIG_PATH, "r", encoding="utf-8") as f: + raw: dict[str, Any] = json.load(f) + config = Qwen3_5TextConfig(**raw) + # The decoder layers dispatch attention through this; a config built by hand leaves it unset. + config._attn_implementation = "sdpa" + return config + + +class Qwen35Encoder(nn.Module): + """Qwen3.5 decoder stack that returns the outputs of selected layers. + + Holds `embed_tokens` and `layers` under the names a checkpoint uses, so a bare-keyed state dict + loads directly. Weights for layers past the deepest one any caller reads may be absent: they are + never run. + """ + + def __init__(self, config: Qwen3_5TextConfig): + super().__init__() + self.config = config + self.embed_tokens = nn.Embedding(config.vocab_size, config.hidden_size) + self.layers = nn.ModuleList([Qwen3_5DecoderLayer(config, i) for i in range(config.num_hidden_layers)]) + self.rotary_emb = Qwen3_5TextRotaryEmbedding(config=config) + + @property + def dtype(self) -> torch.dtype: + return self.embed_tokens.weight.dtype + + @torch.no_grad() + def forward( + self, + input_ids: torch.Tensor, + layer_indices: tuple[int, ...], + last_layer_attention_only: bool = False, + ) -> list[torch.Tensor]: + """Run the stack up to `max(layer_indices)` and return those layers' outputs, in order. + + Args: + input_ids: Token IDs, unpadded. Shape: (batch, seq_len). + layer_indices: Zero-based decoder layers whose output to return. + last_layer_attention_only: Run the deepest requested layer without its MLP sub-block and + return its post-attention residual. What Anima-3.8B's encoder checkpoint computes. + + Returns: + One tensor of shape (batch, seq_len, hidden_size) per entry of `layer_indices`. + """ + if not layer_indices: + raise ValueError("layer_indices must not be empty") + last = max(layer_indices) + if last >= len(self.layers): + raise ValueError(f"Layer {last} requested from a {len(self.layers)}-layer Qwen3.5 model") + + hidden_states = self.embed_tokens(input_ids) + position_ids = torch.arange(input_ids.shape[1], device=input_ids.device).unsqueeze(0) + # Text-only: the three mRoPE axes share one position, which reduces to plain 1-D RoPE. + position_embeddings = self.rotary_emb(hidden_states, position_ids.unsqueeze(0).expand(3, -1, -1)) + + outputs: dict[int, torch.Tensor] = {} + for index in range(last + 1): + layer = self.layers[index] + if index == last and last_layer_attention_only: + hidden_states = hidden_states + self._token_mixer(layer, hidden_states, position_embeddings) + else: + # No mask: a single unpadded prompt, so full attention is plain causal attention and + # the linear-attention layers have nothing to zero. + hidden_states = layer(hidden_states, position_embeddings=position_embeddings, attention_mask=None) + if index in layer_indices: + outputs[index] = hidden_states + return [outputs[index] for index in layer_indices] + + @staticmethod + def _token_mixer( + layer: Qwen3_5DecoderLayer, hidden_states: torch.Tensor, position_embeddings: tuple[torch.Tensor, torch.Tensor] + ) -> torch.Tensor: + normed = layer.input_layernorm(hidden_states) + if layer.layer_type == "linear_attention": + return layer.linear_attn(hidden_states=normed, cache_params=None, attention_mask=None) + mixed, _ = layer.self_attn(hidden_states=normed, position_embeddings=position_embeddings, attention_mask=None) + return mixed diff --git a/invokeai/backend/qwen3_5/tokenizer/tokenizer.json.gz b/invokeai/backend/qwen3_5/tokenizer/tokenizer.json.gz new file mode 100644 index 00000000000..c541b6937e8 Binary files /dev/null and b/invokeai/backend/qwen3_5/tokenizer/tokenizer.json.gz differ diff --git a/invokeai/backend/qwen3_5/tokenizer/tokenizer_config.json b/invokeai/backend/qwen3_5/tokenizer/tokenizer_config.json new file mode 100644 index 00000000000..eda48d3e75a --- /dev/null +++ b/invokeai/backend/qwen3_5/tokenizer/tokenizer_config.json @@ -0,0 +1,305 @@ +{ + "add_prefix_space": false, + "added_tokens_decoder": { + "248044": { + "content": "<|endoftext|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248045": { + "content": "<|im_start|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248046": { + "content": "<|im_end|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248047": { + "content": "<|object_ref_start|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248048": { + "content": "<|object_ref_end|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248049": { + "content": "<|box_start|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248050": { + "content": "<|box_end|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248051": { + "content": "<|quad_start|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248052": { + "content": "<|quad_end|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248053": { + "content": "<|vision_start|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248054": { + "content": "<|vision_end|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248055": { + "content": "<|vision_pad|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248056": { + "content": "<|image_pad|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248057": { + "content": "<|video_pad|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248058": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248059": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248060": { + "content": "<|fim_prefix|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248061": { + "content": "<|fim_middle|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248062": { + "content": "<|fim_suffix|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248063": { + "content": "<|fim_pad|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248064": { + "content": "<|repo_name|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248065": { + "content": "<|file_sep|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248066": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248067": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248068": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248069": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": false + }, + "248070": { + "content": "<|audio_start|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248071": { + "content": "<|audio_end|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248072": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248073": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248074": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248075": { + "content": "", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + }, + "248076": { + "content": "<|audio_pad|>", + "lstrip": false, + "normalized": false, + "rstrip": false, + "single_word": false, + "special": true + } + }, + "additional_special_tokens": [ + "<|im_start|>", + "<|im_end|>", + "<|object_ref_start|>", + "<|object_ref_end|>", + "<|box_start|>", + "<|box_end|>", + "<|quad_start|>", + "<|quad_end|>", + "<|vision_start|>", + "<|vision_end|>", + "<|vision_pad|>", + "<|image_pad|>", + "<|video_pad|>" + ], + "bos_token": null, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if messages[0].role == 'system' %}\n {%- set content = render_content(messages[0].content, false, true)|trim %}\n {%- if content %}\n {{- '\\n\\n' + content }}\n {%- endif %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if messages[0].role == 'system' %}\n {%- set content = render_content(messages[0].content, false, true)|trim %}\n {{- '<|im_start|>system\\n' + content + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if ns.multi_step_tool %}\n {{- raise_exception('No user query found in messages.') }}\n{%- endif %}\n{%- for message in messages %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" %}\n {%- if not loop.first %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- endif %}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- else %}\n {%- if '' in content %}\n {%- set reasoning_content = content.split('')[0].rstrip('\\n').split('')[-1].lstrip('\\n') %}\n {%- set content = content.split('')[-1].lstrip('\\n') %}\n {%- endif %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is defined %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}", + "clean_up_tokenization_spaces": false, + "eos_token": "<|im_end|>", + "errors": "replace", + "model_max_length": 262144, + "pad_token": "<|endoftext|>", + "split_special_tokens": false, + "tokenizer_class": "Qwen2Tokenizer", + "unk_token": null, + "add_bos_token": false, + "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+", + "extra_special_tokens": { + "audio_bos_token": "<|audio_start|>", + "audio_eos_token": "<|audio_end|>", + "audio_token": "<|audio_pad|>", + "image_token": "<|image_pad|>", + "video_token": "<|video_pad|>", + "vision_bos_token": "<|vision_start|>", + "vision_eos_token": "<|vision_end|>" + } +} \ No newline at end of file diff --git a/invokeai/backend/stable_diffusion/diffusion/conditioning_data.py b/invokeai/backend/stable_diffusion/diffusion/conditioning_data.py index bbdfd186eb6..a3e357463f2 100644 --- a/invokeai/backend/stable_diffusion/diffusion/conditioning_data.py +++ b/invokeai/backend/stable_diffusion/diffusion/conditioning_data.py @@ -173,11 +173,22 @@ class AnimaConditioningInfo: t5xxl_weights: Optional[torch.Tensor] = None """Per-token weights for prompt weighting. Shape: (seq_len,). None means uniform weight.""" + qwen35_states: Optional[torch.Tensor] = None + """Qwen3.5 4B hidden states for Anima-3.8B's semantic connector, one row per tapped layer (7, 15, + 23, 31). Shape: (num_layers, seq_len, 2560). None when the prompt was encoded without Qwen3.5.""" + + qwen35_mask: Optional[torch.Tensor] = None + """True for valid Qwen3.5 tokens. Shape: (seq_len,). All False for an empty prompt.""" + def to(self, device: torch.device | None = None, dtype: torch.dtype | None = None): self.qwen3_embeds = self.qwen3_embeds.to(device=device, dtype=dtype) self.t5xxl_ids = self.t5xxl_ids.to(device=device) if self.t5xxl_weights is not None: self.t5xxl_weights = self.t5xxl_weights.to(device=device, dtype=dtype) + if self.qwen35_states is not None: + self.qwen35_states = self.qwen35_states.to(device=device, dtype=dtype) + if self.qwen35_mask is not None: + self.qwen35_mask = self.qwen35_mask.to(device=device) return self diff --git a/invokeai/frontend/api/openapi.json b/invokeai/frontend/api/openapi.json index 658f57011d1..6a3c8162b47 100644 --- a/invokeai/frontend/api/openapi.json +++ b/invokeai/frontend/api/openapi.json @@ -1679,6 +1679,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -2146,6 +2149,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -2613,6 +2619,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -3125,6 +3134,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -3664,6 +3676,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -5060,6 +5075,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -18444,7 +18462,7 @@ "category": "model", "class": "invocation", "classification": "prototype", - "description": "Loads an Anima model, outputting its submodels.\n\nAnima uses:\n- Transformer: Cosmos Predict2 DiT + LLM Adapter (from single-file checkpoint)\n- Qwen3 Encoder: Qwen3 0.6B (standalone single-file)\n- VAE: AutoencoderKLQwenImage / Wan 2.1 VAE (standalone single-file)\n\nThe T5-XXL tokenizer needed for LLM Adapter token IDs is bundled in the package,\nso no T5-XXL encoder model needs to be installed.", + "description": "Loads an Anima model, outputting its submodels.\n\nAnima uses:\n- Transformer: Cosmos Predict2 DiT + LLM Adapter (from single-file checkpoint)\n- Qwen3 Encoder: Qwen3 0.6B (standalone single-file)\n- VAE: AutoencoderKLQwenImage / Wan 2.1 VAE (standalone single-file)\n- Qwen3.5 Encoder: Qwen3.5 4B, for Anima-3.8B only, whose bundled semantic connector reads it\n\nThe T5-XXL tokenizer needed for LLM Adapter token IDs is bundled in the package,\nso no T5-XXL encoder model needs to be installed.", "node_pack": "invokeai", "properties": { "id": { @@ -18500,6 +18518,24 @@ "title": "Qwen3 Encoder", "ui_model_type": ["qwen3_encoder"] }, + "qwen3_5_encoder_model": { + "anyOf": [ + { + "$ref": "#/components/schemas/ModelIdentifierField" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Standalone Qwen3.5 4B Encoder model. Required by Anima-3.8B, ignored by every other Anima model.", + "field_kind": "input", + "input": "direct", + "orig_default": null, + "orig_required": false, + "title": "Qwen3.5 Encoder", + "ui_model_type": ["qwen3_5_encoder"] + }, "type": { "const": "anima_model_loader", "default": "anima_model_loader", @@ -18512,7 +18548,7 @@ "tags": ["model", "anima"], "title": "Main Model - Anima", "type": "object", - "version": "1.4.0", + "version": "1.5.0", "output": { "$ref": "#/components/schemas/AnimaModelLoaderOutput" } @@ -18535,6 +18571,21 @@ "title": "Qwen3 Encoder", "ui_hidden": false }, + "qwen3_5_encoder": { + "anyOf": [ + { + "$ref": "#/components/schemas/Qwen35EncoderField" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Qwen3.5 tokenizer and text encoder. Set only for an Anima-3.8B model.", + "field_kind": "output", + "title": "Qwen3.5 Encoder", + "ui_hidden": false + }, "vae": { "$ref": "#/components/schemas/VAEField", "description": "VAE", @@ -18550,7 +18601,7 @@ "type": "string" } }, - "required": ["output_meta", "transformer", "qwen3_encoder", "vae", "type", "type"], + "required": ["output_meta", "transformer", "qwen3_encoder", "qwen3_5_encoder", "vae", "type", "type"], "title": "AnimaModelLoaderOutput", "type": "object" }, @@ -18634,6 +18685,23 @@ "orig_default": null, "orig_required": false }, + "qwen3_5_encoder": { + "anyOf": [ + { + "$ref": "#/components/schemas/Qwen35EncoderField" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Qwen3.5 tokenizer and text encoder. Connect it for Anima-3.8B, which reads both encoders.", + "field_kind": "input", + "input": "connection", + "orig_default": null, + "orig_required": false, + "title": "Qwen3.5 Encoder" + }, "type": { "const": "anima_text_encoder", "default": "anima_text_encoder", @@ -18646,11 +18714,17 @@ "tags": ["prompt", "conditioning", "anima"], "title": "Prompt - Anima", "type": "object", - "version": "1.4.0", + "version": "1.5.0", "output": { "$ref": "#/components/schemas/AnimaConditioningOutput" } }, + "AnimaVariantType": { + "type": "string", + "enum": ["anima_qwen3", "anima_qwen35"], + "title": "AnimaVariantType", + "description": "Anima model variants, by how the transformer is conditioned.\n\nThe depth of the DiT (28 blocks for the official release, 40 for Anima-2.9B, 52 for Anima-3.8B)\nis not a variant: the loader reads it from the checkpoint and nothing else depends on it." + }, "AnyModelConfig": { "oneOf": [ { @@ -18953,6 +19027,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -69439,6 +69516,10 @@ "const": "checkpoint", "title": "Format", "default": "checkpoint" + }, + "variant": { + "$ref": "#/components/schemas/AnimaVariantType", + "description": "Which text encoders the model is conditioned on." } }, "type": "object", @@ -69459,7 +69540,8 @@ "default_settings", "config_path", "base", - "format" + "format", + "variant" ], "title": "Main_Checkpoint_Anima_Config", "description": "Model config for Anima single-file checkpoint models (safetensors).\n\nAnima is built on NVIDIA Cosmos Predict2 DiT with a custom LLM Adapter\nthat bridges Qwen3 0.6B text encoder outputs to the DiT." @@ -83468,6 +83550,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -84181,6 +84266,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -84779,6 +84867,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -85233,6 +85324,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -85684,6 +85778,12 @@ { "$ref": "#/components/schemas/Qwen3VLVariantType" }, + { + "$ref": "#/components/schemas/Qwen35VariantType" + }, + { + "$ref": "#/components/schemas/AnimaVariantType" + }, { "$ref": "#/components/schemas/Krea2VariantType" }, @@ -85839,6 +85939,7 @@ "qwen3_encoder", "qwen_vl_encoder", "qwen3_vl_encoder", + "qwen3_5_encoder", "mistral_encoder", "wan_t5_encoder", "gemma2_encoder", @@ -86168,6 +86269,9 @@ { "$ref": "#/components/schemas/Qwen3VLEncoder_Qwen3VLEncoder_Config" }, + { + "$ref": "#/components/schemas/Qwen35Encoder_Checkpoint_Config" + }, { "$ref": "#/components/schemas/Qwen3Encoder_Qwen3Encoder_Config" }, @@ -89944,6 +90048,182 @@ "title": "QueueItemsRetriedEvent", "type": "object" }, + "Qwen35EncoderField": { + "description": "Field for the Qwen3.5 text encoder Anima-3.8B's semantic connector reads.", + "properties": { + "tokenizer": { + "$ref": "#/components/schemas/ModelIdentifierField", + "description": "Info to load tokenizer submodel" + }, + "text_encoder": { + "$ref": "#/components/schemas/ModelIdentifierField", + "description": "Info to load text_encoder submodel" + } + }, + "required": ["tokenizer", "text_encoder"], + "title": "Qwen35EncoderField", + "type": "object" + }, + "Qwen35Encoder_Checkpoint_Config": { + "properties": { + "key": { + "type": "string", + "title": "Key", + "description": "A unique key for this model." + }, + "hash": { + "type": "string", + "title": "Hash", + "description": "The hash of the model file(s)." + }, + "path": { + "type": "string", + "title": "Path", + "description": "Path to the model on the filesystem. Relative paths are relative to the Invoke root directory." + }, + "file_size": { + "type": "integer", + "title": "File Size", + "description": "The size of the model in bytes." + }, + "name": { + "type": "string", + "title": "Name", + "description": "Name of the model." + }, + "description": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Description", + "description": "Model description" + }, + "source": { + "type": "string", + "title": "Source", + "description": "The original source of the model (path, URL or repo_id)." + }, + "source_type": { + "$ref": "#/components/schemas/ModelSourceType", + "description": "The type of source" + }, + "source_api_response": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Source Api Response", + "description": "The original API response from the source, as stringified JSON." + }, + "source_url": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Source Url", + "description": "Optional URL for the model (e.g. download page or model page)." + }, + "cover_image": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Cover Image", + "description": "Url for image to preview model" + }, + "config_path": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Config Path", + "description": "Path to the config for this model, if any." + }, + "base": { + "type": "string", + "const": "any", + "title": "Base", + "default": "any" + }, + "type": { + "type": "string", + "const": "qwen3_5_encoder", + "title": "Type", + "default": "qwen3_5_encoder" + }, + "format": { + "type": "string", + "const": "checkpoint", + "title": "Format", + "default": "checkpoint" + }, + "cpu_only": { + "anyOf": [ + { + "type": "boolean" + }, + { + "type": "null" + } + ], + "title": "Cpu Only", + "description": "Whether this model should run on CPU only" + }, + "variant": { + "$ref": "#/components/schemas/Qwen35VariantType", + "description": "Qwen3.5 model size variant" + } + }, + "type": "object", + "required": [ + "key", + "hash", + "path", + "file_size", + "name", + "description", + "source", + "source_type", + "source_api_response", + "source_url", + "cover_image", + "config_path", + "base", + "type", + "format", + "cpu_only", + "variant" + ], + "title": "Qwen35Encoder_Checkpoint_Config", + "description": "Configuration for single-file Qwen3.5 text encoders (safetensors)." + }, + "Qwen35VariantType": { + "type": "string", + "enum": ["qwen3_5_4b"], + "title": "Qwen35VariantType", + "description": "Qwen3.5 text encoder variants, by language-model width.\n\nA family of its own: Qwen3.5 interleaves Gated DeltaNet linear-attention layers with gated full\nattention, so its checkpoints share neither architecture nor key layout with `Qwen3VariantType`,\neven at the same width." + }, "Qwen3EncoderField": { "description": "Field for Qwen3 text encoder used by Z-Image models.", "properties": { @@ -98741,6 +99021,12 @@ { "$ref": "#/components/schemas/Qwen3VLVariantType" }, + { + "$ref": "#/components/schemas/Qwen35VariantType" + }, + { + "$ref": "#/components/schemas/AnimaVariantType" + }, { "$ref": "#/components/schemas/Krea2VariantType" }, @@ -98925,6 +99211,12 @@ { "$ref": "#/components/schemas/Qwen3VLVariantType" }, + { + "$ref": "#/components/schemas/Qwen35VariantType" + }, + { + "$ref": "#/components/schemas/AnimaVariantType" + }, { "$ref": "#/components/schemas/Krea2VariantType" }, @@ -100193,6 +100485,12 @@ { "$ref": "#/components/schemas/Qwen3VLVariantType" }, + { + "$ref": "#/components/schemas/Qwen35VariantType" + }, + { + "$ref": "#/components/schemas/AnimaVariantType" + }, { "$ref": "#/components/schemas/Krea2VariantType" }, diff --git a/invokeai/frontend/api/schema.ts b/invokeai/frontend/api/schema.ts index c4dbf1c0622..fddf71a2711 100644 --- a/invokeai/frontend/api/schema.ts +++ b/invokeai/frontend/api/schema.ts @@ -5206,6 +5206,7 @@ export type components = { * - Transformer: Cosmos Predict2 DiT + LLM Adapter (from single-file checkpoint) * - Qwen3 Encoder: Qwen3 0.6B (standalone single-file) * - VAE: AutoencoderKLQwenImage / Wan 2.1 VAE (standalone single-file) + * - Qwen3.5 Encoder: Qwen3.5 4B, for Anima-3.8B only, whose bundled semantic connector reads it * * The T5-XXL tokenizer needed for LLM Adapter token IDs is bundled in the package, * so no T5-XXL encoder model needs to be installed. @@ -5243,6 +5244,12 @@ export type components = { * @description Standalone Qwen3 0.6B Encoder model. */ qwen3_encoder_model: components["schemas"]["ModelIdentifierField"]; + /** + * Qwen3.5 Encoder + * @description Standalone Qwen3.5 4B Encoder model. Required by Anima-3.8B, ignored by every other Anima model. + * @default null + */ + qwen3_5_encoder_model?: components["schemas"]["ModelIdentifierField"] | null; /** * type * @default anima_model_loader @@ -5265,6 +5272,12 @@ export type components = { * @description Qwen3 tokenizer and text encoder */ qwen3_encoder: components["schemas"]["Qwen3EncoderField"]; + /** + * Qwen3.5 Encoder + * @description Qwen3.5 tokenizer and text encoder. Set only for an Anima-3.8B model. + * @default null + */ + qwen3_5_encoder: components["schemas"]["Qwen35EncoderField"] | null; /** * VAE * @description VAE @@ -5320,6 +5333,12 @@ export type components = { * @default null */ mask?: components["schemas"]["TensorField"] | null; + /** + * Qwen3.5 Encoder + * @description Qwen3.5 tokenizer and text encoder. Connect it for Anima-3.8B, which reads both encoders. + * @default null + */ + qwen3_5_encoder?: components["schemas"]["Qwen35EncoderField"] | null; /** * type * @default anima_text_encoder @@ -5327,7 +5346,16 @@ export type components = { */ type: "anima_text_encoder"; }; - AnyModelConfig: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + /** + * AnimaVariantType + * @description Anima model variants, by how the transformer is conditioned. + * + * The depth of the DiT (28 blocks for the official release, 40 for Anima-2.9B, 52 for Anima-3.8B) + * is not a variant: the loader reads it from the checkpoint and nothing else depends on it. + * @enum {string} + */ + AnimaVariantType: "anima_qwen3" | "anima_qwen35"; + AnyModelConfig: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; /** * AppVersion * @description App Version Response @@ -28787,6 +28815,8 @@ export type components = { * @constant */ format: "checkpoint"; + /** @description Which text encoders the model is conditioned on. */ + variant: components["schemas"]["AnimaVariantType"]; }; /** * Main_Checkpoint_ErnieImage_Config @@ -36051,7 +36081,7 @@ export type components = { * Config * @description The installed model's config */ - config: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + config: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; }; /** * ModelInstallDownloadProgressEvent @@ -36217,7 +36247,7 @@ export type components = { * Config Out * @description After successful installation, this will hold the configuration object. */ - config_out?: (components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]) | null; + config_out?: (components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]) | null; /** * Inplace * @description Leave model in its current location; otherwise install under models directory @@ -36303,7 +36333,7 @@ export type components = { * Config * @description The model's config */ - config: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + config: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; /** * @description The submodel type, if any * @default null @@ -36330,7 +36360,7 @@ export type components = { * Config * @description The model's config */ - config: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + config: components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; /** * @description The submodel type, if any * @default null @@ -36462,7 +36492,7 @@ export type components = { * Variant * @description The variant of the model. */ - variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; + variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Qwen35VariantType"] | components["schemas"]["AnimaVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; /** @description The prediction type of the model. */ prediction_type?: components["schemas"]["SchedulerPredictionType"] | null; /** @@ -36544,7 +36574,7 @@ export type components = { * @description Model type. * @enum {string} */ - ModelType: "onnx" | "main" | "vae" | "lora" | "control_lora" | "controlnet" | "embedding" | "ip_adapter" | "clip_vision" | "clip_embed" | "t2i_adapter" | "t5_encoder" | "qwen3_encoder" | "qwen_vl_encoder" | "qwen3_vl_encoder" | "mistral_encoder" | "wan_t5_encoder" | "gemma2_encoder" | "gemma4_encoder" | "spandrel_image_to_image" | "siglip" | "flux_redux" | "llava_onevision" | "prompt_enhancer" | "text_llm" | "external_image_generator" | "pid_decoder" | "ltx2_duration_head" | "unknown"; + ModelType: "onnx" | "main" | "vae" | "lora" | "control_lora" | "controlnet" | "embedding" | "ip_adapter" | "clip_vision" | "clip_embed" | "t2i_adapter" | "t5_encoder" | "qwen3_encoder" | "qwen_vl_encoder" | "qwen3_vl_encoder" | "qwen3_5_encoder" | "mistral_encoder" | "wan_t5_encoder" | "gemma2_encoder" | "gemma4_encoder" | "spandrel_image_to_image" | "siglip" | "flux_redux" | "llava_onevision" | "prompt_enhancer" | "text_llm" | "external_image_generator" | "pid_decoder" | "ltx2_duration_head" | "unknown"; /** * ModelVariantType * @description Variant type. @@ -36557,7 +36587,7 @@ export type components = { */ ModelsList: { /** Models */ - models: (components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"])[]; + models: (components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"])[]; }; /** * Multiply Integers @@ -38680,6 +38710,114 @@ export type components = { [key: string]: number[]; }; }; + /** + * Qwen35EncoderField + * @description Field for the Qwen3.5 text encoder Anima-3.8B's semantic connector reads. + */ + Qwen35EncoderField: { + /** @description Info to load tokenizer submodel */ + tokenizer: components["schemas"]["ModelIdentifierField"]; + /** @description Info to load text_encoder submodel */ + text_encoder: components["schemas"]["ModelIdentifierField"]; + }; + /** + * Qwen35Encoder_Checkpoint_Config + * @description Configuration for single-file Qwen3.5 text encoders (safetensors). + */ + Qwen35Encoder_Checkpoint_Config: { + /** + * Key + * @description A unique key for this model. + */ + key: string; + /** + * Hash + * @description The hash of the model file(s). + */ + hash: string; + /** + * Path + * @description Path to the model on the filesystem. Relative paths are relative to the Invoke root directory. + */ + path: string; + /** + * File Size + * @description The size of the model in bytes. + */ + file_size: number; + /** + * Name + * @description Name of the model. + */ + name: string; + /** + * Description + * @description Model description + */ + description: string | null; + /** + * Source + * @description The original source of the model (path, URL or repo_id). + */ + source: string; + /** @description The type of source */ + source_type: components["schemas"]["ModelSourceType"]; + /** + * Source Api Response + * @description The original API response from the source, as stringified JSON. + */ + source_api_response: string | null; + /** + * Source Url + * @description Optional URL for the model (e.g. download page or model page). + */ + source_url: string | null; + /** + * Cover Image + * @description Url for image to preview model + */ + cover_image: string | null; + /** + * Config Path + * @description Path to the config for this model, if any. + */ + config_path: string | null; + /** + * Base + * @default any + * @constant + */ + base: "any"; + /** + * Type + * @default qwen3_5_encoder + * @constant + */ + type: "qwen3_5_encoder"; + /** + * Format + * @default checkpoint + * @constant + */ + format: "checkpoint"; + /** + * Cpu Only + * @description Whether this model should run on CPU only + */ + cpu_only: boolean | null; + /** @description Qwen3.5 model size variant */ + variant: components["schemas"]["Qwen35VariantType"]; + }; + /** + * Qwen35VariantType + * @description Qwen3.5 text encoder variants, by language-model width. + * + * A family of its own: Qwen3.5 interleaves Gated DeltaNet linear-attention layers with gated full + * attention, so its checkpoints share neither architecture nor key layout with `Qwen3VariantType`, + * even at the same width. + * @enum {string} + */ + Qwen35VariantType: "qwen3_5_4b"; /** * Qwen3EncoderField * @description Field for Qwen3 text encoder used by Z-Image models. @@ -43297,7 +43435,7 @@ export type components = { type: components["schemas"]["ModelType"]; format?: components["schemas"]["ModelFormat"] | null; /** Variant */ - variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; + variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Qwen35VariantType"] | components["schemas"]["AnimaVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; /** * Is Installed * @default false @@ -43342,7 +43480,7 @@ export type components = { type: components["schemas"]["ModelType"]; format?: components["schemas"]["ModelFormat"] | null; /** Variant */ - variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; + variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Qwen35VariantType"] | components["schemas"]["AnimaVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; /** * Is Installed * @default false @@ -44036,7 +44174,7 @@ export type components = { path_or_prefix: string; model_type: components["schemas"]["ModelType"]; /** Variant */ - variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; + variant?: components["schemas"]["ModelVariantType"] | components["schemas"]["ClipVariantType"] | components["schemas"]["FluxVariantType"] | components["schemas"]["Flux2VariantType"] | components["schemas"]["ZImageVariantType"] | components["schemas"]["QwenImageVariantType"] | components["schemas"]["WanVariantType"] | components["schemas"]["WanLoRAVariantType"] | components["schemas"]["Qwen3VariantType"] | components["schemas"]["Qwen3VLVariantType"] | components["schemas"]["Qwen35VariantType"] | components["schemas"]["AnimaVariantType"] | components["schemas"]["Krea2VariantType"] | components["schemas"]["MiniMaxH3VariantType"] | components["schemas"]["LTX2VariantType"] | components["schemas"]["MistralVariantType"] | components["schemas"]["PiDDecoderVariantType"] | null; }; /** * Subtract Integers @@ -52298,7 +52436,7 @@ export interface operations { [name: string]: unknown; }; content: { - "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; }; }; /** @description Validation Error */ @@ -52330,7 +52468,7 @@ export interface operations { [name: string]: unknown; }; content: { - "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; }; }; /** @description Validation Error */ @@ -52382,7 +52520,7 @@ export interface operations { * "upcast_attention": false * } */ - "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; }; }; /** @description Bad request */ @@ -52496,7 +52634,7 @@ export interface operations { * "upcast_attention": false * } */ - "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; }; }; /** @description Bad request */ @@ -52569,7 +52707,7 @@ export interface operations { * "upcast_attention": false * } */ - "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; }; }; /** @description Bad request */ @@ -53335,7 +53473,7 @@ export interface operations { * "upcast_attention": false * } */ - "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; + "application/json": components["schemas"]["Main_Diffusers_SD1_Config"] | components["schemas"]["Main_Diffusers_SD2_Config"] | components["schemas"]["Main_Diffusers_SDXL_Config"] | components["schemas"]["Main_Diffusers_SDXLRefiner_Config"] | components["schemas"]["Main_Diffusers_SD3_Config"] | components["schemas"]["Main_Diffusers_FLUX_Config"] | components["schemas"]["Main_Diffusers_Flux2_Config"] | components["schemas"]["Main_Diffusers_CogView4_Config"] | components["schemas"]["Main_Diffusers_QwenImage_Config"] | components["schemas"]["Main_Diffusers_Wan_Config"] | components["schemas"]["Main_Diffusers_ZImage_Config"] | components["schemas"]["Main_Diffusers_ErnieImage_Config"] | components["schemas"]["Main_Diffusers_Ideogram4_Config"] | components["schemas"]["Main_Diffusers_Krea2_Config"] | components["schemas"]["Main_Diffusers_MiniMaxH3_Config"] | components["schemas"]["Main_Diffusers_LTX2_Config"] | components["schemas"]["Main_Checkpoint_SD1_Config"] | components["schemas"]["Main_Checkpoint_SD2_Config"] | components["schemas"]["Main_Checkpoint_SDXL_Config"] | components["schemas"]["Main_Checkpoint_SDXLRefiner_Config"] | components["schemas"]["Main_Checkpoint_Flux2_Config"] | components["schemas"]["Main_Checkpoint_FLUX_Config"] | components["schemas"]["Main_Checkpoint_QwenImage_Config"] | components["schemas"]["Main_Checkpoint_Wan_Config"] | components["schemas"]["Main_Checkpoint_ZImage_Config"] | components["schemas"]["Main_Checkpoint_ErnieImage_Config"] | components["schemas"]["Main_Checkpoint_Ideogram4_Config"] | components["schemas"]["Main_Checkpoint_Krea2_Config"] | components["schemas"]["Main_Checkpoint_Anima_Config"] | components["schemas"]["Main_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Main_Checkpoint_LTX2_Config"] | components["schemas"]["Main_BnBNF4_FLUX_Config"] | components["schemas"]["Main_GGUF_Flux2_Config"] | components["schemas"]["Main_GGUF_FLUX_Config"] | components["schemas"]["Main_GGUF_QwenImage_Config"] | components["schemas"]["Main_GGUF_Wan_Config"] | components["schemas"]["Main_GGUF_ZImage_Config"] | components["schemas"]["Main_GGUF_Krea2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_Flux2_Config"] | components["schemas"]["Main_SDNQ_Diffusers_FLUX_Config"] | components["schemas"]["Main_SDNQ_Diffusers_ZImage_Config"] | components["schemas"]["VAE_Checkpoint_SD1_Config"] | components["schemas"]["VAE_Checkpoint_SD2_Config"] | components["schemas"]["VAE_Checkpoint_SDXL_Config"] | components["schemas"]["VAE_Checkpoint_FLUX_Config"] | components["schemas"]["VAE_Checkpoint_SD3_Config"] | components["schemas"]["VAE_Checkpoint_Flux2_Config"] | components["schemas"]["VAE_Checkpoint_Wan_Config"] | components["schemas"]["VAE_Checkpoint_QwenImage_Config"] | components["schemas"]["VAE_Checkpoint_Anima_Config"] | components["schemas"]["VAE_Diffusers_SD1_Config"] | components["schemas"]["VAE_Diffusers_SDXL_Config"] | components["schemas"]["VAE_Diffusers_FLUX_Config"] | components["schemas"]["VAE_Diffusers_SD3_Config"] | components["schemas"]["VAE_Diffusers_Flux2_Config"] | components["schemas"]["VAE_Diffusers_Wan_Config"] | components["schemas"]["PiDDecoder_Checkpoint_FLUX_Config"] | components["schemas"]["PiDDecoder_Checkpoint_Flux2_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SD3_Config"] | components["schemas"]["PiDDecoder_Checkpoint_SDXL_Config"] | components["schemas"]["PiDDecoder_Checkpoint_QwenImage_Config"] | components["schemas"]["ControlNet_Checkpoint_SD1_Config"] | components["schemas"]["ControlNet_Checkpoint_SD2_Config"] | components["schemas"]["ControlNet_Checkpoint_SDXL_Config"] | components["schemas"]["ControlNet_Checkpoint_FLUX_Config"] | components["schemas"]["ControlNet_Checkpoint_ZImage_Config"] | components["schemas"]["ControlNet_Checkpoint_Anima_Config"] | components["schemas"]["ControlNet_Diffusers_SD1_Config"] | components["schemas"]["ControlNet_Diffusers_SD2_Config"] | components["schemas"]["ControlNet_Diffusers_SDXL_Config"] | components["schemas"]["ControlNet_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_SD1_Config"] | components["schemas"]["LoRA_LyCORIS_SD2_Config"] | components["schemas"]["LoRA_LyCORIS_SDXL_Config"] | components["schemas"]["LoRA_LyCORIS_Flux2_Config"] | components["schemas"]["LoRA_LyCORIS_FLUX_Config"] | components["schemas"]["LoRA_LyCORIS_ZImage_Config"] | components["schemas"]["LoRA_LyCORIS_Krea2_Config"] | components["schemas"]["LoRA_LyCORIS_QwenImage_Config"] | components["schemas"]["LoRA_LyCORIS_LTX2_Config"] | components["schemas"]["LoRA_LyCORIS_MiniMaxH3_Config"] | components["schemas"]["LoRA_LyCORIS_Wan_Config"] | components["schemas"]["LoRA_LyCORIS_Anima_Config"] | components["schemas"]["LoRA_OMI_SDXL_Config"] | components["schemas"]["LoRA_OMI_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_SD1_Config"] | components["schemas"]["LoRA_Diffusers_SD2_Config"] | components["schemas"]["LoRA_Diffusers_SDXL_Config"] | components["schemas"]["LoRA_Diffusers_Flux2_Config"] | components["schemas"]["LoRA_Diffusers_FLUX_Config"] | components["schemas"]["LoRA_Diffusers_ZImage_Config"] | components["schemas"]["ControlLoRA_LyCORIS_FLUX_Config"] | components["schemas"]["T5Encoder_T5Encoder_Config"] | components["schemas"]["T5Encoder_BnBLLMint8_Config"] | components["schemas"]["T5Encoder_SDNQ_Config"] | components["schemas"]["T5Encoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_MiniMaxH3_Config"] | components["schemas"]["Qwen3VLEncoder_Checkpoint_Config"] | components["schemas"]["Qwen3VLEncoder_GGUF_Config"] | components["schemas"]["Qwen3VLEncoder_Qwen3VLEncoder_Config"] | components["schemas"]["Qwen35Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_Qwen3Encoder_Config"] | components["schemas"]["Qwen3Encoder_Checkpoint_Config"] | components["schemas"]["Qwen3Encoder_GGUF_Config"] | components["schemas"]["Qwen3Encoder_SDNQ_Folder_Config"] | components["schemas"]["MistralEncoder_Diffusers_Config"] | components["schemas"]["MistralEncoder_Checkpoint_Config"] | components["schemas"]["MistralEncoder_GGUF_Config"] | components["schemas"]["Gemma2Encoder_Gemma2Encoder_Config"] | components["schemas"]["Gemma4Encoder_Gemma4Encoder_LTX2_Config"] | components["schemas"]["Gemma2Encoder_GGUF_Config"] | components["schemas"]["QwenVLEncoder_Diffusers_Config"] | components["schemas"]["QwenVLEncoder_Checkpoint_Config"] | components["schemas"]["WanT5Encoder_WanT5Encoder_Config"] | components["schemas"]["TI_File_SD1_Config"] | components["schemas"]["TI_File_SD2_Config"] | components["schemas"]["TI_File_SDXL_Config"] | components["schemas"]["TI_Folder_SD1_Config"] | components["schemas"]["TI_Folder_SD2_Config"] | components["schemas"]["TI_Folder_SDXL_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD1_Config"] | components["schemas"]["IPAdapter_InvokeAI_SD2_Config"] | components["schemas"]["IPAdapter_InvokeAI_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD1_Config"] | components["schemas"]["IPAdapter_Checkpoint_SD2_Config"] | components["schemas"]["IPAdapter_Checkpoint_SDXL_Config"] | components["schemas"]["IPAdapter_Checkpoint_FLUX_Config"] | components["schemas"]["T2IAdapter_Diffusers_SD1_Config"] | components["schemas"]["T2IAdapter_Diffusers_SDXL_Config"] | components["schemas"]["Spandrel_Checkpoint_Config"] | components["schemas"]["CLIPEmbed_Diffusers_G_Config"] | components["schemas"]["CLIPEmbed_Diffusers_L_Config"] | components["schemas"]["CLIPVision_Diffusers_Config"] | components["schemas"]["SigLIP_Diffusers_Config"] | components["schemas"]["FLUXRedux_Checkpoint_Config"] | components["schemas"]["LTX2DurationHead_Checkpoint_Config"] | components["schemas"]["LlavaOnevision_Diffusers_Config"] | components["schemas"]["TextLLM_Diffusers_Config"] | components["schemas"]["ExternalApiModelConfig"] | components["schemas"]["Unknown_Config"]; }; }; /** @description Bad request */ diff --git a/invokeai/frontend/webv2/src/features/generation/core/__fixtures__/architectureCapabilities.json b/invokeai/frontend/webv2/src/features/generation/core/__fixtures__/architectureCapabilities.json index b854efb1231..7af37b5ae7c 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/__fixtures__/architectureCapabilities.json +++ b/invokeai/frontend/webv2/src/features/generation/core/__fixtures__/architectureCapabilities.json @@ -55,6 +55,62 @@ ] } }, + { + "base": "anima", + "variant": "anima_qwen35", + "features": { + "negative_prompt": { + "visible": true, + "usage": "cfg-gated" + }, + "dimension_grid": 8, + "guidance_label": "CFG", + "guidance_min": 1.0, + "guidance_max": null, + "scheduler_set": "anima", + "scheduler_applies_to_graph": true, + "control_kinds": [], + "max_reference_images": 0, + "reference_images_require_variant": null, + "supports_regional_guidance": true, + "regional_negative": false, + "clip_skip_max": null, + "supports_seamless": false, + "supports_cfg_rescale": false, + "sd_vae_override": false, + "color_compensation": false, + "vae_precision": false + }, + "defaults": { + "vae": null, + "vae_precision": null, + "scheduler": "euler", + "steps": 40, + "cfg_scale": 6.0, + "cfg_rescale_multiplier": null, + "width": 1024, + "height": 1024, + "guidance": null, + "cpu_only": null, + "fp8_storage": null + }, + "vae": { + "accepted": [ + { + "base": "anima", + "latent_channels": null + }, + { + "base": "qwen-image", + "latent_channels": null + }, + { + "base": "wan", + "latent_channels": 16 + } + ] + } + }, { "base": "cogview4", "variant": null, diff --git a/invokeai/frontend/webv2/src/features/generation/core/__snapshots__/generateGraphNodeTypes.json b/invokeai/frontend/webv2/src/features/generation/core/__snapshots__/generateGraphNodeTypes.json index 7fbb0822030..88a3f8d4de9 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/__snapshots__/generateGraphNodeTypes.json +++ b/invokeai/frontend/webv2/src/features/generation/core/__snapshots__/generateGraphNodeTypes.json @@ -247,11 +247,12 @@ }, "anima": { "componentsFilled": { - "diffusers": [ + "qwen3": [ "qwen3EncoderModel", "vae" ], - "standalone-components": [ + "qwen35": [ + "qwen35EncoderModel", "qwen3EncoderModel", "vae" ] @@ -284,7 +285,8 @@ 0 ], "guidance_scale": [ - 4.5 + 4.5, + 6 ], "height": [ 1024 @@ -296,7 +298,8 @@ "euler" ], "steps": [ - 35 + 35, + 40 ], "use_cache": [ true @@ -332,6 +335,7 @@ true ], "model": [], + "qwen3_5_encoder_model": [], "qwen3_encoder_model": [], "use_cache": [ true @@ -339,6 +343,7 @@ "vae_model": [] }, "outputs": [ + "qwen3_5_encoder", "qwen3_encoder", "transformer", "vae" @@ -347,6 +352,7 @@ "anima_text_encoder": { "inputs": [ "prompt", + "qwen3_5_encoder", "qwen3_encoder" ], "literalInputs": { @@ -509,6 +515,7 @@ 3.5, 4, 4.5, + 6, 7 ], "clip_embed_model": [], @@ -583,6 +590,7 @@ "model": [], "qwen_image_qwen_vl_encoder": [], "qwen_image_vae": [], + "qwen3_5_encoder": [], "qwen3_encoder": [], "qwen3_source": [], "qwen3_vl_encoder": [], diff --git a/invokeai/frontend/webv2/src/features/generation/core/baseGenerationPolicies.ts b/invokeai/frontend/webv2/src/features/generation/core/baseGenerationPolicies.ts index af0b37053ac..3868f08f143 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/baseGenerationPolicies.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/baseGenerationPolicies.ts @@ -31,6 +31,7 @@ import type { import { getCompatibleDiffusersComponentSource, isBundledMainForBase, + isAnimaQwen35Encoder, isAnimaQwen3Encoder, isClipVariant, isDiffusersMainForBase, @@ -524,6 +525,7 @@ export const getDefaultGenerateSettings = (model?: GenerateModelConfig): Generat qwen3EncoderModel: null, qwenVLEncoderModel: null, qwen3VLEncoderModel: null, + qwen35EncoderModel: null, wanT5EncoderModel: null, ideogram4UnconditionalModel: null, wanLowNoiseModel: null, @@ -592,6 +594,7 @@ export type GenerateComponentValueKey = | 'qwen3EncoderModel' | 'qwenVLEncoderModel' | 'qwen3VLEncoderModel' + | 'qwen35EncoderModel' | 'wanT5EncoderModel' | 'wanLowNoiseModel' | 'ideogram4UnconditionalModel' @@ -636,6 +639,7 @@ const TYPE_MISTRAL: ModelTaxonomyType[] = ['mistral_encoder']; const TYPE_QWEN3: ModelTaxonomyType[] = ['qwen3_encoder']; const TYPE_QWEN_VL: ModelTaxonomyType[] = ['qwen_vl_encoder']; const TYPE_QWEN3_VL: ModelTaxonomyType[] = ['qwen3_vl_encoder']; +const TYPE_QWEN35: ModelTaxonomyType[] = ['qwen3_5_encoder']; const TYPE_WAN_T5: ModelTaxonomyType[] = ['wan_t5_encoder']; const TYPE_PID_DECODER: ModelTaxonomyType[] = ['pid_decoder']; const TYPE_GEMMA2: ModelTaxonomyType[] = ['gemma2_encoder']; @@ -699,6 +703,16 @@ const qwen3VlEncoderSlot = (helpText: string, filter: GenerateComponentFilter): filter, }); +const qwen35EncoderSlot = (helpText: string): ComponentSlotPolicy => + slot({ + key: 'qwen35EncoderModel', + label: 'Qwen3.5 Encoder', + modelTypes: TYPE_QWEN35, + valueKind: 'component', + helpText, + filter: wrapFilter(isAnimaQwen35Encoder), + }); + /** Expose PiD slots before enabling it; Z-Image uses FLUX's decoder. */ const pidDecoderSlot = (): ComponentSlotPolicy => slot({ @@ -1107,6 +1121,16 @@ const getBaseComponentSectionPolicy = ( required: () => true, missingMessage: 'Generate needs a Qwen3 Encoder for Anima models.', }, + // Anima-3.8B reads a second encoder through its bundled semantic connector; no other Anima does. + ...(model.variant === 'anima_qwen35' + ? [ + { + ...qwen35EncoderSlot('Anima-3.8B reads Qwen3.5 4B beside Qwen3 0.6B. Required for this model.'), + required: () => true, + missingMessage: 'Generate needs a Qwen3.5 Encoder for Anima-3.8B.', + }, + ] + : []), { ...vaeSlot( 'Anima decodes with the 16-channel Wan 2.1 VAE, which may be installed under the Anima, Qwen-Image, or Wan base, so all three are listed. Required for Anima models.', @@ -1153,6 +1177,7 @@ const getComponentPolicyContext = (model: GenerateModelConfig, settings: Generat qwen3EncoderModel: settings.qwen3EncoderModel, qwenVLEncoderModel: settings.qwenVLEncoderModel, qwen3VLEncoderModel: settings.qwen3VLEncoderModel, + qwen35EncoderModel: settings.qwen35EncoderModel, wanT5EncoderModel: settings.wanT5EncoderModel, wanLowNoiseModel: settings.wanLowNoiseModel, ideogram4UnconditionalModel: settings.ideogram4UnconditionalModel, @@ -1171,6 +1196,7 @@ const COMPONENT_SETTING_LABELS: Record = { qwen3EncoderModel: 'Qwen3 Encoder', qwenVLEncoderModel: 'Qwen VL Encoder', qwen3VLEncoderModel: 'Qwen3-VL Encoder', + qwen35EncoderModel: 'Qwen3.5 Encoder', wanT5EncoderModel: 'Wan T5 Encoder', wanLowNoiseModel: 'Low-noise expert', ideogram4UnconditionalModel: 'Transformer (Unconditional)', diff --git a/invokeai/frontend/webv2/src/features/generation/core/canvas/addRegionalGuidance.ts b/invokeai/frontend/webv2/src/features/generation/core/canvas/addRegionalGuidance.ts index b8891ab0afd..4939bba2532 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/canvas/addRegionalGuidance.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/canvas/addRegionalGuidance.ts @@ -157,8 +157,10 @@ const copyEncoderFields = (base: RegionalGuidanceBase, modelVariant?: string | n case 'krea-2': return ['qwen3_vl_encoder']; case 'z-image': - case 'anima': return ['qwen3_encoder']; + case 'anima': + // Copied only where the global prompt has it: Anima-3.8B's second encoder. + return ['qwen3_encoder', 'qwen3_5_encoder']; case 'sd-1': case 'sd-2': return ['clip']; diff --git a/invokeai/frontend/webv2/src/features/generation/core/componentCompatibility.ts b/invokeai/frontend/webv2/src/features/generation/core/componentCompatibility.ts index b0c17622634..85c4bf40a3b 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/componentCompatibility.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/componentCompatibility.ts @@ -74,6 +74,10 @@ export const isClipVariant = export const isAnimaQwen3Encoder: GenerateComponentFilter = (model) => model.type === 'qwen3_encoder' && model.variant === 'qwen3_06b'; +/** Anima-3.8B's semantic connector was trained on the Qwen3.5 4B. */ +export const isAnimaQwen35Encoder: GenerateComponentFilter = (model) => + model.type === 'qwen3_5_encoder' && model.variant === 'qwen3_5_4b'; + export const isNonAnimaQwen3Encoder: GenerateComponentFilter = (model) => model.type === 'qwen3_encoder' && model.variant !== 'qwen3_06b'; diff --git a/invokeai/frontend/webv2/src/features/generation/core/graph.ts b/invokeai/frontend/webv2/src/features/generation/core/graph.ts index 87a4168e40f..02be3aae118 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/graph.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/graph.ts @@ -29,6 +29,7 @@ import { } from './baseGenerationPolicies'; import { getCompatibleDiffusersComponentSource, + isAnimaQwen35Encoder, isAnimaQwen3Encoder, isBundledMainForBase, isErnieImageMistralEncoder, @@ -1125,6 +1126,11 @@ const buildAnimaGraph = ( settings.qwen3EncoderModel && isAnimaQwen3Encoder(settings.qwen3EncoderModel) ? settings.qwen3EncoderModel : null, 'Qwen3 Encoder' ); + // Anima-3.8B's bundled semantic connector also reads Qwen3.5; no other Anima model does. + const qwen35EncoderModel = + model.variant === 'anima_qwen35' + ? requireComponent(getCompatibleComponent(settings.qwen35EncoderModel, isAnimaQwen35Encoder), 'Qwen3.5 Encoder') + : null; const graph: BackendGraphContract = { edges: [], id: createId('anima_graph'), nodes: {} }; const { negativePrompt, positivePrompt, seed } = addPromptAndSeedNodes(graph); const scheduler = coerceSchedulerForGraph(model, settings.scheduler); @@ -1134,6 +1140,7 @@ const buildAnimaGraph = ( id: 'model_loader', model, qwen3_encoder_model: qwen3EncoderModel, + ...(qwen35EncoderModel ? { qwen3_5_encoder_model: qwen35EncoderModel } : {}), type: 'anima_model_loader', vae_model: vaeModel, }); @@ -1162,6 +1169,10 @@ const buildAnimaGraph = ( addEdge(graph, loraSource, 'transformer', denoise, 'transformer'); addEdge(graph, loraSource, 'qwen3_encoder', posCond, 'qwen3_encoder'); + if (qwen35EncoderModel) { + // Straight from the loader: the LoRA collection loader passes Qwen3.5 through untouched, so it has no output for it. + addEdge(graph, modelLoader, 'qwen3_5_encoder', posCond, 'qwen3_5_encoder'); + } addEdge(graph, modelLoader, 'vae', output, 'vae'); addEdge(graph, positivePrompt, 'value', posCond, 'prompt'); addEdge(graph, posCond, 'conditioning', posCondCollect, 'item'); @@ -1169,6 +1180,9 @@ const buildAnimaGraph = ( if (negCond && negCondCollect) { addEdge(graph, loraSource, 'qwen3_encoder', negCond, 'qwen3_encoder'); + if (qwen35EncoderModel) { + addEdge(graph, modelLoader, 'qwen3_5_encoder', negCond, 'qwen3_5_encoder'); + } addEdge(graph, negativePrompt, 'value', negCond, 'prompt'); addEdge(graph, negCond, 'conditioning', negCondCollect, 'item'); addEdge(graph, negCondCollect, 'collection', denoise, 'negative_conditioning'); @@ -1178,6 +1192,7 @@ const buildAnimaGraph = ( addEdge(graph, denoise, 'latents', output, 'latents'); addMetadata(graph, output, settings, model, 'anima_txt2img', projectSettings, { qwen3_encoder: qwen3EncoderModel, + ...(qwen35EncoderModel ? { qwen3_5_encoder: qwen35EncoderModel } : {}), scheduler, vae: vaeModel, }); diff --git a/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.test.ts b/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.test.ts index e506673ec1f..d825794bee5 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.test.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.test.ts @@ -108,8 +108,9 @@ describe('generate graph coverage', () => { // Require nonempty cross-base coverage after runtime seeding. expect(checked).toEqual( expect.arrayContaining([ - 'anima/standalone-components:qwen-image/undefined', - 'anima/standalone-components:wan/16', + 'anima/qwen3:qwen-image/undefined', + 'anima/qwen3:wan/16', + 'anima/qwen35:qwen-image/undefined', 'krea-2/standalone-components:anima/undefined', 'qwen-image/standalone-components:anima/undefined', 'z-image/standalone-components:flux/undefined', diff --git a/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.testing.ts b/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.testing.ts index eb2d03de4cf..6340317d653 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.testing.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/graphCoverage.testing.ts @@ -28,6 +28,11 @@ const DEFAULT_SHAPES: readonly ModelShape[] = [ /** Check per-base shape overrides against supported bases. */ export const SHAPE_OVERRIDES: Partial> = { + // Anima ships single files only; Anima-3.8B's variant adds a second encoder. + anima: [ + { label: 'qwen3', overrides: { format: 'checkpoint', variant: 'anima_qwen3' } }, + { label: 'qwen35', overrides: { format: 'checkpoint', variant: 'anima_qwen35' } }, + ], // Ideogram's standalone fixture is a conditional checkpoint, not GGUF. 'ideogram-4': [ { label: 'diffusers', overrides: { format: 'diffusers' } }, @@ -58,6 +63,7 @@ const CANDIDATE_VARIANTS = [ 'ministral3_3b', 'qwen3_vl_4b', 'qwen3_vl_8b', + 'qwen3_5_4b', ] as const; /** Ideogram 4's two transformer branches; every other main leaves the field unset. */ const CANDIDATE_BRANCHES = [undefined, 'conditional', 'unconditional'] as const; diff --git a/invokeai/frontend/webv2/src/features/generation/core/settings.ts b/invokeai/frontend/webv2/src/features/generation/core/settings.ts index 2399ad20e42..9e223a9825c 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/settings.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/settings.ts @@ -469,6 +469,7 @@ export const cloneGenerateWidgetValues = ( qwen3EncoderModel: values.qwen3EncoderModel ? { ...values.qwen3EncoderModel } : null, qwenVLEncoderModel: values.qwenVLEncoderModel ? { ...values.qwenVLEncoderModel } : null, qwen3VLEncoderModel: values.qwen3VLEncoderModel ? { ...values.qwen3VLEncoderModel } : null, + qwen35EncoderModel: values.qwen35EncoderModel ? { ...values.qwen35EncoderModel } : null, wanT5EncoderModel: values.wanT5EncoderModel ? { ...values.wanT5EncoderModel } : null, wanLowNoiseModel: values.wanLowNoiseModel ? { ...values.wanLowNoiseModel } : null, ideogram4UnconditionalModel: values.ideogram4UnconditionalModel ? { ...values.ideogram4UnconditionalModel } : null, @@ -559,6 +560,7 @@ export const syncGenerateWidgetValuesWithModels = ( mistralEncoderModel: syncModelIdentifierWithModels(values.mistralEncoderModel, modelsByKey), qwen3EncoderModel: syncModelIdentifierWithModels(values.qwen3EncoderModel, modelsByKey), qwenVLEncoderModel: syncModelIdentifierWithModels(values.qwenVLEncoderModel, modelsByKey), + qwen35EncoderModel: syncModelIdentifierWithModels(values.qwen35EncoderModel, modelsByKey), referenceImages: syncReferenceImagesWithModels(values.referenceImages, models), t5EncoderModel: syncModelIdentifierWithModels(values.t5EncoderModel, modelsByKey), vae: isVaeModelConfig(vae) ? vae : values.vae, @@ -574,6 +576,7 @@ export const syncGenerateWidgetValuesWithModels = ( nextValues.mistralEncoderModel === values.mistralEncoderModel && nextValues.qwen3EncoderModel === values.qwen3EncoderModel && nextValues.qwenVLEncoderModel === values.qwenVLEncoderModel && + nextValues.qwen35EncoderModel === values.qwen35EncoderModel && nextValues.referenceImages === values.referenceImages && nextValues.componentSourceModel === values.componentSourceModel && nextValues.modelKey === values.modelKey @@ -750,6 +753,7 @@ export const normalizeGenerateSettings = (values: unknown): GenerateSettings | n qwen3EncoderModel: getModelIdentifierOrNull(values.qwen3EncoderModel), qwenVLEncoderModel: getModelIdentifierOrNull(values.qwenVLEncoderModel), qwen3VLEncoderModel: getModelIdentifierOrNull(values.qwen3VLEncoderModel), + qwen35EncoderModel: getModelIdentifierOrNull(values.qwen35EncoderModel), wanT5EncoderModel: getModelIdentifierOrNull(values.wanT5EncoderModel), wanLowNoiseModel: getMainModelOrNull(values.wanLowNoiseModel), ideogram4UnconditionalModel: getMainModelOrNull(values.ideogram4UnconditionalModel), diff --git a/invokeai/frontend/webv2/src/features/generation/core/types.ts b/invokeai/frontend/webv2/src/features/generation/core/types.ts index 8631409294c..e2aff4fbb4a 100644 --- a/invokeai/frontend/webv2/src/features/generation/core/types.ts +++ b/invokeai/frontend/webv2/src/features/generation/core/types.ts @@ -229,6 +229,8 @@ export interface GenerateSettings { qwenVLEncoderModel: ComponentModelConfig | null; /** Krea-2's text encoder. Distinct from `qwenVLEncoderModel` (Qwen2.5-VL). */ qwen3VLEncoderModel: ComponentModelConfig | null; + /** Anima-3.8B's second text encoder, read by its bundled semantic connector. */ + qwen35EncoderModel: ComponentModelConfig | null; /** Wan 2.2's UMT5-XXL text encoder. */ wanT5EncoderModel: ComponentModelConfig | null; /** The low-noise expert is optional; the selected expert can span the full schedule. */ diff --git a/invokeai/frontend/webv2/src/features/generation/ui/GenerateComponentsSection.tsx b/invokeai/frontend/webv2/src/features/generation/ui/GenerateComponentsSection.tsx index ac05164395c..f4252f9c4e1 100644 --- a/invokeai/frontend/webv2/src/features/generation/ui/GenerateComponentsSection.tsx +++ b/invokeai/frontend/webv2/src/features/generation/ui/GenerateComponentsSection.tsx @@ -36,6 +36,7 @@ const getComponentPolicyContext = (model: GenerateModelConfig, settings: Generat qwen3EncoderModel: settings.qwen3EncoderModel, qwenVLEncoderModel: settings.qwenVLEncoderModel, qwen3VLEncoderModel: settings.qwen3VLEncoderModel, + qwen35EncoderModel: settings.qwen35EncoderModel, wanT5EncoderModel: settings.wanT5EncoderModel, wanLowNoiseModel: settings.wanLowNoiseModel, ideogram4UnconditionalModel: settings.ideogram4UnconditionalModel, diff --git a/invokeai/frontend/webv2/src/features/models/core/relationships.ts b/invokeai/frontend/webv2/src/features/models/core/relationships.ts index d5fdf769c74..e185ff768c6 100644 --- a/invokeai/frontend/webv2/src/features/models/core/relationships.ts +++ b/invokeai/frontend/webv2/src/features/models/core/relationships.ts @@ -25,6 +25,8 @@ export const NULL_BASE_ALLOWANCES: Readonly = new Set = { '5b': 'Wan 2.2 5B LoRA', a14b: 'Wan 2.2 A14B LoRA', + anima_qwen3: 'Anima', + anima_qwen35: 'Anima + Qwen3.5 (Anima-3.8B)', cow_mistral3_small: 'cow-mistral3-small (FLUX.2)', depth: 'Depth', dev: 'FLUX Dev', @@ -131,6 +134,7 @@ export const MODEL_VARIANT_LABELS: Record = { qwen3_06b: 'Qwen3 0.6B', qwen3_4b: 'Qwen3 4B', qwen3_8b: 'Qwen3 8B', + qwen3_5_4b: 'Qwen3.5 4B (Anima-3.8B)', qwen3_vl_4b: 'Qwen3-VL 4B (Krea-2)', qwen3_vl_8b: 'Qwen3-VL 8B (Ideogram 4)', ref2va: 'MiniMax H3 Ref2VA', @@ -150,6 +154,7 @@ export const getModelVariantLabel = (variant: string): string => MODEL_VARIANT_L // Mirrors the backend's per-class variant enums (taxonomy.py): main models // key their variants off the base, and a few non-main types carry their own. const MAIN_VARIANTS_BY_BASE: Record = { + anima: ['anima_qwen3', 'anima_qwen35'], flux: ['schnell', 'dev', 'dev_fill'], flux2: ['klein_4b', 'klein_4b_base', 'klein_9b', 'klein_9b_base', 'dev'], 'krea-2': ['krea2_turbo', 'krea2_base'], @@ -169,6 +174,7 @@ const VARIANTS_BY_TYPE: Record = { mistral_encoder: ['cow_mistral3_small', 'mistral3_24b', 'ministral3_3b'], pid_decoder: ['res2k_sr4x', 'res2kto4k_sr4x'], qwen3_encoder: ['qwen3_4b', 'qwen3_8b', 'qwen3_06b'], + qwen3_5_encoder: ['qwen3_5_4b'], // Required variant configs need explicit choices; a fallback None would fail database validation. qwen3_vl_encoder: ['qwen3_vl_4b', 'qwen3_vl_8b'], }; diff --git a/invokeai/frontend/webv2/src/features/models/core/types.ts b/invokeai/frontend/webv2/src/features/models/core/types.ts index deb851c51b5..fc09a8111f5 100644 --- a/invokeai/frontend/webv2/src/features/models/core/types.ts +++ b/invokeai/frontend/webv2/src/features/models/core/types.ts @@ -46,6 +46,7 @@ export type ModelTaxonomyType = | 'qwen3_encoder' | 'qwen_vl_encoder' | 'qwen3_vl_encoder' + | 'qwen3_5_encoder' | 'wan_t5_encoder' | 'mistral_encoder' | 'gemma2_encoder' diff --git a/invokeai/frontend/webv2/src/features/models/ui/detail/CpuOnlySetting.tsx b/invokeai/frontend/webv2/src/features/models/ui/detail/CpuOnlySetting.tsx index 8fe8340cef7..cba8fe20972 100644 --- a/invokeai/frontend/webv2/src/features/models/ui/detail/CpuOnlySetting.tsx +++ b/invokeai/frontend/webv2/src/features/models/ui/detail/CpuOnlySetting.tsx @@ -25,6 +25,7 @@ const CPU_ONLY_TYPES: ReadonlySet = new Set([ 'gemma2_encoder', 'mistral_encoder', 'qwen3_vl_encoder', + 'qwen3_5_encoder', ]); export const supportsCpuOnlySetting = (model: Pick): boolean => CPU_ONLY_TYPES.has(model.type); diff --git a/invokeai/frontend/webv2/src/features/workflow/core/modelRequirements.ts b/invokeai/frontend/webv2/src/features/workflow/core/modelRequirements.ts index 87f82ed5520..db97a6fa583 100644 --- a/invokeai/frontend/webv2/src/features/workflow/core/modelRequirements.ts +++ b/invokeai/frontend/webv2/src/features/workflow/core/modelRequirements.ts @@ -69,6 +69,7 @@ const MODEL_TYPE_LABELS: Record = { pid_decoder: 'PID decoder', qwen3_encoder: 'Qwen3 encoder', qwen3_vl_encoder: 'Qwen3 VL encoder', + qwen3_5_encoder: 'Qwen3.5 encoder', qwen_vl_encoder: 'Qwen VL encoder', siglip: 'SigLIP', spandrel_image_to_image: 'upscaler', diff --git a/invokeai/frontend/webv2/src/workbench/canvas-operations/importGalleryImages.test.ts b/invokeai/frontend/webv2/src/workbench/canvas-operations/importGalleryImages.test.ts index 6b44e6e8bad..891ff98dfb2 100644 --- a/invokeai/frontend/webv2/src/workbench/canvas-operations/importGalleryImages.test.ts +++ b/invokeai/frontend/webv2/src/workbench/canvas-operations/importGalleryImages.test.ts @@ -108,6 +108,7 @@ const setModel = (project: Project, base: GenerateWidgetValues['model']['base']) qwen3EncoderModel: null, qwenVLEncoderModel: null, qwen3VLEncoderModel: null, + qwen35EncoderModel: null, wanT5EncoderModel: null, wanLowNoiseModel: null, ideogram4UnconditionalModel: null, diff --git a/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.test.ts b/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.test.ts index fcd3c97ec21..b96af23f17b 100644 --- a/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.test.ts +++ b/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.test.ts @@ -110,6 +110,7 @@ const createValues = (overrides: Partial = {}): GenerateWi qwen3EncoderModel: null, qwenVLEncoderModel: null, qwen3VLEncoderModel: null, + qwen35EncoderModel: null, wanT5EncoderModel: null, wanLowNoiseModel: null, ideogram4UnconditionalModel: null, @@ -918,6 +919,7 @@ describe('remix round trip', () => { 'ideogram4UnconditionalModel', 'mistralEncoderModel', 'pidDecoderModel', + 'qwen35EncoderModel', 'qwen3EncoderModel', 'qwen3VLEncoderModel', 'qwenVLEncoderModel', diff --git a/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.ts b/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.ts index cded1adbaa0..ea297e0f8c9 100644 --- a/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.ts +++ b/invokeai/frontend/webv2/src/workbench/image-actions/imageRecall.ts @@ -410,6 +410,7 @@ type RecalledComponentSetting = keyof Pick< | 'pidDecoderModel' | 'qwen3EncoderModel' | 'qwen3VLEncoderModel' + | 'qwen35EncoderModel' | 'qwenVLEncoderModel' | 't5EncoderModel' | 'wanLowNoiseModel' @@ -439,6 +440,7 @@ const RECALLED_COMPONENTS: readonly RecalledComponent[] = [ { metadataKey: 'pid_decoder', setting: 'pidDecoderModel' }, { metadataKey: 'qwen3_encoder', setting: 'qwen3EncoderModel' }, { metadataKey: 'qwen3_vl_encoder', setting: 'qwen3VLEncoderModel' }, + { metadataKey: 'qwen3_5_encoder', setting: 'qwen35EncoderModel' }, { metadataKey: 'qwen_image_qwen_vl_encoder', setting: 'qwenVLEncoderModel' }, { metadataKey: 't5_encoder', setting: 't5EncoderModel' }, { metadataKey: 'wan_t5_encoder_model', setting: 'wanT5EncoderModel' }, diff --git a/invokeai/frontend/webv2/src/workbench/workbenchState.test.ts b/invokeai/frontend/webv2/src/workbench/workbenchState.test.ts index c6d4d403311..4486b6b23ee 100644 --- a/invokeai/frontend/webv2/src/workbench/workbenchState.test.ts +++ b/invokeai/frontend/webv2/src/workbench/workbenchState.test.ts @@ -236,6 +236,7 @@ const createGenerateValues = (overrides: Partial = {}): Ge qwen3EncoderModel: null, qwenVLEncoderModel: null, qwen3VLEncoderModel: null, + qwen35EncoderModel: null, wanT5EncoderModel: null, wanLowNoiseModel: null, ideogram4UnconditionalModel: null, diff --git a/pyproject.toml b/pyproject.toml index 6975b28d4fa..36b23f1ea67 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -329,6 +329,7 @@ version = { attr = "invokeai.version.__version__" } "invokeai.app.assets" = ["**/*.png"] "invokeai.backend.qwen3" = ["tokenizer/*.json", "tokenizer/*.json.gz"] "invokeai.backend.qwen3_vl" = ["tokenizer/*.json", "tokenizer/*.json.gz", "config/*.json"] +"invokeai.backend.qwen3_5" = ["*.json", "tokenizer/*.json", "tokenizer/*.json.gz"] "invokeai.backend.qwen2_5_vl" = ["*.json", "tokenizer/*.json", "tokenizer/*.json.gz"] # Read at load time by the MiniMax H3 single-file encoder loader. Without this the file is # absent from a wheel install and that loader raises FileNotFoundError. diff --git a/tests/app/invocations/test_anima_qwen3_5_conditioning.py b/tests/app/invocations/test_anima_qwen3_5_conditioning.py new file mode 100644 index 00000000000..34fb231a0c6 --- /dev/null +++ b/tests/app/invocations/test_anima_qwen3_5_conditioning.py @@ -0,0 +1,177 @@ +"""Anima-3.8B's Qwen3.5 conditioning through the prompt and denoise nodes. + +The semantic connector depends on the timestep, so the denoise node must recompute its context at every +step, with that step's sigma -- in float32, because the connector scales sigma by 1000 before embedding +it -- while every other Anima model keeps computing its context once. +""" + +import subprocess +import sys +from contextlib import contextmanager, nullcontext +from types import SimpleNamespace +from unittest.mock import MagicMock + +import pytest +import torch + +from invokeai.app.invocations.anima.anima_denoise import AnimaDenoiseInvocation +from invokeai.app.invocations.fields import AnimaConditioningField +from invokeai.app.invocations.text_encoder.anima_text_encoder import AnimaTextEncoderInvocation +from invokeai.backend.anima.semantic_connector import QWEN35_LAYER_INDICES +from invokeai.backend.stable_diffusion.diffusion.conditioning_data import AnimaConditioningInfo, ConditioningFieldData + + +class _Loaded: + def __init__(self, model, compute_device=torch.device("cpu")): + self.model = model + self.compute_device = compute_device + + @contextmanager + def model_on_device(self, working_mem_bytes=None): + yield (None, self.model) + + +# --- prompt node -------------------------------------------------------------------------------------- + + +class _FakeQwen35Tokenizer: + pad_token_id = 248044 + + def encode(self, prompt: str, add_special_tokens: bool) -> list[int]: + assert add_special_tokens is False + return [] if not prompt else [5, 6, 7] + + +class _FakeQwen35Encoder(torch.nn.Module): + def __init__(self) -> None: + super().__init__() + self.calls: list[tuple[list[int], tuple[int, ...], bool]] = [] + + def forward(self, input_ids, layer_indices, last_layer_attention_only=False): + self.calls.append((input_ids[0].tolist(), layer_indices, last_layer_attention_only)) + return [torch.full((1, input_ids.shape[1], 8), float(i)) for i in layer_indices] + + +def _encode_qwen3_5(monkeypatch, prompt: str) -> tuple[torch.Tensor, torch.Tensor, _FakeQwen35Encoder]: + module = "invokeai.app.invocations.text_encoder.anima_text_encoder" + monkeypatch.setattr(f"{module}.PreTrainedTokenizerBase", _FakeQwen35Tokenizer) + # The node imports Qwen35Encoder when it runs, from its own module. + monkeypatch.setattr("invokeai.backend.qwen3_5.qwen3_5_encoder.Qwen35Encoder", _FakeQwen35Encoder) + encoder = _FakeQwen35Encoder() + context = MagicMock() + context.models.load.side_effect = [_Loaded(_FakeQwen35Tokenizer()), _Loaded(encoder)] + invocation = AnimaTextEncoderInvocation.model_construct( + prompt=prompt, qwen3_5_encoder=SimpleNamespace(tokenizer=object(), text_encoder=object()) + ) + states, mask = invocation._encode_qwen3_5(context) + return states, mask, encoder + + +def test_qwen3_5_is_read_at_the_trained_layers_with_an_attention_only_last_layer(monkeypatch) -> None: + states, mask, encoder = _encode_qwen3_5(monkeypatch, "1girl, Miku") + + assert encoder.calls == [([5, 6, 7], QWEN35_LAYER_INDICES, True)] + assert states.shape == (len(QWEN35_LAYER_INDICES), 3, 8) + assert mask.tolist() == [True, True, True] + + +def test_the_prompt_node_does_not_import_transformers_qwen3_5_at_startup() -> None: + """Every app start imports every node module; transformers' qwen3_5 modeling costs ~0.25 s of it.""" + program = ( + "import sys\n" + "import invokeai.app.invocations.text_encoder.anima_text_encoder\n" + "loaded = [m for m in sys.modules if m.startswith('transformers.models.qwen3_5')]\n" + "assert not loaded, loaded\n" + "print('OK')\n" + ) + result = subprocess.run([sys.executable, "-c", program], capture_output=True, text=True, timeout=300) + assert result.returncode == 0, f"stderr:\n{result.stderr[-3000:]}" + assert "OK" in result.stdout + + +def test_an_empty_prompt_is_one_masked_padding_token(monkeypatch) -> None: + states, mask, encoder = _encode_qwen3_5(monkeypatch, "") + + assert encoder.calls[0][0] == [_FakeQwen35Tokenizer.pad_token_id] + assert states.shape[1] == 1 + assert mask.tolist() == [False] + + +# --- denoise node ------------------------------------------------------------------------------------- + + +class _FakeTransformer(torch.nn.Module): + """Records what the adapter is asked for; predicts zero velocity.""" + + def __init__(self, connector: bool) -> None: + super().__init__() + self.has_semantic_connector = connector + self.adapter_calls: list[tuple[torch.Tensor | None, bool]] = [] + self.patch_spatial = 2 + self.blocks = torch.nn.ModuleList([torch.nn.Identity() for _ in range(28)]) + + def preprocess_text_embeds(self, text_embeds, text_ids, t5xxl_weights=None, *, semantic_states=None, + semantic_mask=None, timesteps=None): # fmt: skip + self.adapter_calls.append((timesteps, semantic_states is not None)) + return torch.zeros(1, 512, 1024) + + def forward(self, x, timesteps, context, **kwargs): + return torch.zeros_like(x) + + +def _conditioning(with_qwen3_5: bool) -> ConditioningFieldData: + return ConditioningFieldData( + conditionings=[ + AnimaConditioningInfo( + qwen3_embeds=torch.zeros(3, 1024), + t5xxl_ids=torch.zeros(3, dtype=torch.long), + qwen35_states=torch.zeros(4, 2, 2560) if with_qwen3_5 else None, + qwen35_mask=torch.ones(2, dtype=torch.bool) if with_qwen3_5 else None, + ) + ] + ) + + +def _denoise(monkeypatch, *, connector: bool, with_qwen3_5: bool, steps: int = 4) -> _FakeTransformer: + module = "invokeai.app.invocations.anima.anima_denoise" + monkeypatch.setattr(f"{module}.TorchDevice.choose_torch_device", lambda: torch.device("cpu")) + monkeypatch.setattr(f"{module}.TorchDevice.choose_anima_inference_dtype", lambda _d: torch.bfloat16) + monkeypatch.setattr(f"{module}.LayerPatcher.apply_smart_model_patches", lambda **_kw: nullcontext()) + monkeypatch.setattr(f"{module}.patch_anima_for_regional_prompting", lambda *_a: nullcontext()) + transformer = _FakeTransformer(connector) + context = MagicMock() + context.models.load.return_value = _Loaded(transformer) + context.conditioning.load.side_effect = lambda _name: _conditioning(with_qwen3_5) + + invocation = AnimaDenoiseInvocation.model_construct( + latents=None, noise=None, denoise_mask=None, denoising_start=0.0, denoising_end=1.0, add_noise=True, + transformer=SimpleNamespace(transformer=object(), loras=[]), + positive_conditioning=AnimaConditioningField(conditioning_name="pos"), + negative_conditioning=AnimaConditioningField(conditioning_name="neg"), + guidance_scale=6.0, width=64, height=64, steps=steps, seed=0, control_lllite=None, scheduler="euler", + ) # fmt: skip + invocation._run_diffusion(context) + return transformer + + +def test_the_connector_context_is_recomputed_at_every_step_in_float32(monkeypatch) -> None: + transformer = _denoise(monkeypatch, connector=True, with_qwen3_5=True, steps=4) + sigmas = AnimaDenoiseInvocation.model_construct()._get_sigmas(4) + + # Positive and negative, once per step; the first step reuses what was built before the loop. + timesteps = [t for t, _ in transformer.adapter_calls] + assert len(timesteps) == 2 * 4 + assert all(t is not None and t.dtype == torch.float32 for t in timesteps) + assert [round(t.item(), 6) for t in timesteps[::2]] == [round(s, 6) for s in sigmas[:4]] + assert all(has_states for _, has_states in transformer.adapter_calls) + + +def test_a_model_without_the_connector_computes_its_context_once(monkeypatch) -> None: + transformer = _denoise(monkeypatch, connector=False, with_qwen3_5=False) + + assert transformer.adapter_calls == [(None, False), (None, False)] + + +def test_the_connector_without_qwen3_5_conditioning_is_refused(monkeypatch) -> None: + with pytest.raises(ValueError, match="encoded without Qwen3.5"): + _denoise(monkeypatch, connector=True, with_qwen3_5=False) diff --git a/tests/app/services/shared/sqlite_migrator/migrations/test_migration_2026_10_01_add_anima_variant.py b/tests/app/services/shared/sqlite_migrator/migrations/test_migration_2026_10_01_add_anima_variant.py new file mode 100644 index 00000000000..6c12251fe32 --- /dev/null +++ b/tests/app/services/shared/sqlite_migrator/migrations/test_migration_2026_10_01_add_anima_variant.py @@ -0,0 +1,107 @@ +"""The Anima variant backfill. + +Without it every Anima model installed before the variant field existed fails validation on read and is +*skipped* -- the models disappear from the model list with nothing in the UI to say why. +""" + +import json +import sqlite3 +from logging import getLogger +from pathlib import Path +from types import SimpleNamespace + +import pytest +import torch +from safetensors.torch import save_file + +from invokeai.app.services.shared.sqlite_migrator.migrations.migration_2026_10_01_add_anima_variant import ( + build_migration, +) +from invokeai.backend.model_manager.taxonomy import AnimaVariantType, BaseModelType, ModelFormat, ModelType + + +@pytest.fixture +def cursor() -> sqlite3.Cursor: + connection = sqlite3.connect(":memory:") + connection.execute("CREATE TABLE models (id TEXT PRIMARY KEY, config TEXT NOT NULL);") + return connection.cursor() + + +def _insert(cursor: sqlite3.Cursor, model_id: str, config: dict) -> None: + cursor.execute("INSERT INTO models (id, config) VALUES (?, ?);", (model_id, json.dumps(config))) + + +def _config(cursor: sqlite3.Cursor, model_id: str) -> dict: + cursor.execute("SELECT config FROM models WHERE id = ?;", (model_id,)) + return json.loads(cursor.fetchone()[0]) + + +def _run(cursor: sqlite3.Cursor, models_path: Path) -> None: + app_config = SimpleNamespace(models_path=models_path) + build_migration(app_config, getLogger(__name__)).callback(cursor) # type: ignore[arg-type,misc] + + +def _anima_main(path: str) -> dict: + return { + "type": ModelType.Main.value, + "base": BaseModelType.Anima.value, + "format": ModelFormat.Checkpoint.value, + "name": Path(path).stem, + "path": path, + } + + +def test_a_legacy_anima_model_becomes_qwen3(cursor: sqlite3.Cursor, tmp_path: Path) -> None: + checkpoint = tmp_path / "anima-base-v1.0.safetensors" + save_file({"net.blocks.0.mlp.layer1.weight": torch.zeros(1)}, checkpoint) + _insert(cursor, "base", _anima_main(str(checkpoint))) + + _run(cursor, tmp_path) + + assert _config(cursor, "base")["variant"] == AnimaVariantType.Qwen3.value + + +def test_an_installed_3_8b_bundle_becomes_qwen35(cursor: sqlite3.Cursor, tmp_path: Path) -> None: + # It identified as a plain Anima before the variant existed; its header still says what it is. + # A relative path resolves against the models directory, as the record service resolves it. + (tmp_path / "anima").mkdir() + save_file( + {"net.blocks.0.mlp.layer1.weight": torch.zeros(1), "net.anima_v2_connector.query_tokens": torch.zeros(1)}, + tmp_path / "anima" / "Anima-3.8B-v1.1.safetensors", + ) + _insert(cursor, "bundle", _anima_main("anima/Anima-3.8B-v1.1.safetensors")) + + _run(cursor, tmp_path) + + assert _config(cursor, "bundle")["variant"] == AnimaVariantType.Qwen35.value + + +def test_a_missing_file_becomes_qwen3(cursor: sqlite3.Cursor, tmp_path: Path) -> None: + _insert(cursor, "gone", _anima_main(str(tmp_path / "gone.safetensors"))) + + _run(cursor, tmp_path) + + assert _config(cursor, "gone")["variant"] == AnimaVariantType.Qwen3.value + + +def test_a_recorded_variant_is_left_alone(cursor: sqlite3.Cursor, tmp_path: Path) -> None: + _insert(cursor, "done", {**_anima_main(str(tmp_path / "x.safetensors")), "variant": AnimaVariantType.Qwen35.value}) + + _run(cursor, tmp_path) + + assert _config(cursor, "done")["variant"] == AnimaVariantType.Qwen35.value + + +def test_other_models_are_untouched(cursor: sqlite3.Cursor, tmp_path: Path) -> None: + # Anima's LLLite adapters share the base and carry no variant; other mains mean something else by it. + _insert( + cursor, + "lllite", + {"type": ModelType.ControlNet.value, "base": BaseModelType.Anima.value, "format": "checkpoint"}, + ) + _insert(cursor, "krea", {"type": ModelType.Main.value, "base": BaseModelType.Krea2.value, "format": "checkpoint"}) + + _run(cursor, tmp_path) + + assert "variant" not in _config(cursor, "lllite") + assert "variant" not in _config(cursor, "krea") diff --git a/tests/backend/anima/test_block_layout.py b/tests/backend/anima/test_block_layout.py new file mode 100644 index 00000000000..2c86d5a20ab --- /dev/null +++ b/tests/backend/anima/test_block_layout.py @@ -0,0 +1,63 @@ +"""Where the depth-expanded Anima finetunes keep the blocks of the model they grew from.""" + +import pytest + +from invokeai.backend.anima import block_layout +from invokeai.backend.anima.block_layout import adapter_block_positions + +BASE_TO_2_9B = block_layout._POSITIONS[(28, 40)] +FROM_2_9B_TO_3_8B = block_layout._POSITIONS[(40, 52)] +BASE_TO_3_8B = block_layout._POSITIONS[(28, 52)] + + +@pytest.mark.parametrize(("source", "target"), [(28, 40), (40, 52), (28, 52)]) +def test_every_source_block_has_its_own_target_block_in_order(source: int, target: int) -> None: + positions = block_layout._POSITIONS[(source, target)] + assert len(positions) == source + assert list(positions) == sorted(set(positions)) + assert positions[0] == 0 and positions[-1] < target + + +def test_the_new_blocks_sit_where_the_checkpoints_have_them() -> None: + """Measured: in 2.9B, the blocks that are not bit-identical to an Anima base block; in 3.8B, the blocks at a cosine + similarity of about 0.6 to every 2.9B block.""" + assert sorted(set(range(40)) - set(BASE_TO_2_9B)) == [2, 5, 8, 11, 14, 17, 21, 24, 27, 30, 33, 36] + assert sorted(set(range(52)) - set(FROM_2_9B_TO_3_8B)) == list(range(3, 48, 4)) + + +def test_the_base_blocks_still_bit_identical_in_3_8b_are_where_the_composed_table_puts_them() -> None: + """Ten blocks of Anima base-v1.0 are bit-identical in Anima-3.8B v1.1 -- an independent check of the composition.""" + measured = {5: 9, 7: 13, 10: 20, 14: 26, 15: 29, 19: 37, 21: 41, 22: 42, 23: 45, 27: 51} + assert {base: BASE_TO_3_8B[base] for base in measured} == measured + + +@pytest.mark.parametrize( + ("max_block_index", "target_depth", "expected"), + [ + (27, 40, BASE_TO_2_9B), + (27, 52, BASE_TO_3_8B), + (39, 52, FROM_2_9B_TO_3_8B), + # A base LoRA that leaves out the last blocks still comes from a 28-block model. + (20, 40, BASE_TO_2_9B), + ], + ids=["base-on-2.9B", "base-on-3.8B", "2.9B-on-3.8B", "partial-base-on-2.9B"], +) +def test_a_shallower_adapter_moves_to_the_blocks_it_was_trained_on(max_block_index, target_depth, expected) -> None: + assert adapter_block_positions(max_block_index, target_depth) == expected + + +@pytest.mark.parametrize( + ("max_block_index", "target_depth"), + [(27, 28), (39, 40), (51, 52), (39, 28), (51, 40), (60, 52), (27, 30)], + ids=[ + "base-on-base", + "2.9B-on-2.9B", + "3.8B-on-3.8B", + "deeper-adapter", + "3.8B-on-2.9B", + "unknown-depth", + "no-layout", + ], +) +def test_everything_else_addresses_blocks_by_index(max_block_index, target_depth) -> None: + assert adapter_block_positions(max_block_index, target_depth) is None diff --git a/tests/backend/anima/test_control_net_lllite.py b/tests/backend/anima/test_control_net_lllite.py index f09b7f1a48f..f64dd88e98b 100644 --- a/tests/backend/anima/test_control_net_lllite.py +++ b/tests/backend/anima/test_control_net_lllite.py @@ -3,6 +3,7 @@ multi-adapter composition, and the conditioning image preprocessing helpers.""" import os +import re from pathlib import Path import pytest @@ -762,3 +763,47 @@ def test_from_state_dict_real_file() -> None: assert not torch.equal(y, plain_linear(linear, x)) finally: module.unbind() + + +# ---------------------------------------------------------------------------- +# Depth-expanded finetunes +# ---------------------------------------------------------------------------- + + +def _adapter_for_blocks(*indices: int) -> AnimaControlNetLLLite: + """The synthetic 2-block adapter, its blocks renamed to ``indices``.""" + renamed = {} + for key, value in make_synthetic_state_dict().items(): + match = re.match(r"^lllite_dit_blocks_(\d+)_", key) + if match: + key = f"lllite_dit_blocks_{indices[int(match.group(1))]}_{key[match.end() :]}" + renamed[key] = value + return AnimaControlNetLLLite.from_state_dict(renamed, None) + + +def _is_wrapped(model: AnimaControlNetLLLite, linear: nn.Linear) -> bool: + x = torch.randn(1, 4, IN_DIM, generator=torch.Generator().manual_seed(5)) + model.set_multiplier(1.0) + model.set_cond_image(matching_cond_image()) + return not torch.equal(linear(x), plain_linear(linear, x)) + + +def test_a_base_adapter_binds_to_the_base_blocks_of_anima_2_9b() -> None: + """Anima base's block 27 is Anima-2.9B's block 39 (see ``invokeai.backend.anima.block_layout``).""" + model = _adapter_for_blocks(0, 27) + transformer = FakeTransformer(IN_DIM, 40) + model.apply_to(transformer) + + assert _is_wrapped(model, transformer.blocks[39].self_attn.q_proj) + assert not _is_wrapped(model, transformer.blocks[27].self_attn.q_proj) + assert _is_wrapped(model, transformer.blocks[0].self_attn.q_proj) + model.restore() + + +def test_a_base_adapter_binds_by_index_on_anima_base() -> None: + model = _adapter_for_blocks(0, 27) + transformer = FakeTransformer(IN_DIM, 28) + model.apply_to(transformer) + + assert _is_wrapped(model, transformer.blocks[27].self_attn.q_proj) + model.restore() diff --git a/tests/backend/anima/test_semantic_connector.py b/tests/backend/anima/test_semantic_connector.py new file mode 100644 index 00000000000..6c39cf5f4db --- /dev/null +++ b/tests/backend/anima/test_semantic_connector.py @@ -0,0 +1,175 @@ +"""Anima-3.8B's semantic connector: layout against the real bundle, and the behaviour the port relies on. + +The port was checked numerically against the reference implementation on the real weights (bit-identical +in fp32). These tests pin what that check cannot keep pinned: that the module tree still matches the +checkpoint's key layout, that the header is read into the right hyperparameters, and the edge cases the +denoise loop depends on -- the timestep dependence, and an empty prompt's fully masked Qwen3.5 source. +""" + +import re + +import accelerate +import pytest +import torch + +from invokeai.backend.anima.anima_transformer import AnimaTransformer, LLMAdapter, masked_sdpa +from invokeai.backend.anima.semantic_connector import AnimaSemanticConnector, AnimaSemanticConnectorConfig +from invokeai.backend.model_manager.load.model_loaders.anima import ( + ANIMA_TRANSFORMER_CONFIG, + _filter_non_model_keys, + _strip_anima_bundle_prefix, + anima_transformer_config, +) +from tests.backend.model_manager.load.state_dicts.anima_2_9b_keys import state_dict_keys as anima_2_9b_keys +from tests.backend.model_manager.load.state_dicts.anima_3_8b_connector_keys import ( + ANIMA_3_8B_NUM_BLOCKS, + connector_keys, + metadata, +) + +_CONNECTOR_PREFIX = "net.anima_v2_connector." +_BLOCK_KEY = re.compile(r"^net\.blocks\.(\d+)\.(.+)$") + + +class TestConfigFromMetadata: + def test_real_bundle_header(self) -> None: + config = AnimaSemanticConnectorConfig.from_metadata(metadata) + assert config == AnimaSemanticConnectorConfig( + num_queries=64, + resampler_blocks=6, + resampler_dim=2048, + resampler_heads=16, + mlp_hidden_dim=5632, + layer_indices=(7, 15, 23, 31), + ) + + def test_a_header_without_connector_keys_gets_the_trained_defaults(self) -> None: + assert AnimaSemanticConnectorConfig.from_metadata({}) == AnimaSemanticConnectorConfig() + + def test_a_different_connector_architecture_is_refused(self) -> None: + with pytest.raises(ValueError, match="Unsupported Anima semantic connector"): + AnimaSemanticConnectorConfig.from_metadata( + {"anima_v2_adapter_architecture": "anima_progressive_qwen35_cross_adapter_v1"} + ) + + +def test_connector_module_tree_matches_the_real_bundle() -> None: + with accelerate.init_empty_weights(): + connector = AnimaSemanticConnector(AnimaSemanticConnectorConfig.from_metadata(metadata)) + expected = {k.removeprefix(_CONNECTOR_PREFIX): shape for k, shape in connector_keys.items()} + actual = {k: list(v.shape) for k, v in connector.state_dict().items()} + assert actual == expected + + +def test_transformer_built_from_the_bundle_takes_every_key() -> None: + """52 blocks plus the connector, from the key layout the loader sees after its prefix passes.""" + block_suffixes = { + m.group(2): shape for k, shape in anima_2_9b_keys.items() if (m := _BLOCK_KEY.match(k)) and m.group(1) == "0" + } + sd = {k: torch.empty(s, device="meta") for k, s in anima_2_9b_keys.items() if not _BLOCK_KEY.match(k)} + for index in range(ANIMA_3_8B_NUM_BLOCKS): + for suffix, shape in block_suffixes.items(): + sd[f"net.blocks.{index}.{suffix}"] = torch.empty(shape, device="meta") + sd.update({k: torch.empty(s, device="meta") for k, s in connector_keys.items()}) + sd = _filter_non_model_keys(_strip_anima_bundle_prefix(sd)) + + with accelerate.init_empty_weights(): + model = AnimaTransformer( + **anima_transformer_config(sd), semantic_connector=AnimaSemanticConnectorConfig.from_metadata(metadata) + ) + result = model.load_state_dict(sd, strict=False, assign=True) + + assert len(model.blocks) == ANIMA_3_8B_NUM_BLOCKS + assert model.has_semantic_connector + assert result.unexpected_keys == [] + assert result.missing_keys == [] + + +def test_fp8_storage_casts_the_whole_connector() -> None: + """No connector module is kept out of FP8 Storage: keeping its timestep path in bf16 was measured to change + nothing (see `AnimaTransformer.__init__`), so a model with the connector declares the same patterns as one + without.""" + with accelerate.init_empty_weights(): + expanded = AnimaTransformer(**ANIMA_TRANSFORMER_CONFIG, semantic_connector=AnimaSemanticConnectorConfig()) + assert expanded.has_semantic_connector + assert expanded._skip_layerwise_casting_patterns is AnimaTransformer._skip_layerwise_casting_patterns + kept = [ + name + for name, module in expanded.named_modules() + if name.startswith("anima_v2_connector.") + and isinstance(module, torch.nn.Linear) + and any(re.search(p, name) for p in expanded._skip_layerwise_casting_patterns) + ] + assert kept == [] + + +class TestMaskedSdpa: + def test_rows_with_keys_match_plain_sdpa(self) -> None: + q, k, v = (torch.randn(1, 2, 3, 8) for _ in range(3)) + mask = torch.tensor([True, False, True, True]).reshape(1, 1, 1, 4) + k, v = torch.randn(1, 2, 4, 8), torch.randn(1, 2, 4, 8) + expected = torch.nn.functional.scaled_dot_product_attention(q, k, v, attn_mask=mask) + assert torch.allclose(masked_sdpa(q, k, v, mask), expected) + + def test_a_fully_masked_row_attends_to_nothing(self) -> None: + q, k, v = torch.randn(2, 2, 3, 8), torch.randn(2, 2, 4, 8), torch.randn(2, 2, 4, 8) + mask = torch.tensor([[True] * 4, [False] * 4]).reshape(2, 1, 1, 4) + out = masked_sdpa(q, k, v, mask) + assert torch.equal(out[1], torch.zeros_like(out[1])) + assert torch.allclose(out[0], torch.nn.functional.scaled_dot_product_attention(q[:1], k[:1], v[:1])) + + +def _tiny_connector() -> tuple[LLMAdapter, AnimaSemanticConnector]: + torch.manual_seed(0) + config = AnimaSemanticConnectorConfig( + num_queries=4, + resampler_blocks=2, + resampler_dim=32, + resampler_heads=4, + mlp_hidden_dim=48, + semantic_source_dim=24, + ) + adapter = LLMAdapter(vocab_size=50, dim=32, num_layers=2, num_heads=4) + connector = AnimaSemanticConnector(config, model_dim=32, num_heads=4, num_adapter_blocks=2) + for parameter in [*adapter.parameters(), *connector.parameters()]: + torch.nn.init.normal_(parameter, std=0.2) + return adapter.eval(), connector.eval() + + +@torch.no_grad() +def test_connector_output_changes_with_the_timestep() -> None: + adapter, connector = _tiny_connector() + source, ids = torch.randn(1, 5, 32), torch.randint(0, 50, (1, 6)) + states = [torch.randn(1, 7, 24) for _ in range(4)] + mask = torch.ones(1, 7, dtype=torch.bool) + + early = connector(adapter, source, ids, states, mask, torch.tensor([0.9])) + late = connector(adapter, source, ids, states, mask, torch.tensor([0.1])) + + assert early.shape == (1, 6, 32) + assert not torch.allclose(early, late) + + +@torch.no_grad() +def test_an_empty_prompt_contributes_no_qwen3_5_signal_through_the_anchor() -> None: + """A fully masked source: finite output, and the anchor's states make no difference.""" + adapter, connector = _tiny_connector() + source, ids = torch.randn(1, 5, 32), torch.randint(0, 50, (1, 6)) + mask = torch.zeros(1, 1, dtype=torch.bool) + t = torch.tensor([0.5]) + + one = connector(adapter, source, ids, [torch.randn(1, 1, 24) for _ in range(4)], mask, t) + other = connector(adapter, source, ids, [torch.randn(1, 1, 24) for _ in range(4)], mask, t) + + assert torch.isfinite(one).all() + # The resampler's own state still evolves (self-attention, MLP) but reads nothing from Qwen3.5 either. + assert torch.allclose(one, other) + + +def test_a_transformer_with_the_connector_refuses_to_run_without_qwen3_5() -> None: + with accelerate.init_empty_weights(): + model = AnimaTransformer( + **{**ANIMA_TRANSFORMER_CONFIG, "num_blocks": 1}, semantic_connector=AnimaSemanticConnectorConfig() + ) + with pytest.raises(ValueError, match="needs Qwen3.5 conditioning"): + model.preprocess_text_embeds(torch.empty(1, 3, 1024), torch.zeros(1, 3, dtype=torch.long)) diff --git a/tests/backend/architectures/test_variants.py b/tests/backend/architectures/test_variants.py index df728397b59..2375f97b1d5 100644 --- a/tests/backend/architectures/test_variants.py +++ b/tests/backend/architectures/test_variants.py @@ -21,6 +21,7 @@ from invokeai.backend.model_manager.configs.base import Config_Base from invokeai.backend.model_manager.configs.factory import AnyModelConfig # noqa: F401 (registers every config class) from invokeai.backend.model_manager.taxonomy import ( + AnimaVariantType, AnyVariant, BaseModelType, ClipVariantType, @@ -28,14 +29,17 @@ ModelType, Qwen3VariantType, Qwen3VLVariantType, + Qwen35VariantType, variant_type_adapter, ) -BASE_AGNOSTIC_VARIANT_ENUMS = frozenset({ClipVariantType, Qwen3VariantType, Qwen3VLVariantType, MistralVariantType}) +BASE_AGNOSTIC_VARIANT_ENUMS = frozenset( + {ClipVariantType, Qwen3VariantType, Qwen3VLVariantType, Qwen35VariantType, MistralVariantType} +) """Variant enums that cannot be declared by any architecture. -All four sit on `base=Any` configs -- CLIP embedders and the Qwen3, Qwen3-VL and Mistral text -encoders are components shared across architectures, not architectures. `Any` is a sentinel the +All of them sit on `base=Any` configs -- CLIP embedders and the Qwen3, Qwen3-VL, Qwen3.5 and Mistral +text encoders are components shared across architectures, not architectures. `Any` is a sentinel the registry refuses to register, so these are named here to keep the completeness check against `AnyVariant` total rather than quietly partial. """ @@ -223,7 +227,14 @@ def test_the_cases_that_motivate_the_model_type_dimension( def test_an_architecture_without_variants_declares_nothing() -> None: - # CogView4, ERNIE-Image, Ideogram 4 and Anima model no variants at all. An empty facet would be + # CogView4, ERNIE-Image and Ideogram 4 model no variants at all. An empty facet would be # indistinguishable from an absent one, so they omit it. - for base in (BaseModelType.CogView4, BaseModelType.ErnieImage, BaseModelType.Ideogram4, BaseModelType.Anima): + for base in (BaseModelType.CogView4, BaseModelType.ErnieImage, BaseModelType.Ideogram4): assert get_variant_enum(base, ModelType.Main) is None + + +def test_anima_variants_label_mains_only() -> None: + # The variant says which encoders a main model needs. Anima's LLLite control adapters share the + # base and carry no variant. + assert get_variant_enum(BaseModelType.Anima, ModelType.Main) is AnimaVariantType + assert get_variant_enum(BaseModelType.Anima, ModelType.ControlNet) is None diff --git a/tests/backend/model_manager/configs/test_anima_finetune_identification.py b/tests/backend/model_manager/configs/test_anima_finetune_identification.py new file mode 100644 index 00000000000..fc70879eb41 --- /dev/null +++ b/tests/backend/model_manager/configs/test_anima_finetune_identification.py @@ -0,0 +1,124 @@ +"""Identification of the depth-expanded Anima finetunes and of the Qwen3.5 encoder Anima-3.8B reads. + +Written against small safetensors files carrying the real key names (and, where identification reads +it, the real header), because identification reads both through `ModelOnDisk`. +""" + +from pathlib import Path + +import pytest +import torch +from safetensors.torch import save_file + +from invokeai.backend.model_manager.configs.factory import AnyModelConfig, ModelConfigFactory +from invokeai.backend.model_manager.configs.identification_utils import InvalidMatchError +from invokeai.backend.model_manager.configs.main import Main_Checkpoint_Anima_Config, _get_anima_variant +from invokeai.backend.model_manager.configs.qwen3_5_encoder import Qwen35Encoder_Checkpoint_Config +from invokeai.backend.model_manager.configs.qwen3_encoder import Qwen3Encoder_Checkpoint_Config +from invokeai.backend.model_manager.model_on_disk import ModelOnDisk +from invokeai.backend.model_manager.taxonomy import AnimaVariantType, Qwen35VariantType +from tests.backend.model_manager.load.state_dicts.anima_3_8b_connector_keys import metadata as anima_3_8b_metadata + +_ANIMA_KEYS = [ + "net.llm_adapter.blocks.0.cross_attn.k_norm.weight", + "net.blocks.0.adaln_modulation_cross_attn.1.weight", + "net.t_embedder.1.linear_1.weight", + "net.x_embedder.proj.1.weight", + "net.final_layer.adaln_modulation.1.weight", +] +_CONNECTOR_KEY = "net.anima_v2_connector.quality_anchor.layer_mix_logits" + + +def _write(path: Path, keys: dict[str, list[int]], metadata: dict[str, str] | None = None) -> ModelOnDisk: + save_file({k: torch.zeros(shape) for k, shape in keys.items()}, path, metadata=metadata) + return ModelOnDisk(path) + + +def _identify(mod: ModelOnDisk, override_fields: dict | None = None) -> AnyModelConfig | None: + """Identify through the real factory: every config class gets its say, so a double match shows.""" + result = ModelConfigFactory.from_model_on_disk(mod, override_fields, allow_unknown=False) + matched = [type(v).__name__ for v in result.details.values() if not isinstance(v, Exception)] + assert len(matched) <= 1, f"more than one config matched: {matched}" + return result.config + + +def _anima(tmp_path: Path, *, connector: bool, metadata: dict[str, str] | None = None) -> ModelOnDisk: + keys = {k: [1] for k in _ANIMA_KEYS} + if connector: + keys[_CONNECTOR_KEY] = [6, 4] + return _write(tmp_path / "anima.safetensors", keys, metadata) + + +class TestAnimaVariant: + @pytest.mark.parametrize("prefix", ["", "net.", "model.diffusion_model."]) + def test_a_bundled_connector_needs_qwen3_5(self, prefix: str) -> None: + sd = {f"{prefix}blocks.0.mlp.layer1.weight": None, f"{prefix}anima_v2_connector.query_tokens": None} + assert _get_anima_variant(sd) is AnimaVariantType.Qwen35 + + def test_everything_else_is_qwen3_only(self) -> None: + # Any depth: the official 28 blocks and Anima-2.9B's 40 are the same variant. + assert _get_anima_variant({"net.blocks.39.mlp.layer1.weight": None}) is AnimaVariantType.Qwen3 + + def test_official_release_and_2_9b_identify_as_qwen3(self, tmp_path: Path) -> None: + config = _identify(_anima(tmp_path, connector=False)) + assert isinstance(config, Main_Checkpoint_Anima_Config) + assert config.variant is AnimaVariantType.Qwen3 + + def test_3_8b_v1_1_identifies_as_qwen35(self, tmp_path: Path) -> None: + mod = _anima(tmp_path, connector=True, metadata=anima_3_8b_metadata) + config = _identify(mod) + assert isinstance(config, Main_Checkpoint_Anima_Config) + assert config.variant is AnimaVariantType.Qwen35 + # Default settings follow the variant (the reference workflow's CFG 6 / 40 steps). + assert config.default_settings is not None + assert (config.default_settings.steps, config.default_settings.cfg_scale) == (40, 6.0) + + def test_3_8b_v1_0_without_its_adapter_is_refused_not_installed(self, tmp_path: Path) -> None: + # The v1.0 DiT is a plain 52-block Anima on the outside; only its header says it was trained + # jointly with a Qwen3.5 adapter it does not carry. + mod = _anima(tmp_path, connector=False, metadata={"qwen35_joint_dit_blocks": "[0, 1, 2]"}) + result = ModelConfigFactory.from_model_on_disk(mod, allow_unknown=True) + # Final, not "unknown": nothing is written to the database. + assert result.config is None + refusal = result.details[Main_Checkpoint_Anima_Config.__name__] + assert isinstance(refusal, InvalidMatchError) + assert "v1.0" in str(refusal) + + def test_an_explicit_variant_override_wins(self, tmp_path: Path) -> None: + mod = _anima(tmp_path, connector=False) + config = _identify(mod, {"variant": AnimaVariantType.Qwen35}) + assert isinstance(config, Main_Checkpoint_Anima_Config) + assert config.variant is AnimaVariantType.Qwen35 + + +# The keys of Anima-3.8B's `qwen35_4b.safetensors` that identification reads, at token extents except +# where the width is the signal. +_QWEN3_5_KEYS = { + "embed_tokens.weight": [4, 2560], + "layers.0.linear_attn.in_proj_qkv.weight": [4, 4], + "layers.3.self_attn.q_proj.weight": [4, 4], + "layers.3.self_attn.q_norm.weight": [4], +} + + +class TestQwen35Encoder: + @pytest.mark.parametrize("prefix", ["", "model.", "model.language_model."]) + def test_qwen3_5_4b_is_identified(self, tmp_path: Path, prefix: str) -> None: + mod = _write(tmp_path / "q35.safetensors", {f"{prefix}{k}": s for k, s in _QWEN3_5_KEYS.items()}) + config = _identify(mod) + assert isinstance(config, Qwen35Encoder_Checkpoint_Config) + assert config.variant is Qwen35VariantType.Qwen35_4B + + def test_an_unknown_width_is_not_a_match(self, tmp_path: Path) -> None: + keys = {**_QWEN3_5_KEYS, "embed_tokens.weight": [4, 1024]} + assert _identify(_write(tmp_path / "q35.safetensors", keys)) is None + + def test_a_qwen3_encoder_is_not_a_match(self, tmp_path: Path) -> None: + keys = {"model.embed_tokens.weight": [4, 2560], "model.layers.0.self_attn.q_norm.weight": [4]} + assert isinstance(_identify(_write(tmp_path / "q3.safetensors", keys)), Qwen3Encoder_Checkpoint_Config) + + def test_the_qwen3_config_refuses_qwen3_5_at_the_same_width(self, tmp_path: Path) -> None: + # Both 4Bs are 2560 wide and a `model.`-prefixed Qwen3.5 export satisfies every Qwen3 heuristic; + # only the linear-attention layers tell them apart. + mod = _write(tmp_path / "q35.safetensors", {f"model.{k}": s for k, s in _QWEN3_5_KEYS.items()}) + assert isinstance(_identify(mod), Qwen35Encoder_Checkpoint_Config) diff --git a/tests/backend/model_manager/load/state_dicts/anima_2_9b_keys.py b/tests/backend/model_manager/load/state_dicts/anima_2_9b_keys.py new file mode 100644 index 00000000000..19d45046275 --- /dev/null +++ b/tests/backend/model_manager/load/state_dicts/anima_2_9b_keys.py @@ -0,0 +1,181 @@ +"""Representative key layout of the Anima-2.9B depth-expanded single-file checkpoint. + +Captured from `Gazingstars123/Anima-2.9B` `Anima-2.9B-preview-v1.safetensors` (bf16). The full +checkpoint has 928 tensors under the `net.` prefix: 40 DiT blocks of 20 tensors each, against 28 in +the official release, grown by interleaved insertion with every other dimension unchanged. This +fixture keeps the first and the last DiT block and every non-DiT-block key -- among them the three +`pos_embedder.*` buffers this export serializes and the official one does not. +""" + +ANIMA_2_9B_NUM_BLOCKS = 40 + +state_dict_keys: dict[str, list[int]] = { + "net.blocks.0.adaln_modulation_cross_attn.1.weight": [256, 2048], + "net.blocks.0.adaln_modulation_cross_attn.2.weight": [6144, 256], + "net.blocks.0.adaln_modulation_mlp.1.weight": [256, 2048], + "net.blocks.0.adaln_modulation_mlp.2.weight": [6144, 256], + "net.blocks.0.adaln_modulation_self_attn.1.weight": [256, 2048], + "net.blocks.0.adaln_modulation_self_attn.2.weight": [6144, 256], + "net.blocks.0.cross_attn.k_norm.weight": [128], + "net.blocks.0.cross_attn.k_proj.weight": [2048, 1024], + "net.blocks.0.cross_attn.output_proj.weight": [2048, 2048], + "net.blocks.0.cross_attn.q_norm.weight": [128], + "net.blocks.0.cross_attn.q_proj.weight": [2048, 2048], + "net.blocks.0.cross_attn.v_proj.weight": [2048, 1024], + "net.blocks.0.mlp.layer1.weight": [8192, 2048], + "net.blocks.0.mlp.layer2.weight": [2048, 8192], + "net.blocks.0.self_attn.k_norm.weight": [128], + "net.blocks.0.self_attn.k_proj.weight": [2048, 2048], + "net.blocks.0.self_attn.output_proj.weight": [2048, 2048], + "net.blocks.0.self_attn.q_norm.weight": [128], + "net.blocks.0.self_attn.q_proj.weight": [2048, 2048], + "net.blocks.0.self_attn.v_proj.weight": [2048, 2048], + "net.blocks.39.adaln_modulation_cross_attn.1.weight": [256, 2048], + "net.blocks.39.adaln_modulation_cross_attn.2.weight": [6144, 256], + "net.blocks.39.adaln_modulation_mlp.1.weight": [256, 2048], + "net.blocks.39.adaln_modulation_mlp.2.weight": [6144, 256], + "net.blocks.39.adaln_modulation_self_attn.1.weight": [256, 2048], + "net.blocks.39.adaln_modulation_self_attn.2.weight": [6144, 256], + "net.blocks.39.cross_attn.k_norm.weight": [128], + "net.blocks.39.cross_attn.k_proj.weight": [2048, 1024], + "net.blocks.39.cross_attn.output_proj.weight": [2048, 2048], + "net.blocks.39.cross_attn.q_norm.weight": [128], + "net.blocks.39.cross_attn.q_proj.weight": [2048, 2048], + "net.blocks.39.cross_attn.v_proj.weight": [2048, 1024], + "net.blocks.39.mlp.layer1.weight": [8192, 2048], + "net.blocks.39.mlp.layer2.weight": [2048, 8192], + "net.blocks.39.self_attn.k_norm.weight": [128], + "net.blocks.39.self_attn.k_proj.weight": [2048, 2048], + "net.blocks.39.self_attn.output_proj.weight": [2048, 2048], + "net.blocks.39.self_attn.q_norm.weight": [128], + "net.blocks.39.self_attn.q_proj.weight": [2048, 2048], + "net.blocks.39.self_attn.v_proj.weight": [2048, 2048], + "net.final_layer.adaln_modulation.1.weight": [256, 2048], + "net.final_layer.adaln_modulation.2.weight": [4096, 256], + "net.final_layer.linear.weight": [64, 2048], + "net.llm_adapter.blocks.0.cross_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.0.cross_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.0.cross_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.0.cross_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.0.cross_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.0.cross_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.0.mlp.0.bias": [4096], + "net.llm_adapter.blocks.0.mlp.0.weight": [4096, 1024], + "net.llm_adapter.blocks.0.mlp.2.bias": [1024], + "net.llm_adapter.blocks.0.mlp.2.weight": [1024, 4096], + "net.llm_adapter.blocks.0.norm_cross_attn.weight": [1024], + "net.llm_adapter.blocks.0.norm_mlp.weight": [1024], + "net.llm_adapter.blocks.0.norm_self_attn.weight": [1024], + "net.llm_adapter.blocks.0.self_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.0.self_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.0.self_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.0.self_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.0.self_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.0.self_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.cross_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.1.cross_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.cross_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.cross_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.1.cross_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.cross_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.mlp.0.bias": [4096], + "net.llm_adapter.blocks.1.mlp.0.weight": [4096, 1024], + "net.llm_adapter.blocks.1.mlp.2.bias": [1024], + "net.llm_adapter.blocks.1.mlp.2.weight": [1024, 4096], + "net.llm_adapter.blocks.1.norm_cross_attn.weight": [1024], + "net.llm_adapter.blocks.1.norm_mlp.weight": [1024], + "net.llm_adapter.blocks.1.norm_self_attn.weight": [1024], + "net.llm_adapter.blocks.1.self_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.1.self_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.self_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.self_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.1.self_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.1.self_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.cross_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.2.cross_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.cross_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.cross_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.2.cross_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.cross_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.mlp.0.bias": [4096], + "net.llm_adapter.blocks.2.mlp.0.weight": [4096, 1024], + "net.llm_adapter.blocks.2.mlp.2.bias": [1024], + "net.llm_adapter.blocks.2.mlp.2.weight": [1024, 4096], + "net.llm_adapter.blocks.2.norm_cross_attn.weight": [1024], + "net.llm_adapter.blocks.2.norm_mlp.weight": [1024], + "net.llm_adapter.blocks.2.norm_self_attn.weight": [1024], + "net.llm_adapter.blocks.2.self_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.2.self_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.self_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.self_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.2.self_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.2.self_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.cross_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.3.cross_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.cross_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.cross_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.3.cross_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.cross_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.mlp.0.bias": [4096], + "net.llm_adapter.blocks.3.mlp.0.weight": [4096, 1024], + "net.llm_adapter.blocks.3.mlp.2.bias": [1024], + "net.llm_adapter.blocks.3.mlp.2.weight": [1024, 4096], + "net.llm_adapter.blocks.3.norm_cross_attn.weight": [1024], + "net.llm_adapter.blocks.3.norm_mlp.weight": [1024], + "net.llm_adapter.blocks.3.norm_self_attn.weight": [1024], + "net.llm_adapter.blocks.3.self_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.3.self_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.self_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.self_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.3.self_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.3.self_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.cross_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.4.cross_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.cross_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.cross_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.4.cross_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.cross_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.mlp.0.bias": [4096], + "net.llm_adapter.blocks.4.mlp.0.weight": [4096, 1024], + "net.llm_adapter.blocks.4.mlp.2.bias": [1024], + "net.llm_adapter.blocks.4.mlp.2.weight": [1024, 4096], + "net.llm_adapter.blocks.4.norm_cross_attn.weight": [1024], + "net.llm_adapter.blocks.4.norm_mlp.weight": [1024], + "net.llm_adapter.blocks.4.norm_self_attn.weight": [1024], + "net.llm_adapter.blocks.4.self_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.4.self_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.self_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.self_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.4.self_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.4.self_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.cross_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.5.cross_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.cross_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.cross_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.5.cross_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.cross_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.mlp.0.bias": [4096], + "net.llm_adapter.blocks.5.mlp.0.weight": [4096, 1024], + "net.llm_adapter.blocks.5.mlp.2.bias": [1024], + "net.llm_adapter.blocks.5.mlp.2.weight": [1024, 4096], + "net.llm_adapter.blocks.5.norm_cross_attn.weight": [1024], + "net.llm_adapter.blocks.5.norm_mlp.weight": [1024], + "net.llm_adapter.blocks.5.norm_self_attn.weight": [1024], + "net.llm_adapter.blocks.5.self_attn.k_norm.weight": [64], + "net.llm_adapter.blocks.5.self_attn.k_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.self_attn.o_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.self_attn.q_norm.weight": [64], + "net.llm_adapter.blocks.5.self_attn.q_proj.weight": [1024, 1024], + "net.llm_adapter.blocks.5.self_attn.v_proj.weight": [1024, 1024], + "net.llm_adapter.embed.weight": [32128, 1024], + "net.llm_adapter.norm.weight": [1024], + "net.llm_adapter.out_proj.bias": [1024], + "net.llm_adapter.out_proj.weight": [1024, 1024], + "net.pos_embedder.dim_spatial_range": [21], + "net.pos_embedder.dim_temporal_range": [22], + "net.pos_embedder.seq": [256], + "net.t_embedder.1.linear_1.weight": [2048, 2048], + "net.t_embedder.1.linear_2.weight": [6144, 2048], + "net.t_embedding_norm.weight": [2048], + "net.x_embedder.proj.1.weight": [2048, 68], +} diff --git a/tests/backend/model_manager/load/state_dicts/anima_3_8b_connector_keys.py b/tests/backend/model_manager/load/state_dicts/anima_3_8b_connector_keys.py new file mode 100644 index 00000000000..65ade5f8b69 --- /dev/null +++ b/tests/backend/model_manager/load/state_dicts/anima_3_8b_connector_keys.py @@ -0,0 +1,218 @@ +"""The semantic connector Anima-3.8B v1.1 bundles, as its checkpoint records it. + +Captured from `lylogummy/Anima-3.8B` `difussion_models/Anima-3.8B-v1.1.safetensors` (bf16, 1358 tensors): +all 190 tensors under `net.anima_v2_connector.`, and the header entries the loader and identification +read. The DiT itself is 52 blocks with the per-block layout of the official release (see +`anima_2_9b_keys.py`); `qwen35_joint_dit_blocks` is the header key that marks a checkpoint trained jointly +with Qwen3.5 -- present here and in the unbundled v1.0 DiT. +""" + +ANIMA_3_8B_NUM_BLOCKS = 52 + +metadata: dict[str, str] = { + "anima_v2_adapter_architecture": "anima_qwen35_quality_anchored_semantic_connector_v2", + "anima_v2_adapter_layer_indices": "[7, 15, 23, 31]", + "anima_v2_adapter_semantic_query_tokens": "64", + "anima_v2_adapter_semantic_resampler_blocks": "6", + "anima_v2_adapter_semantic_resampler_dim": "2048", + "anima_v2_adapter_semantic_resampler_heads": "16", + "anima_v2_adapter_semantic_resampler_mlp_hidden_dim": "5632", + "architecture": "anima_3_8b_semantic_connector_v2_bundle", + "anima_v2_bundle_format": "1", + "anima_v2_connector_prefix": "net.anima_v2_connector.", + "new_block_count": "52", + "qwen35_joint_dit_blocks": "[0, 1, 2, 4, 5, 6, 7, 8, 10, 11, 12, 14, 16, 17, 18, 19, 21, 22, 24, 25, 27, 28, 30, 31, 32, 33, 34, 36, 38, 40, 43, 44, 46, 49, 50]", +} + +connector_keys: dict[str, list[int]] = { + "net.anima_v2_connector.quality_anchor.layer_mix_logits": [6, 4], + "net.anima_v2_connector.quality_anchor.query_norms.0.weight": [1024], + "net.anima_v2_connector.quality_anchor.query_norms.1.weight": [1024], + "net.anima_v2_connector.quality_anchor.query_norms.2.weight": [1024], + "net.anima_v2_connector.quality_anchor.query_norms.3.weight": [1024], + "net.anima_v2_connector.quality_anchor.query_norms.4.weight": [1024], + "net.anima_v2_connector.quality_anchor.query_norms.5.weight": [1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.0.k_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.0.k_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.0.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.0.q_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.0.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.0.v_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.1.k_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.1.k_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.1.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.1.q_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.1.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.1.v_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.2.k_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.2.k_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.2.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.2.q_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.2.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.2.v_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.3.k_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.3.k_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.3.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.3.q_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.3.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.3.v_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.4.k_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.4.k_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.4.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.4.q_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.4.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.4.v_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.5.k_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.5.k_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.semantic_attentions.5.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.5.q_norm.weight": [64], + "net.anima_v2_connector.quality_anchor.semantic_attentions.5.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.quality_anchor.semantic_attentions.5.v_proj.weight": [1024, 2560], + "net.anima_v2_connector.quality_anchor.source_norms.0.weight": [2560], + "net.anima_v2_connector.quality_anchor.source_norms.1.weight": [2560], + "net.anima_v2_connector.quality_anchor.source_norms.2.weight": [2560], + "net.anima_v2_connector.quality_anchor.source_norms.3.weight": [2560], + "net.anima_v2_connector.quality_anchor.source_norms.4.weight": [2560], + "net.anima_v2_connector.quality_anchor.source_norms.5.weight": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.0.cross_attention.k_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.0.cross_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.0.cross_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.0.cross_attention.v_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.0.mlp_in.weight": [11264, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.0.mlp_out.weight": [2048, 5632], + "net.anima_v2_connector.semantic_resampler.blocks.0.self_attention.k_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.0.self_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.0.self_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.0.self_attention.v_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.0.source_norm.bias": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.0.source_norm.weight": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.0.time_modulation.bias": [12288], + "net.anima_v2_connector.semantic_resampler.blocks.0.time_modulation.weight": [12288, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.cross_attention.k_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.1.cross_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.cross_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.cross_attention.v_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.1.mlp_in.weight": [11264, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.mlp_out.weight": [2048, 5632], + "net.anima_v2_connector.semantic_resampler.blocks.1.self_attention.k_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.self_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.self_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.self_attention.v_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.1.source_norm.bias": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.1.source_norm.weight": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.1.time_modulation.bias": [12288], + "net.anima_v2_connector.semantic_resampler.blocks.1.time_modulation.weight": [12288, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.cross_attention.k_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.2.cross_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.cross_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.cross_attention.v_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.2.mlp_in.weight": [11264, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.mlp_out.weight": [2048, 5632], + "net.anima_v2_connector.semantic_resampler.blocks.2.self_attention.k_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.self_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.self_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.self_attention.v_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.2.source_norm.bias": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.2.source_norm.weight": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.2.time_modulation.bias": [12288], + "net.anima_v2_connector.semantic_resampler.blocks.2.time_modulation.weight": [12288, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.cross_attention.k_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.3.cross_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.cross_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.cross_attention.v_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.3.mlp_in.weight": [11264, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.mlp_out.weight": [2048, 5632], + "net.anima_v2_connector.semantic_resampler.blocks.3.self_attention.k_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.self_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.self_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.self_attention.v_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.3.source_norm.bias": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.3.source_norm.weight": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.3.time_modulation.bias": [12288], + "net.anima_v2_connector.semantic_resampler.blocks.3.time_modulation.weight": [12288, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.cross_attention.k_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.4.cross_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.cross_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.cross_attention.v_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.4.mlp_in.weight": [11264, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.mlp_out.weight": [2048, 5632], + "net.anima_v2_connector.semantic_resampler.blocks.4.self_attention.k_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.self_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.self_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.self_attention.v_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.4.source_norm.bias": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.4.source_norm.weight": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.4.time_modulation.bias": [12288], + "net.anima_v2_connector.semantic_resampler.blocks.4.time_modulation.weight": [12288, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.cross_attention.k_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.5.cross_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.cross_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.cross_attention.v_proj.weight": [2048, 2560], + "net.anima_v2_connector.semantic_resampler.blocks.5.mlp_in.weight": [11264, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.mlp_out.weight": [2048, 5632], + "net.anima_v2_connector.semantic_resampler.blocks.5.self_attention.k_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.self_attention.o_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.self_attention.q_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.self_attention.v_proj.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.blocks.5.source_norm.bias": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.5.source_norm.weight": [2560], + "net.anima_v2_connector.semantic_resampler.blocks.5.time_modulation.bias": [12288], + "net.anima_v2_connector.semantic_resampler.blocks.5.time_modulation.weight": [12288, 2048], + "net.anima_v2_connector.semantic_resampler.layer_embeddings": [4, 1, 2560], + "net.anima_v2_connector.semantic_resampler.output_norm.bias": [2048], + "net.anima_v2_connector.semantic_resampler.output_norm.weight": [2048], + "net.anima_v2_connector.semantic_resampler.output_projection.weight": [1024, 2048], + "net.anima_v2_connector.semantic_resampler.query_tokens": [1, 64, 2048], + "net.anima_v2_connector.semantic_resampler.time_mlp.0.bias": [2048], + "net.anima_v2_connector.semantic_resampler.time_mlp.0.weight": [2048, 2048], + "net.anima_v2_connector.semantic_resampler.time_mlp.2.bias": [2048], + "net.anima_v2_connector.semantic_resampler.time_mlp.2.weight": [2048, 2048], + "net.anima_v2_connector.v2_attentions.0.k_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.0.k_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.0.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.0.q_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.0.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.0.v_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.1.k_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.1.k_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.1.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.1.q_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.1.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.1.v_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.2.k_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.2.k_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.2.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.2.q_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.2.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.2.v_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.3.k_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.3.k_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.3.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.3.q_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.3.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.3.v_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.4.k_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.4.k_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.4.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.4.q_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.4.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.4.v_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.5.k_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.5.k_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.5.o_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.5.q_norm.weight": [64], + "net.anima_v2_connector.v2_attentions.5.q_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_attentions.5.v_proj.weight": [1024, 1024], + "net.anima_v2_connector.v2_query_norms.0.weight": [1024], + "net.anima_v2_connector.v2_query_norms.1.weight": [1024], + "net.anima_v2_connector.v2_query_norms.2.weight": [1024], + "net.anima_v2_connector.v2_query_norms.3.weight": [1024], + "net.anima_v2_connector.v2_query_norms.4.weight": [1024], + "net.anima_v2_connector.v2_query_norms.5.weight": [1024], + "net.anima_v2_connector.v2_semantic_norms.0.weight": [1024], + "net.anima_v2_connector.v2_semantic_norms.1.weight": [1024], + "net.anima_v2_connector.v2_semantic_norms.2.weight": [1024], + "net.anima_v2_connector.v2_semantic_norms.3.weight": [1024], + "net.anima_v2_connector.v2_semantic_norms.4.weight": [1024], + "net.anima_v2_connector.v2_semantic_norms.5.weight": [1024], +} diff --git a/tests/backend/model_manager/load/test_anima_depth.py b/tests/backend/model_manager/load/test_anima_depth.py new file mode 100644 index 00000000000..d9e4052fc1f --- /dev/null +++ b/tests/backend/model_manager/load/test_anima_depth.py @@ -0,0 +1,109 @@ +"""The Anima loader builds the transformer at the depth of the checkpoint it loads. + +Anima-2.9B (40 blocks) and Anima-3.8B (52 blocks) are depth-expanded finetunes of the official +28-block release. A transformer built at the official depth loads the first 28 blocks of such a +checkpoint and drops the rest as unexpected keys under `strict=False` -- logged at DEBUG only, so the +failure is a degraded image, not an error. These tests pin the depth detection to the real key +layouts and check that a model built from it takes every key. +""" + +import re + +import accelerate +import pytest +import torch + +from invokeai.backend.anima.anima_transformer import AnimaTransformer +from invokeai.backend.model_manager.load.model_loaders.anima import ( + ANIMA_TRANSFORMER_CONFIG, + _filter_non_model_keys, + _strip_anima_bundle_prefix, + anima_transformer_config, + count_anima_dit_blocks, +) +from tests.backend.model_manager.load.state_dicts.anima_2_9b_keys import ANIMA_2_9B_NUM_BLOCKS +from tests.backend.model_manager.load.state_dicts.anima_2_9b_keys import state_dict_keys as anima_2_9b_keys +from tests.backend.model_manager.load.state_dicts.anima_comfyui_keys import state_dict_keys as anima_keys + +_BLOCK_KEY = re.compile(r"^net\.blocks\.(\d+)\.(.+)$") + + +def _full_depth_meta_state_dict(fixture: dict[str, list[int]], num_blocks: int) -> dict[str, torch.Tensor]: + """Expand a fixture's block-0 keys to `num_blocks` blocks, as meta tensors. + + Meta, because the real extents of 40 blocks are billions of elements and nothing here reads a + value. + """ + block_suffixes = { + m.group(2): shape for key, shape in fixture.items() if (m := _BLOCK_KEY.match(key)) and m.group(1) == "0" + } + sd = {key: torch.empty(shape, device="meta") for key, shape in fixture.items() if not _BLOCK_KEY.match(key)} + for index in range(num_blocks): + for suffix, shape in block_suffixes.items(): + sd[f"net.blocks.{index}.{suffix}"] = torch.empty(shape, device="meta") + return sd + + +def _prepare(sd: dict[str, torch.Tensor]) -> dict[str, torch.Tensor]: + """The two key passes the loader runs before it builds the model.""" + return _filter_non_model_keys(_strip_anima_bundle_prefix(sd)) + + +class TestCountAnimaDitBlocks: + def test_official_release_has_28_blocks(self) -> None: + sd = _prepare(_full_depth_meta_state_dict(anima_keys, 28)) + assert count_anima_dit_blocks(sd) == 28 == ANIMA_TRANSFORMER_CONFIG["num_blocks"] + + def test_anima_2_9b_has_40_blocks(self) -> None: + sd = _prepare(_full_depth_meta_state_dict(anima_2_9b_keys, ANIMA_2_9B_NUM_BLOCKS)) + assert count_anima_dit_blocks(sd) == 40 + + def test_llm_adapter_blocks_are_not_counted(self) -> None: + # `llm_adapter.blocks.` shares the `blocks.` segment; only a key that starts with it counts. + sd = {"blocks.0.mlp.layer1.weight": torch.empty(0), "llm_adapter.blocks.5.mlp.0.weight": torch.empty(0)} + assert count_anima_dit_blocks(sd) == 1 + + def test_gap_in_block_indices_is_refused(self) -> None: + sd = {f"blocks.{i}.mlp.layer1.weight": torch.empty(0) for i in (0, 1, 3)} + with pytest.raises(ValueError, match=r"gaps .*\[2\]"): + count_anima_dit_blocks(sd) + + def test_state_dict_without_blocks_is_refused(self) -> None: + with pytest.raises(ValueError, match="no DiT blocks"): + count_anima_dit_blocks({"final_layer.linear.weight": torch.empty(0)}) + + +class TestRealAnima29BLayout: + """What the captured 2.9B header says, beyond the block count.""" + + def test_last_block_has_the_same_layout_as_the_first(self) -> None: + first = {m.group(2): s for k, s in anima_2_9b_keys.items() if (m := _BLOCK_KEY.match(k)) and m.group(1) == "0"} + last = { + m.group(2): s + for k, s in anima_2_9b_keys.items() + if (m := _BLOCK_KEY.match(k)) and m.group(1) == str(ANIMA_2_9B_NUM_BLOCKS - 1) + } + assert first and first == last + + def test_every_non_block_tensor_matches_the_official_release(self) -> None: + # Only the depth changed: the LLM adapter, embedders and final layer are the official shapes. + official = {k: s for k, s in anima_keys.items() if not _BLOCK_KEY.match(k)} + expanded = {k: s for k, s in anima_2_9b_keys.items() if not _BLOCK_KEY.match(k)} + assert {k: s for k, s in expanded.items() if not k.startswith("net.pos_embedder.")} == official + + +@pytest.mark.parametrize( + ("fixture", "num_blocks"), + [(anima_keys, 28), (anima_2_9b_keys, ANIMA_2_9B_NUM_BLOCKS)], + ids=["official-28", "anima-2.9b-40"], +) +def test_model_built_at_detected_depth_takes_every_key(fixture: dict[str, list[int]], num_blocks: int) -> None: + sd = _prepare(_full_depth_meta_state_dict(fixture, num_blocks)) + + with accelerate.init_empty_weights(): + model = AnimaTransformer(**anima_transformer_config(sd)) + result = model.load_state_dict(sd, strict=False, assign=True) + + assert len(model.blocks) == num_blocks + assert result.unexpected_keys == [] + assert result.missing_keys == [] diff --git a/tests/backend/model_manager/load/test_anima_fp8_wiring.py b/tests/backend/model_manager/load/test_anima_fp8_wiring.py index 485e96dbb64..b9b66bed8ca 100644 --- a/tests/backend/model_manager/load/test_anima_fp8_wiring.py +++ b/tests/backend/model_manager/load/test_anima_fp8_wiring.py @@ -121,6 +121,10 @@ def test_single_file_loader_applies_fp8_layerwise_casting(monkeypatch, tmp_path) monkeypatch.setattr(anima_transformer_module, "AnimaTransformer", _TinyAnimaTransformer) monkeypatch.setattr(safetensors.torch, "load_file", lambda _path: {"weight": torch.ones(2, 2)}) + # The stand-in state dict has no DiT blocks to read a depth from, and the tiny model takes no config. + monkeypatch.setattr( + "invokeai.backend.model_manager.load.model_loaders.anima.anima_transformer_config", lambda _sd: {} + ) monkeypatch.setattr( "invokeai.backend.model_manager.load.model_loaders.anima.TorchDevice.choose_torch_device", lambda: torch.device("cpu"), diff --git a/tests/backend/patches/lora_conversions/test_anima_lora_conversion_utils.py b/tests/backend/patches/lora_conversions/test_anima_lora_conversion_utils.py index 3d7972cda38..3bdfd6771d6 100644 --- a/tests/backend/patches/lora_conversions/test_anima_lora_conversion_utils.py +++ b/tests/backend/patches/lora_conversions/test_anima_lora_conversion_utils.py @@ -8,9 +8,11 @@ from invokeai.backend.patches.lora_conversions.anima_lora_conversion_utils import ( _convert_kohya_te_key, _convert_kohya_unet_key, + anima_lora_for_depth, is_state_dict_likely_anima_lora, lora_model_from_anima_state_dict, ) +from invokeai.backend.patches.model_patch_raw import ModelPatchRaw from tests.backend.patches.lora_conversions.lora_state_dicts.anima_lora_kohya_format import ( state_dict_keys as anima_kohya_keys, ) @@ -228,3 +230,65 @@ def test_empty_state_dict_returns_empty_model(): """An empty state dict should produce a ModelPatchRaw with no layers.""" lora_model = lora_model_from_anima_state_dict({}) assert len(lora_model.layers) == 0 + + +# --- depth-expanded finetunes ------------------------------------------------------------------------ + + +def _patch(*keys: str) -> ModelPatchRaw: + return ModelPatchRaw(layers={key: object() for key in keys}) # type: ignore[misc] + + +T = ANIMA_LORA_TRANSFORMER_PREFIX + + +def test_a_base_lora_moves_to_the_base_blocks_of_anima_2_9b() -> None: + original = _patch( + f"{T}blocks.0.self_attn.q_proj", + f"{T}blocks.2.cross_attn.k_proj", + f"{T}blocks.27.mlp.layer2", + f"{T}llm_adapter.blocks.3.self_attn.q_proj", + f"{ANIMA_LORA_QWEN3_PREFIX}layers.2.self_attn.q_proj", + ) + moved, moved_from = anima_lora_for_depth(original, 40) + + assert moved_from == 28 + assert sorted(moved.layers) == sorted( + [ + f"{T}blocks.0.self_attn.q_proj", + f"{T}blocks.3.cross_attn.k_proj", + f"{T}blocks.39.mlp.layer2", + # The LLM adapter's own blocks and the text encoder are not DiT blocks. + f"{T}llm_adapter.blocks.3.self_attn.q_proj", + f"{ANIMA_LORA_QWEN3_PREFIX}layers.2.self_attn.q_proj", + ] + ) + # The same layer objects, and the cached LoRA itself untouched. + assert set(map(id, moved.layers.values())) == set(map(id, original.layers.values())) + assert f"{T}blocks.27.mlp.layer2" in original.layers + + +def test_a_base_lora_moves_to_the_base_blocks_of_anima_3_8b() -> None: + moved, moved_from = anima_lora_for_depth(_patch(f"{T}blocks.2.mlp.layer1", f"{T}blocks.27.mlp.layer1"), 52) + assert moved_from == 28 + assert sorted(moved.layers) == [f"{T}blocks.4.mlp.layer1", f"{T}blocks.51.mlp.layer1"] + + +def test_an_anima_2_9b_lora_moves_to_its_blocks_in_anima_3_8b() -> None: + moved, moved_from = anima_lora_for_depth(_patch(f"{T}blocks.3.mlp.layer1", f"{T}blocks.39.mlp.layer1"), 52) + assert moved_from == 40 + assert sorted(moved.layers) == [f"{T}blocks.4.mlp.layer1", f"{T}blocks.51.mlp.layer1"] + + +@pytest.mark.parametrize( + ("keys", "depth"), + [ + ((f"{T}blocks.27.mlp.layer1",), 28), + ((f"{T}blocks.39.mlp.layer1",), 40), + ((f"{T}llm_adapter.blocks.0.self_attn.q_proj", f"{ANIMA_LORA_QWEN3_PREFIX}layers.0.mlp.up_proj"), 40), + ], + ids=["base-on-base", "2.9B-on-2.9B", "no-dit-blocks"], +) +def test_a_lora_for_this_depth_is_returned_as_it_is(keys, depth) -> None: + original = _patch(*keys) + assert anima_lora_for_depth(original, depth) == (original, None) diff --git a/tests/backend/qwen3_5/test_qwen3_5_encoder.py b/tests/backend/qwen3_5/test_qwen3_5_encoder.py new file mode 100644 index 00000000000..e3013084ebf --- /dev/null +++ b/tests/backend/qwen3_5/test_qwen3_5_encoder.py @@ -0,0 +1,130 @@ +"""`Qwen35Encoder` runs the `transformers` Qwen3.5 decoder layers in its own loop. + +The loop exists so the encoder can stop at the deepest requested layer and return intermediate +outputs without the final norm. These tests pin that it computes exactly what `Qwen3_5TextModel` +computes for the same weights, and what the attention-only last layer means. (Against the reference +implementation and the real Anima-3.8B encoder it was checked separately: fp32 relative error ~1e-6.) +""" + +import pytest +import torch +from transformers.models.qwen3_5.configuration_qwen3_5 import Qwen3_5TextConfig +from transformers.models.qwen3_5.modeling_qwen3_5 import Qwen3_5TextModel + +from invokeai.backend.qwen3_5.qwen3_5_encoder import ( + Qwen35Encoder, + load_bundled_qwen3_5_tokenizer, + qwen3_5_4b_text_config, +) + + +def _tiny_config() -> Qwen3_5TextConfig: + config = Qwen3_5TextConfig( + vocab_size=64, + hidden_size=32, + intermediate_size=48, + num_hidden_layers=8, + num_attention_heads=4, + num_key_value_heads=2, + head_dim=16, + linear_num_key_heads=2, + linear_num_value_heads=4, + linear_key_head_dim=8, + linear_value_head_dim=8, + linear_conv_kernel_dim=4, + full_attention_interval=4, + rope_parameters={ + "rope_type": "default", + "rope_theta": 10000.0, + "partial_rotary_factor": 0.25, + "mrope_section": [1, 1, 0], + "mrope_interleaved": True, + }, + ) + config._attn_implementation = "sdpa" + return config + + +@pytest.fixture +def models() -> tuple[Qwen3_5TextModel, Qwen35Encoder]: + torch.manual_seed(0) + config = _tiny_config() + reference = Qwen3_5TextModel(config).eval() + for parameter in reference.parameters(): + torch.nn.init.normal_(parameter, std=0.1) + encoder = Qwen35Encoder(config).eval() + encoder.load_state_dict({k: v for k, v in reference.state_dict().items() if not k.startswith("norm.")}, strict=True) + return reference, encoder + + +def _layer_outputs(reference: Qwen3_5TextModel, input_ids: torch.Tensor) -> list[torch.Tensor]: + """Each decoder layer's output, captured by hook -- not through `output_hidden_states`, whose last entry is normed.""" + captured: list[torch.Tensor] = [] + hooks = [layer.register_forward_hook(lambda _m, _i, out: captured.append(out)) for layer in reference.layers] + try: + with torch.no_grad(): + reference(input_ids=input_ids) + finally: + for hook in hooks: + hook.remove() + return captured + + +def test_layer_outputs_match_the_transformers_model(models) -> None: + reference, encoder = models + input_ids = torch.randint(0, 64, (1, 70)) # longer than one 64-token delta-rule chunk + expected = _layer_outputs(reference, input_ids) + + outputs = encoder(input_ids, (1, 3, 7)) + + for index, output in zip((1, 3, 7), outputs, strict=True): + torch.testing.assert_close(output, expected[index]) + + +def test_attention_only_last_layer_skips_exactly_its_mlp(models) -> None: + reference, encoder = models + input_ids = torch.randint(0, 64, (1, 9)) + expected = _layer_outputs(reference, input_ids) + last = encoder.layers[7] + with torch.no_grad(): + residual = expected[6] + mlp_term = last.mlp(last.post_attention_layernorm(_post_attention(last, residual))) + + (output,) = encoder(input_ids, (7,), last_layer_attention_only=True) + + torch.testing.assert_close(output + mlp_term, expected[7]) + + +def _post_attention(layer, hidden: torch.Tensor) -> torch.Tensor: + rotary = Qwen35Encoder(_tiny_config()).rotary_emb + position_ids = torch.arange(hidden.shape[1]).unsqueeze(0).unsqueeze(0).expand(3, -1, -1) + return hidden + Qwen35Encoder._token_mixer(layer, hidden, rotary(hidden, position_ids)) + + +def test_layers_past_the_deepest_requested_one_are_not_run(models) -> None: + _, encoder = models + calls: list[int] = [] + for index, layer in enumerate(encoder.layers): + layer.register_forward_hook(lambda *_a, i=index: calls.append(i)) + + encoder(torch.randint(0, 64, (1, 5)), (2,)) + + assert calls == [0, 1, 2] + + +def test_the_vendored_4b_config_is_the_published_one() -> None: + config = qwen3_5_4b_text_config() + assert (config.hidden_size, config.num_hidden_layers, config.vocab_size) == (2560, 32, 248320) + assert [i for i, t in enumerate(config.layer_types) if t == "full_attention"] == [3, 7, 11, 15, 19, 23, 27, 31] + assert config.rope_parameters["partial_rotary_factor"] == 0.25 + assert config._attn_implementation == "sdpa" + + +def test_the_vendored_tokenizer_encodes_offline() -> None: + tokenizer = load_bundled_qwen3_5_tokenizer() + # Token ids pinned from the tokenizer the reference extension ships (Qwen/Qwen3.5-4B @ 851bf6e). + assert tokenizer.encode("1girl, Miku from Vocaloid", add_special_tokens=False) == [ + 16, 27620, 11, 380, 36974, 494, 93948, 573, + ] # fmt: skip + assert tokenizer.encode("", add_special_tokens=False) == [] + assert tokenizer.pad_token == "<|endoftext|>" diff --git a/tests/test_package_data.py b/tests/test_package_data.py new file mode 100644 index 00000000000..5d4d287b607 --- /dev/null +++ b/tests/test_package_data.py @@ -0,0 +1,45 @@ +"""Every data file a backend module reads at runtime ships in the wheel. + +Vendored tokenizers and configs are read from the installed package, so a file that +`[tool.setuptools.package-data]` does not match is absent from a wheel install: the loader raises +`FileNotFoundError`, or -- for a tokenizer missing only its vocabulary -- encodes every prompt to +nothing. A source checkout has every file, so nothing but this test notices. +""" + +import tomllib +from fnmatch import fnmatch +from pathlib import Path + +REPO_ROOT = Path(__file__).resolve().parents[1] +BACKEND = REPO_ROOT / "invokeai" / "backend" +DATA_SUFFIXES = (".json", ".json.gz") + + +def _package_data() -> dict[str, list[str]]: + with open(REPO_ROOT / "pyproject.toml", "rb") as f: + return tomllib.load(f)["tool"]["setuptools"]["package-data"] + + +def _is_shipped(path: Path, package_data: dict[str, list[str]]) -> bool: + """True if a package-data glob of `path`'s package, or of any package above it, matches it.""" + relative = path.relative_to(REPO_ROOT) + parts = relative.parts + for depth in range(len(parts) - 1, 0, -1): + package = ".".join(parts[:depth]) + member = "/".join(parts[depth:]) + for pattern in package_data.get(package, []): + # setuptools globs: `**` spans directories, `*` stays within one. + if fnmatch(member, pattern) and (pattern.count("/") == member.count("/") or "**" in pattern): + return True + return False + + +def test_every_vendored_backend_data_file_is_package_data() -> None: + package_data = _package_data() + vendored = [ + p for p in BACKEND.rglob("*") if p.is_file() and p.name.endswith(DATA_SUFFIXES) and "__pycache__" not in p.parts + ] + assert vendored, "found no vendored data files -- the scan is broken" + + missing = sorted(str(p.relative_to(REPO_ROOT)) for p in vendored if not _is_shipped(p, package_data)) + assert not missing, "not matched by [tool.setuptools.package-data] in pyproject.toml:\n " + "\n ".join(missing)