Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
41 changes: 37 additions & 4 deletions docs/src/content/docs/Users Guide/03.Models/Local Models/anima.mdx
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
title: Anima
description: Generate anime-style images with Anima, a 2B Cosmos Predict2 diffusion transformer with an LLM adapter, including its schedulers, regional prompting and ControlNet-LLLite adapters.
lastUpdated: 2026-09-30
lastUpdated: 2026-10-02
sidebar:
order: 1
---
Expand All @@ -12,7 +12,8 @@ built-in **LLM adapter** that translates the encoder's output for the transforme
16-channel Wan 2.1 / Qwen Image VAE.

InvokeAI supports Anima for text-to-image, and on the Canvas for image-to-image, inpainting, outpainting
and regional prompts.
and regional prompts. Two community finetunes are supported as well: **Anima-2.9B** and **Anima-3.8B**
(see [below](#community-finetunes)).

## License

Expand Down Expand Up @@ -53,6 +54,38 @@ three are listed. FLUX VAEs are **not** compatible.

Anima also needs a T5-XXL tokenizer, which ships with InvokeAI; no T5 model has to be installed.

## Community finetunes

Both finetunes keep Anima's VAE, schedulers and Canvas support, and install from the Starter Models.

| Model | What changes | Components |
| ----- | ------------ | ---------- |
| **Anima-2.9B** (Preview v1) | 40 transformer blocks instead of 28, trained on 1.7M more images (cutoff July 2026). An int8 build (~3.1 GB) halves the memory of the bf16 one (~5.8 GB); it is not faster. | Same as Anima |
| **Anima-3.8B** (v1.1) | 52 blocks, plus a **semantic connector** bundled in the checkpoint that adds a second text encoder, **Qwen3.5 4B**, for better prompt adherence, multi-character binding and mixed natural-language/tag prompts. | Qwen3 0.6B **and** Qwen3.5 4B encoder, VAE |

With an Anima-3.8B model selected, the **Components** section shows a third picker, **Qwen3.5 Encoder**,
and generation needs it. Its starter model (~4.8 GB) installs together with Anima-3.8B. The connector
runs once per denoising step, so Anima-3.8B is slower per step than Anima-2.9B (about 1.3 against 1.7
iterations per second at 832×1216 on an RTX 4090), and its Qwen3.5 encoder adds memory while the prompt
is encoded. **Reset all to model defaults** sets 40 steps at CFG 6, the settings of the author's reference
workflow.

[FP8 Storage](/configuration/optimization/fp8-storage/) changes Anima-3.8B's images more than Anima's: at
the same seed the composition can change, where Anima keeps it and only details differ.

Only the v1.1 checkpoint of Anima-3.8B is supported. The earlier v1.0 transformer needs a separate adapter
file and is refused at install with a message that says so.

Both finetunes are derivatives of Anima and fall under the same non-commercial license.

:::note[LoRAs and LLLite adapters]
LoRAs and ControlNet-LLLite adapters address the transformer's blocks by position, and the finetunes
insert their extra blocks between the original ones. Invoke moves an adapter trained on Anima to the
blocks the finetune kept from Anima, and one trained on Anima-2.9B to its blocks in Anima-3.8B.
Anima-2.9B keeps Anima's blocks unchanged, so Anima adapters work there as on Anima itself.
Anima-3.8B changed them slightly and adds the semantic connector, so results there can differ more.
:::

## Generation settings

Selecting an Anima model applies these defaults:
Expand Down Expand Up @@ -110,8 +143,8 @@ Base 1.0; the pose adapter is notably weak. Only the Inpainting and Sketch adapt

| Node | Purpose |
| ---------------------------- | --------------------------------------------------------- |
| **Main Model - Anima** | Loads the transformer, Qwen3 encoder and VAE |
| **Prompt - Anima** | Encodes a prompt, with an optional regional mask |
| **Main Model - Anima** | Loads the transformer, Qwen3 encoder and VAE, and for Anima-3.8B the Qwen3.5 encoder |
| **Prompt - Anima** | Encodes a prompt, with an optional regional mask; connect the Qwen3.5 encoder for Anima-3.8B |
| **Denoise - Anima** | Runs sampling; accepts img2img latents, masks and LLLite adapters |
| **Image to Latents - Anima** | VAE encode |
| **Latents to Image - Anima** | VAE decode |
Expand Down
119 changes: 108 additions & 11 deletions docs/src/content/docs/contributing/new-model-integration.mdx

Large diffs are not rendered by default.

209 changes: 139 additions & 70 deletions invokeai/app/invocations/anima/anima_denoise.py

Large diffs are not rendered by default.

41 changes: 39 additions & 2 deletions invokeai/app/invocations/anima/anima_model_loader.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,5 @@
from typing import Optional

from invokeai.app.invocations.baseinvocation import (
BaseInvocation,
BaseInvocationOutput,
Expand All @@ -9,12 +11,13 @@
from invokeai.app.invocations.model import (
ModelIdentifierField,
Qwen3EncoderField,
Qwen35EncoderField,
TransformerField,
VAEField,
)
from invokeai.app.services.shared.invocation_context import InvocationContext
from invokeai.backend.architectures import accepted_vae_bases
from invokeai.backend.model_manager.taxonomy import BaseModelType, ModelType, SubModelType
from invokeai.backend.model_manager.taxonomy import AnimaVariantType, BaseModelType, ModelType, SubModelType


@invocation_output("anima_model_loader_output")
Expand All @@ -23,6 +26,11 @@ class AnimaModelLoaderOutput(BaseInvocationOutput):

transformer: TransformerField = OutputField(description=FieldDescriptions.transformer, title="Transformer")
qwen3_encoder: Qwen3EncoderField = OutputField(description=FieldDescriptions.qwen3_encoder, title="Qwen3 Encoder")
qwen3_5_encoder: Optional[Qwen35EncoderField] = OutputField(
default=None,
description=f"{FieldDescriptions.qwen3_5_encoder}. Set only for an Anima-3.8B model.",
title="Qwen3.5 Encoder",
)
vae: VAEField = OutputField(description=FieldDescriptions.vae, title="VAE")


Expand All @@ -31,7 +39,7 @@ class AnimaModelLoaderOutput(BaseInvocationOutput):
title="Main Model - Anima",
tags=["model", "anima"],
category="model",
version="1.4.0",
version="1.5.0",
classification=Classification.Prototype,
)
class AnimaModelLoaderInvocation(BaseInvocation):
Expand All @@ -41,6 +49,7 @@ class AnimaModelLoaderInvocation(BaseInvocation):
- Transformer: Cosmos Predict2 DiT + LLM Adapter (from single-file checkpoint)
- Qwen3 Encoder: Qwen3 0.6B (standalone single-file)
- VAE: AutoencoderKLQwenImage / Wan 2.1 VAE (standalone single-file)
- Qwen3.5 Encoder: Qwen3.5 4B, for Anima-3.8B only, whose bundled semantic connector reads it

The T5-XXL tokenizer needed for LLM Adapter token IDs is bundled in the package,
so no T5-XXL encoder model needs to be installed.
Expand Down Expand Up @@ -69,6 +78,14 @@ class AnimaModelLoaderInvocation(BaseInvocation):
title="Qwen3 Encoder",
)

qwen3_5_encoder_model: Optional[ModelIdentifierField] = InputField(
default=None,
description="Standalone Qwen3.5 4B Encoder model. Required by Anima-3.8B, ignored by every other Anima model.",
input=Input.Direct,
ui_model_type=ModelType.Qwen35Encoder,
title="Qwen3.5 Encoder",
)

def invoke(self, context: InvocationContext) -> AnimaModelLoaderOutput:
# Transformer always comes from the main model
transformer = self.model.model_copy(update={"submodel_type": SubModelType.Transformer})
Expand All @@ -83,5 +100,25 @@ def invoke(self, context: InvocationContext) -> AnimaModelLoaderOutput:
return AnimaModelLoaderOutput(
transformer=TransformerField(transformer=transformer, loras=[]),
qwen3_encoder=Qwen3EncoderField(tokenizer=qwen3_tokenizer, text_encoder=qwen3_encoder),
qwen3_5_encoder=self._qwen3_5_encoder(context),
vae=VAEField(vae=vae),
)

def _qwen3_5_encoder(self, context: InvocationContext) -> Optional[Qwen35EncoderField]:
"""The Qwen3.5 encoder if the main model reads one, else None. Refuses a model that needs one and has none."""
variant = getattr(context.models.get_config(self.model), "variant", None)
if variant != AnimaVariantType.Qwen35:
if self.qwen3_5_encoder_model is not None:
context.logger.warning(
f"{self.model.name} is not conditioned on Qwen3.5; ignoring the selected Qwen3.5 encoder."
)
return None
if self.qwen3_5_encoder_model is None:
raise ValueError(
f"{self.model.name} bundles a Qwen3.5 semantic connector and needs a Qwen3.5 4B encoder. "
"Install one (qwen35_4b.safetensors) and select it as the Qwen3.5 Encoder."
)
return Qwen35EncoderField(
tokenizer=self.qwen3_5_encoder_model.model_copy(update={"submodel_type": SubModelType.Tokenizer}),
text_encoder=self.qwen3_5_encoder_model.model_copy(update={"submodel_type": SubModelType.TextEncoder}),
)
1 change: 1 addition & 0 deletions invokeai/app/invocations/fields.py
Original file line number Diff line number Diff line change
Expand Up @@ -163,6 +163,7 @@ class FieldDescriptions:
glm_encoder = "GLM (THUDM) tokenizer and text encoder"
qwen3_encoder = "Qwen3 tokenizer and text encoder"
qwen3_vl_encoder = "Qwen3-VL tokenizer and text encoder"
qwen3_5_encoder = "Qwen3.5 tokenizer and text encoder"
mistral_encoder = "Mistral tokenizer/processor and text encoder"
clip_embed_model = "CLIP Embed loader"
clip_g_model = "CLIP-G Embed loader"
Expand Down
7 changes: 7 additions & 0 deletions invokeai/app/invocations/model.py
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,13 @@ class Qwen3VLEncoderField(BaseModel):
loras: List[LoRAField] = Field(default_factory=list, description="LoRAs to apply on model loading")


class Qwen35EncoderField(BaseModel):
"""Field for the Qwen3.5 text encoder Anima-3.8B's semantic connector reads."""

tokenizer: ModelIdentifierField = Field(description="Info to load tokenizer submodel")
text_encoder: ModelIdentifierField = Field(description="Info to load text_encoder submodel")


class WanT5EncoderField(BaseModel):
"""Field for the UMT5-XXL text encoder used by Wan 2.2 models."""

Expand Down
67 changes: 65 additions & 2 deletions invokeai/app/invocations/text_encoder/anima_text_encoder.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,10 @@
Both outputs are stored together in AnimaConditioningInfo and used by
the LLM Adapter inside the transformer during denoising.

Anima-3.8B additionally reads Qwen3.5 4B hidden states (layers 7, 15, 23, 31) through the semantic
connector bundled in its transformer. They are encoded here when a Qwen3.5 encoder is connected and
stored beside the Qwen3 ones; the connector itself runs inside the denoising loop, per step.

Key differences from Z-Image text encoder:
- Anima uses Qwen3 0.6B (base model, NOT instruct) — no chat template
- Anima additionally tokenizes with T5-XXL tokenizer to get token IDs
Expand All @@ -28,9 +32,10 @@
TensorField,
UIComponent,
)
from invokeai.app.invocations.model import Qwen3EncoderField
from invokeai.app.invocations.model import Qwen3EncoderField, Qwen35EncoderField
from invokeai.app.invocations.primitives import AnimaConditioningOutput
from invokeai.app.services.shared.invocation_context import InvocationContext
from invokeai.backend.anima.semantic_connector import QWEN35_LAYER_INDICES
from invokeai.backend.patches.layer_patcher import LayerPatcher, PatchSpec
from invokeai.backend.patches.lora_conversions.anima_lora_constants import ANIMA_LORA_QWEN3_PREFIX
from invokeai.backend.patches.model_patch_raw import ModelPatchRaw
Expand All @@ -52,13 +57,16 @@
# Qwen3 0.6B supports 32K context but the LLM Adapter doesn't need that much.
QWEN3_MAX_SEQ_LEN = 8192

# The reference encoder for Anima-3.8B's semantic connector truncates the Qwen3.5 prompt here.
QWEN3_5_MAX_SEQ_LEN = 1024


@invocation(
"anima_text_encoder",
title="Prompt - Anima",
tags=["prompt", "conditioning", "anima"],
category="conditioning",
version="1.4.0",
version="1.5.0",
classification=Classification.Prototype,
idle_gpu_offloadable=True,
)
Expand All @@ -80,10 +88,17 @@ class AnimaTextEncoderInvocation(BaseInvocation):
default=None,
description="A mask defining the region that this conditioning prompt applies to.",
)
qwen3_5_encoder: Qwen35EncoderField | None = InputField(
default=None,
title="Qwen3.5 Encoder",
description=f"{FieldDescriptions.qwen3_5_encoder}. Connect it for Anima-3.8B, which reads both encoders.",
input=Input.Connection,
)

@torch.no_grad()
def invoke(self, context: InvocationContext) -> AnimaConditioningOutput:
qwen3_embeds, t5xxl_ids, t5xxl_weights = self._encode_prompt(context)
qwen35_states, qwen35_mask = self._encode_qwen3_5(context) if self.qwen3_5_encoder else (None, None)

# Move to CPU for storage
qwen3_embeds = qwen3_embeds.detach().to("cpu")
Expand All @@ -96,6 +111,8 @@ def invoke(self, context: InvocationContext) -> AnimaConditioningOutput:
qwen3_embeds=qwen3_embeds,
t5xxl_ids=t5xxl_ids,
t5xxl_weights=t5xxl_weights,
qwen35_states=qwen35_states,
qwen35_mask=qwen35_mask,
)
]
)
Expand Down Expand Up @@ -219,6 +236,52 @@ def _encode_prompt(

return qwen3_embeds, t5xxl_ids, None

def _encode_qwen3_5(self, context: InvocationContext) -> tuple[torch.Tensor, torch.Tensor]:
"""Encode the prompt with Qwen3.5 the way Anima-3.8B's reference encoder does.

Returns:
Tuple of (states, mask), on the CPU.
- states: Shape (num_layers, seq_len, 2560), one row per layer in `QWEN35_LAYER_INDICES`.
- mask: Shape (seq_len,), True for real tokens.
"""
assert self.qwen3_5_encoder is not None
context.util.signal_progress("Running Qwen3.5 text encoder")
tokenizer_info = context.models.load(self.qwen3_5_encoder.tokenizer)
with tokenizer_info.model_on_device() as (_, tokenizer):
if not isinstance(tokenizer, PreTrainedTokenizerBase):
raise TypeError(f"Expected PreTrainedTokenizerBase for tokenizer, got {type(tokenizer).__name__}.")
# No template, no special tokens -- the prompt's tokens as they are.
token_ids = tokenizer.encode(self.prompt, add_special_tokens=False)
pad_token_id = tokenizer.pad_token_id
if len(token_ids) > QWEN3_5_MAX_SEQ_LEN:
logger.warning(f"Prompt was truncated to {QWEN3_5_MAX_SEQ_LEN} tokens for the Qwen3.5 encoder.")
token_ids = token_ids[:QWEN3_5_MAX_SEQ_LEN]
# An empty prompt becomes one padding token that nothing attends to, as in the reference, so the
# connector adds no Qwen3.5 signal for it. (The reference also masks every token from the first
# occurrence of id 151643 on -- the Qwen3 padding id, which in Qwen3.5's vocabulary is the
# ordinary token " 내용". That is not reproduced.)
is_empty = not token_ids
if is_empty:
if pad_token_id is None:
raise ValueError("The Qwen3.5 tokenizer has no padding token to encode an empty prompt with.")
token_ids = [pad_token_id]
mask = torch.full((len(token_ids),), not is_empty, dtype=torch.bool)

# Imported here: transformers' qwen3_5 modeling module costs ~0.25 s, which every app start would
# otherwise pay for a node that only Anima-3.8B uses.
from invokeai.backend.qwen3_5.qwen3_5_encoder import Qwen35Encoder

encoder_info = context.models.load(self.qwen3_5_encoder.text_encoder)
with encoder_info.model_on_device() as (_, encoder):
if not isinstance(encoder, Qwen35Encoder):
raise TypeError(f"Expected a Qwen3.5 encoder, got {type(encoder).__name__}.")
input_ids = torch.tensor([token_ids], device=encoder_info.compute_device)
# The connector was trained on what Anima-3.8B's encoder checkpoint computes, and that file
# ships its last tapped layer without the MLP.
states = encoder(input_ids, QWEN35_LAYER_INDICES, last_layer_attention_only=True)
stacked = torch.cat(states, dim=0).detach().to("cpu")
return stacked, mask

def _lora_iterator(self, context: InvocationContext) -> Iterator[PatchSpec]:
"""Iterate over LoRA models to apply to the Qwen3 text encoder."""
for lora in self.qwen3_encoder.loras:
Expand Down
4 changes: 4 additions & 0 deletions invokeai/app/services/model_records/model_records_base.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
from invokeai.backend.model_manager.configs.lora import LoraModelDefaultSettings
from invokeai.backend.model_manager.configs.main import MainModelDefaultSettings
from invokeai.backend.model_manager.taxonomy import (
AnimaVariantType,
BaseModelType,
ClipVariantType,
Flux2VariantType,
Expand All @@ -36,6 +37,7 @@
PiDDecoderVariantType,
Qwen3VariantType,
Qwen3VLVariantType,
Qwen35VariantType,
QwenImageVariantType,
SchedulerPredictionType,
WanLoRAVariantType,
Expand Down Expand Up @@ -152,6 +154,8 @@ def validate_source_url(cls, v: Any) -> Optional[str]:
| WanLoRAVariantType
| Qwen3VariantType
| Qwen3VLVariantType
| Qwen35VariantType
| AnimaVariantType
| Krea2VariantType
| MiniMaxH3VariantType
| LTX2VariantType
Expand Down
Loading
Loading