Skip to content

feat(anima): support Anima-2.9B and Anima-3.8B finetunes - #9639

Merged
joshistoast merged 7 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/anima-expanded-finetunes
Oct 2, 2026
Merged

joshistoast merged 7 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/anima-expanded-finetunes

Conversation

@Pfannkuchensack

@Pfannkuchensack Pfannkuchensack commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

Two community finetunes of Anima now install and generate correctly: Anima-2.9B and Anima-3.8B.

Anima-2.9B (40 DiT blocks instead of 28) already installed as Anima, but the loader built the transformer from a hard-coded 28-block config and loaded with strict=False. It filled the first 28 blocks and dropped the other 12 (240 tensors) as unexpected keys, logged at DEBUG only. The result was a degraded image and no error. The loader now reads the depth from the checkpoint and refuses gaps. The int8_tensorwise build from the same repo loads too, through install_int8_convrot_layers.

Anima-3.8B v1.1 (52 blocks) bundles a "semantic connector" into its checkpoint that reads a second text encoder, Qwen3.5 4B. The connector is timestep-aware, so its context changes at every step. Until now the checkpoint installed as Anima, loaded 28 blocks and dropped the connector. Now:

  • AnimaVariantType (anima_qwen3, anima_qwen35), detected from the bundled anima_v2_connector.* keys
  • new ModelType.Qwen35Encoder (qwen3_5_encoder): single-file config and loader built on the transformers qwen3_5 decoder layers, with the tokenizer vendored from Qwen/Qwen3.5-4B @ 851bf6e (Apache-2.0). The encoder checkpoint is not a stock export: it ships layer 31 without its MLP, and the encoder reproduces that exactly.
  • a port of the connector (invokeai/backend/anima/semantic_connector.py), whose module names match the checkpoint
  • Prompt - Anima encodes Qwen3.5 when its new optional input is connected. Denoise - Anima recomputes the connector context at every step, for the positive, the negative and every regional conditioning, with a float32 sigma. The connector scales sigma by 1000 before embedding it, and bf16 rounding would shift its high frequencies.
  • webv2: a Qwen3.5 Encoder slot shown only for the anima_qwen35 variant, plus graph, regional guidance and recall wiring
  • the earlier Anima-3.8B v1.0 transformer, which needs a separate adapter file, is refused at install with an InvalidMatchError that names v1.1

LoRAs and ControlNet-LLLite adapters. Both finetunes insert their new blocks between the original ones, so an adapter trained on Anima patched different blocks from the third one on. Measured from the checkpoints:

  • Anima-2.9B holds all 28 blocks of Anima base-v1.0 bit for bit, with its 12 new blocks at 2, 5, 8, …, 36.
  • Anima-3.8B holds the 40 blocks of 2.9B nearly unchanged (cosine similarity ≥ 0.9997, 11 bit-identical), with 12 new blocks at 3, 7, …, 47.

invokeai/backend/anima/block_layout.py records the tables. A LoRA's DiT keys and an LLLite adapter's bindings now move to the blocks of the model they were trained on: Anima adapters on 2.9B and 3.8B, and 2.9B adapters on 3.8B. The adapter's depth is read off the highest block it addresses.

Also included:

  • starter models for both finetunes (2.9B bf16 and int8, 3.8B v1.1) and the Qwen3.5 encoder, with the non-commercial license noted
  • the Anima user guide page
  • three additions to new-model-integration.mdx, from the gaps this change hit:
    • a new section, Extending an existing architecture
    • a section on persisted records and migrations
    • corrections to the webv2 checklists

Related Issues / Discussions

QA Instructions

Automated

  • uv run pytest -n 8 tests/backend/model_manager tests/backend/architectures tests/backend/anima tests/backend/qwen3_5 tests/backend/util tests/backend/patches tests/app/invocations tests/app/services/shared tests/app/services/model_records tests/test_package_data.py: 6254 passed, 1 failed. The failure, test_16_channel_vae_loader.py::test_an_ldm_layout_file_is_converted_and_loses_no_tensor, reads an LFS fixture that was not pulled in the test worktree; it is unrelated.
  • ruff check . and ruff format --check .: clean.
  • webv2: lint:tsc, lint:oxc, format:check and architecture:check pass. Vitest over src/features/generation, src/workbench, src/features/models and src/features/workflow: 7277 passed, 2 failed. Both failures are in image-map (clusterStats, indexProgress) and come from German-locale digit grouping on the test machine; unrelated.
  • Regenerated: openapi.json/schema.ts, the capabilities fixture, and the graph-coverage snapshot. The docs build passes, and check-docs-data shows no diff.

After review

  • Startup import: the prompt node imports the Qwen3.5 encoder where it runs, not at module level. That spares every app start transformers' qwen3_5 modeling module, about 0.25 s. A subprocess test fails on the previous module-level import.

  • Migration on real databases: the full chain ran on copies of three populated databases, using init_db as the app does, and every model record read back afterwards.

    • A v7 dev root with Anima base, preview and a tcfp8 build: all three became anima_qwen3.
    • A v7 dev root with Anima base fp8: it became anima_qwen3.
    • The production database of a v6 install, 69 MiB with 75 models: the chain ran in 0.3 s and all 75 records read back. It holds no Anima main model.
    • Two Ideogram-4 GGUF records from another branch's build were invalid before and after; they are unrelated.
  • FP8 Storage and the connector's timestep path (Anima-3.8B, 832×1216, 30 steps, three seeds; every run fully resident and bit-reproducible):

    PSNR against bf16
    timestep path kept in bf16 (the previous exception) 8.4 / 10.3 / 9.6 dB
    everything in FP8 8.4 / 10.2 / 9.6 dB
    whole connector in bf16 8.3 / 10.1 / 9.6 dB
    Anima base, FP8 against bf16 15.9 / 15.2 / 14.0 dB

    The exception changed nothing measurable and cost 152 MiB (159M parameters), so it is gone. FP8 Storage changes Anima-3.8B's composition through its DiT blocks, not the connector; the user guide now says so.

  • LLLite through the block layout (anima-lllite-any-test-like-v2, trained on Anima). Control: a grayscale render. Prompt: different content. The metric is the edge correlation between output and control:

    no adapter adapter by index (before) adapter on the kept blocks (now)
    Anima base 0.126 0.451 0.451
    Anima-2.9B 0.121 0.187 0.474
    Anima-3.8B 0.081 0.106 0.421

    By index, the finetunes all but ignored the adapter; on the kept blocks they follow it as Anima base does. The LoRA path shares the tables and is unit-tested (key moves, untouched LLM adapter and text encoder keys, cached LoRA not modified); no Anima LoRA was available locally for an end-to-end run.

Numerical checks against the reference implementation, on the real weights

  • Qwen3.5 encoder, fp32: relative L2 ~1.5e-6 at layers 7/15/23/31, for 27- and 114-token prompts. In bf16 the error is ~1e-2, the same as the reference's own bf16 error.
  • Connector: bit-identical (relative L2 0.0) at sigma 1.0, 0.75, 0.3 and 0.02. An empty prompt (one masked Qwen3.5 token) also matches the reference, on both the math and the efficient SDPA backends.

End to end: RTX 4090, fresh root, models installed in place through the API, 832×1216, 30 steps, Euler, seed 1234.

Model Result
Anima Base 1.0 unchanged (regression check)
Anima-2.9B bf16 coherent, all 40 blocks loaded
Anima-2.9B int8 nearly identical to bf16; 640 of 640 layers kept int8, 2.9 GB in VRAM instead of 5.6 GB
Anima-3.8B v1.1, CFG 6 coherent, with clearly better spatial and object binding
Anima-3.8B, empty negative prompt coherent
Anima-3.8B without a Qwen3.5 encoder refused by the model loader with a message naming the missing encoder

Speed, single runs only (not a benchmark): ~1.7 it/s for 2.9B against ~1.3 it/s for 3.8B, which also runs the connector every step.

Browser (built webv2 against the same server):

  • The Qwen3.5 slot appears only when Anima-3.8B is selected. Invoke stays disabled until both encoders are chosen.
  • Generation works, and the image metadata records qwen3_5_encoder.
  • Reset all to model defaults sets 40 steps / CFG 6, the settings of the author's reference workflow.

Review

Remaining risks and limitations:

  • The block layout tables are keyed by depth (40, 52). These are the only depth-expanded Anima models today. A future finetune with one of these depths but another layout would need its own entry.
  • An adapter's own depth is read off the highest block it addresses. An adapter trained on Anima-2.9B that addresses only blocks below 28 would be taken for an Anima adapter.
  • On Anima-3.8B, the blocks an Anima adapter lands on were changed slightly, and the semantic connector conditions every block. Adapters work there (see the LLLite table), but results can differ more than on 2.9B.
  • The reference masks every Qwen3.5 token from the first id 151643 on. That id is the Qwen3 padding token, but in Qwen3.5's vocabulary it is the ordinary token " 내용". This behavior is not reproduced; empty prompts behave like the reference.
  • The author's workflow samples with res_multistep and a beta schedule, which Anima's scheduler set does not offer.
  • The int8 path is new for Anima and has been validated on the Anima-2.9B int8 file only.

Compatibility / Rollout

  • Migration 2026_10_01_add_anima_variant (depends on 2026_09_26_add_workflow_revision): variant is now a required field on Anima main records. Without the migration, every installed Anima model would fail validation on read and vanish from the model list. The migration reads each checkpoint's header, so a 3.8B installed on an earlier v7 build becomes anima_qwen35. A record whose file cannot be read gets anima_qwen3.
  • New persisted values: model type qwen3_5_encoder, variants anima_qwen3, anima_qwen35 and qwen3_5_4b.
  • Node versions: anima_model_loader 1.5.0 and anima_text_encoder 1.5.0. The new inputs and outputs are optional, so existing workflows keep validating.
  • Packaging: invokeai.backend.qwen3_5 is added to package-data. A new tests/test_package_data.py checks that every vendored *.json/*.json.gz under invokeai/backend ships in the wheel.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

- Read the DiT depth from the checkpoint instead of hard-coding 28 blocks
  (Anima-2.9B silently lost 12 of 40 blocks); load int8_tensorwise builds
- Anima-3.8B v1.1: AnimaVariantType, bundled semantic connector recomputed
  per denoise step, new Qwen3.5 encoder model type with vendored tokenizer
- Migration 2026_10_01_add_anima_variant for installed Anima records
- webv2: Qwen3.5 encoder slot and graph wiring for the anima_qwen35 variant
- Starter models, user guide, and integration guide section on extending
  an existing architecture; package-data guard test
@github-actions github-actions Bot added CI-CD Continuous integration / Continuous delivery docker api python PRs that change python files Root invocations PRs that change invocations backend PRs that change backend files services PRs that change app services frontend-deps PRs that change frontend dependencies frontend PRs that change frontend files docs PRs that change docs labels Oct 1, 2026
@github-actions github-actions Bot added python-tests PRs that change python tests python-deps PRs that change python dependencies labels Oct 2, 2026
@Pfannkuchensack
Pfannkuchensack marked this pull request as ready for review October 2, 2026 01:53

@joshistoast joshistoast left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • anima_text_encoder.py now imports transformers.models.qwen3_5 at module level, so it loads at every app startup. There's precedent (segment_anything.py), but a lazy import would avoid it.
  • The author's own disclosures still hold: the migration hasn't been run on a populated production database, the FP8 handling of the connector's timestep path is unmeasured, existing Anima LoRAs and LLLite adapters will patch the wrong blocks on the finetunes (documented, not remapped), and #9506 will need a small merge with this.

The prompt node imported invokeai.backend.qwen3_5 at module level, which pulls in transformers'
qwen3_5 modeling module (~0.25 s) on every app start for a node input only Anima-3.8B uses. It is now
imported where the encoder runs, as the loader already did. A subprocess test pins that importing the
node leaves transformers.models.qwen3_5 unloaded.
…finetunes kept

Anima-2.9B and Anima-3.8B insert new DiT blocks between the existing ones, so an adapter trained on Anima
patched different blocks from the third one on. Measured from the checkpoints:
- Anima-2.9B holds all 28 blocks of Anima base-v1.0 bit for bit; its 12 new blocks sit at
  2, 5, 8, ..., 36.
- Anima-3.8B holds the 40 blocks of 2.9B nearly unchanged (cosine similarity >= 0.9997, 11 bit-identical),
  with 12 new blocks at 3, 7, ..., 47. Ten Anima base blocks are still bit-identical there, all where the
  composed table puts them.

invokeai/backend/anima/block_layout.py records the tables. A LoRA's DiT block keys and an LLLite
adapter's bindings move to the blocks of the model they were trained on, whose depth is read off the
highest block they address: Anima adapters on 2.9B and 3.8B, 2.9B adapters on 3.8B. The cached LoRA is
not modified; the denoise node applies a re-keyed copy that shares its layers.
The connector's timestep path (159M parameters) was kept in bf16 by analogy with t_embedder, unmeasured.
Measured on Anima-3.8B, 30 steps, three seeds, PSNR against bf16:
- with that exception: 8.4/10.3/9.6 dB;
- without it: 8.4/10.2/9.6 dB;
- with the whole connector in bf16: 8.3/10.1/9.6 dB.
The exception changes nothing and costs 152 MiB, so it goes. FP8 Storage moves Anima-3.8B's composition
through its DiT blocks (Anima base stays at 14-16 dB); the user guide now says so.

@joshistoast joshistoast left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

works good

@joshistoast
joshistoast enabled auto-merge October 2, 2026 04:07
@joshistoast
joshistoast merged commit 77a7611 into invoke-ai:main Oct 2, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

api backend PRs that change backend files CI-CD Continuous integration / Continuous delivery docker docs PRs that change docs frontend PRs that change frontend files frontend-deps PRs that change frontend dependencies invocations PRs that change invocations python PRs that change python files python-deps PRs that change python dependencies python-tests PRs that change python tests Root services PRs that change app services

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants