Skip to content
This repository was archived by the owner on Oct 4, 2026. It is now read-only.

feat(anima): support Anima-2.9B and Anima-3.8B finetunes - #394

Closed
Pfannkuchensack wants to merge 1 commit into
upstream-mergefrom
feat/anima-expanded-finetunes
Closed

Pfannkuchensack wants to merge 1 commit into
upstream-mergefrom
feat/anima-expanded-finetunes

Conversation

@Pfannkuchensack

Copy link
Copy Markdown
Member

Summary

Two community finetunes of Anima now install and generate correctly: Anima-2.9B and Anima-3.8B.

Anima-2.9B (40 DiT blocks instead of 28) already installed as Anima, but the loader built the transformer from a hard-coded 28-block config and loaded with strict=False. It filled the first 28 blocks and dropped the other 12 (240 tensors) as unexpected keys, logged at DEBUG only. The result was a degraded image and no error. The loader now reads the depth from the checkpoint and refuses gaps. The int8_tensorwise build from the same repo loads too, through install_int8_convrot_layers.

Anima-3.8B v1.1 (52 blocks) bundles a "semantic connector" into its checkpoint that reads a second text encoder, Qwen3.5 4B. The connector is timestep-aware, so its context changes at every step. Until now the checkpoint installed as Anima, loaded 28 blocks and dropped the connector. Now:

  • AnimaVariantType (anima_qwen3, anima_qwen35), detected from the bundled anima_v2_connector.* keys
  • new ModelType.Qwen35Encoder (qwen3_5_encoder): single-file config and loader built on the transformers qwen3_5 decoder layers, with the tokenizer vendored from Qwen/Qwen3.5-4B @ 851bf6e (Apache-2.0). The encoder checkpoint is not a stock export: it ships layer 31 without its MLP, and the encoder reproduces that exactly.
  • a port of the connector (invokeai/backend/anima/semantic_connector.py), whose module names match the checkpoint
  • Prompt - Anima encodes Qwen3.5 when its new optional input is connected. Denoise - Anima recomputes the connector context at every step, for the positive, the negative and every regional conditioning, with a float32 sigma. The connector scales sigma by 1000 before embedding it, and bf16 rounding would shift its high frequencies.
  • webv2: a Qwen3.5 Encoder slot shown only for the anima_qwen35 variant, plus graph, regional guidance and recall wiring
  • the earlier Anima-3.8B v1.0 transformer, which needs a separate adapter file, is refused at install with an InvalidMatchError that names v1.1

Also included:

  • starter models for both finetunes (2.9B bf16 and int8, 3.8B v1.1) and the Qwen3.5 encoder, with the non-commercial license noted
  • the Anima user guide page
  • three additions to new-model-integration.mdx, from the gaps this change hit:
    • a new section, Extending an existing architecture
    • a section on persisted records and migrations
    • corrections to the webv2 checklists

Related Issues / Discussions

QA Instructions

Automated

  • uv run pytest -n 8 tests/backend/model_manager tests/backend/architectures tests/backend/anima tests/backend/qwen3_5 tests/backend/util tests/backend/patches tests/app/invocations tests/app/services/shared tests/app/services/model_records tests/test_package_data.py: 6254 passed, 1 failed. The failure, test_16_channel_vae_loader.py::test_an_ldm_layout_file_is_converted_and_loses_no_tensor, reads an LFS fixture that was not pulled in the test worktree; it is unrelated.
  • ruff check . and ruff format --check .: clean.
  • webv2: lint:tsc, lint:oxc, format:check and architecture:check pass. Vitest over src/features/generation, src/workbench, src/features/models and src/features/workflow: 7277 passed, 2 failed. Both failures are in image-map (clusterStats, indexProgress) and come from German-locale digit grouping on the test machine; unrelated.
  • Regenerated: openapi.json/schema.ts, the capabilities fixture, and the graph-coverage snapshot. The docs build passes, and check-docs-data shows no diff.

Numerical checks against the reference implementation, on the real weights

  • Qwen3.5 encoder, fp32: relative L2 ~1.5e-6 at layers 7/15/23/31, for 27- and 114-token prompts. In bf16 the error is ~1e-2, the same as the reference's own bf16 error.
  • Connector: bit-identical (relative L2 0.0) at sigma 1.0, 0.75, 0.3 and 0.02. An empty prompt (one masked Qwen3.5 token) also matches the reference, on both the math and the efficient SDPA backends.

End to end: RTX 4090, fresh root, models installed in place through the API, 832×1216, 30 steps, Euler, seed 1234.

Model Result
Anima Base 1.0 unchanged (regression check)
Anima-2.9B bf16 coherent, all 40 blocks loaded
Anima-2.9B int8 nearly identical to bf16; 640 of 640 layers kept int8, 2.9 GB in VRAM instead of 5.6 GB
Anima-3.8B v1.1, CFG 6 coherent, with clearly better spatial and object binding
Anima-3.8B, empty negative prompt coherent
Anima-3.8B without a Qwen3.5 encoder refused by the model loader with a message naming the missing encoder

Speed, single runs only (not a benchmark): ~1.7 it/s for 2.9B against ~1.3 it/s for 3.8B, which also runs the connector every step.

Browser (built webv2 against the same server):

  • The Qwen3.5 slot appears only when Anima-3.8B is selected. Invoke stays disabled until both encoders are chosen.
  • Generation works, and the image metadata records qwen3_5_encoder.
  • Reset all to model defaults sets 40 steps / CFG 6, the settings of the author's reference workflow.

The migration is unit-tested against an in-memory models table and ran on a fresh database. It has not been run against a populated production database.

Review

Remaining risks and limitations:

  • LoRAs and ControlNet-LLLite adapters trained on Anima address blocks by index. The finetunes insert blocks between the original ones, so from the third block on such an adapter patches different blocks. This is documented in the user guide; adapters are not remapped.
  • FP8 Storage keeps the connector's timestep path (time_mlp, time_modulation) in bf16, by analogy with t_embedder. This is not measured, and no fp8 quality comparison was run on either finetune.
  • The reference masks every Qwen3.5 token from the first id 151643 on. That id is the Qwen3 padding token, but in Qwen3.5's vocabulary it is the ordinary token " 내용". This behavior is not reproduced; empty prompts behave like the reference.
  • The author's workflow samples with res_multistep and a beta schedule, which Anima's scheduler set does not offer.
  • The int8 path is new for Anima and has been validated on the Anima-2.9B int8 file only.

Compatibility / Rollout

  • Migration 2026_10_01_add_anima_variant (depends on 2026_09_26_add_workflow_revision): variant is now a required field on Anima main records. Without the migration, every installed Anima model would fail validation on read and vanish from the model list. The migration reads each checkpoint's header, so a 3.8B installed on an earlier v7 build becomes anima_qwen35. A record whose file cannot be read gets anima_qwen3.
  • New persisted values: model type qwen3_5_encoder, variants anima_qwen3, anima_qwen35 and qwen3_5_4b.
  • Node versions: anima_model_loader 1.5.0 and anima_text_encoder 1.5.0. The new inputs and outputs are optional, so existing workflows keep validating.
  • Packaging: invokeai.backend.qwen3_5 is added to package-data. A new tests/test_package_data.py checks that every vendored *.json/*.json.gz under invokeai/backend ships in the wheel.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

- Read the DiT depth from the checkpoint instead of hard-coding 28 blocks
  (Anima-2.9B silently lost 12 of 40 blocks); load int8_tensorwise builds
- Anima-3.8B v1.1: AnimaVariantType, bundled semantic connector recomputed
  per denoise step, new Qwen3.5 encoder model type with vendored tokenizer
- Migration 2026_10_01_add_anima_variant for installed Anima records
- webv2: Qwen3.5 encoder slot and graph wiring for the anima_qwen35 variant
- Starter models, user guide, and integration guide section on extending
  an existing architecture; package-data guard test
@Pfannkuchensack

Copy link
Copy Markdown
Member Author

Opened against the wrong repository. Moved to invoke-ai#9639 (stacked on invoke-ai#9613).

@Pfannkuchensack
Pfannkuchensack deleted the feat/anima-expanded-finetunes branch October 1, 2026 23:00
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant