feat(anima): support Anima-2.9B and Anima-3.8B finetunes - #9639
Merged
joshistoast merged 7 commits intoOct 2, 2026
Merged
Conversation
- Read the DiT depth from the checkpoint instead of hard-coding 28 blocks (Anima-2.9B silently lost 12 of 40 blocks); load int8_tensorwise builds - Anima-3.8B v1.1: AnimaVariantType, bundled semantic connector recomputed per denoise step, new Qwen3.5 encoder model type with vendored tokenizer - Migration 2026_10_01_add_anima_variant for installed Anima records - webv2: Qwen3.5 encoder slot and graph wiring for the anima_qwen35 variant - Starter models, user guide, and integration guide section on extending an existing architecture; package-data guard test
5 of 7 tasks
Pfannkuchensack
marked this pull request as ready for review
October 2, 2026 01:53
Pfannkuchensack
requested review from
JPPhoto,
blessedcoolant,
dunkeroni and
lstein
as code owners
October 2, 2026 01:53
joshistoast
requested changes
Oct 2, 2026
joshistoast
left a comment
Collaborator
There was a problem hiding this comment.
- anima_text_encoder.py now imports transformers.models.qwen3_5 at module level, so it loads at every app startup. There's precedent (segment_anything.py), but a lazy import would avoid it.
- The author's own disclosures still hold: the migration hasn't been run on a populated production database, the FP8 handling of the connector's timestep path is unmeasured, existing Anima LoRAs and LLLite adapters will patch the wrong blocks on the finetunes (documented, not remapped), and #9506 will need a small merge with this.
The prompt node imported invokeai.backend.qwen3_5 at module level, which pulls in transformers' qwen3_5 modeling module (~0.25 s) on every app start for a node input only Anima-3.8B uses. It is now imported where the encoder runs, as the loader already did. A subprocess test pins that importing the node leaves transformers.models.qwen3_5 unloaded.
…finetunes kept Anima-2.9B and Anima-3.8B insert new DiT blocks between the existing ones, so an adapter trained on Anima patched different blocks from the third one on. Measured from the checkpoints: - Anima-2.9B holds all 28 blocks of Anima base-v1.0 bit for bit; its 12 new blocks sit at 2, 5, 8, ..., 36. - Anima-3.8B holds the 40 blocks of 2.9B nearly unchanged (cosine similarity >= 0.9997, 11 bit-identical), with 12 new blocks at 3, 7, ..., 47. Ten Anima base blocks are still bit-identical there, all where the composed table puts them. invokeai/backend/anima/block_layout.py records the tables. A LoRA's DiT block keys and an LLLite adapter's bindings move to the blocks of the model they were trained on, whose depth is read off the highest block they address: Anima adapters on 2.9B and 3.8B, 2.9B adapters on 3.8B. The cached LoRA is not modified; the denoise node applies a re-keyed copy that shares its layers.
The connector's timestep path (159M parameters) was kept in bf16 by analogy with t_embedder, unmeasured. Measured on Anima-3.8B, 30 steps, three seeds, PSNR against bf16: - with that exception: 8.4/10.3/9.6 dB; - without it: 8.4/10.2/9.6 dB; - with the whole connector in bf16: 8.3/10.1/9.6 dB. The exception changes nothing and costs 152 MiB, so it goes. FP8 Storage moves Anima-3.8B's composition through its DiT blocks (Anima base stays at 14-16 dB); the user guide now says so.
joshistoast
enabled auto-merge
October 2, 2026 04:07
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two community finetunes of Anima now install and generate correctly: Anima-2.9B and Anima-3.8B.
Anima-2.9B (40 DiT blocks instead of 28) already installed as Anima, but the loader built the transformer from a hard-coded 28-block config and loaded with
strict=False. It filled the first 28 blocks and dropped the other 12 (240 tensors) as unexpected keys, logged at DEBUG only. The result was a degraded image and no error. The loader now reads the depth from the checkpoint and refuses gaps. The int8_tensorwise build from the same repo loads too, throughinstall_int8_convrot_layers.Anima-3.8B v1.1 (52 blocks) bundles a "semantic connector" into its checkpoint that reads a second text encoder, Qwen3.5 4B. The connector is timestep-aware, so its context changes at every step. Until now the checkpoint installed as Anima, loaded 28 blocks and dropped the connector. Now:
AnimaVariantType(anima_qwen3,anima_qwen35), detected from the bundledanima_v2_connector.*keysModelType.Qwen35Encoder(qwen3_5_encoder): single-file config and loader built on thetransformersqwen3_5decoder layers, with the tokenizer vendored fromQwen/Qwen3.5-4B@851bf6e(Apache-2.0). The encoder checkpoint is not a stock export: it ships layer 31 without its MLP, and the encoder reproduces that exactly.invokeai/backend/anima/semantic_connector.py), whose module names match the checkpointanima_qwen35variant, plus graph, regional guidance and recall wiringInvalidMatchErrorthat names v1.1LoRAs and ControlNet-LLLite adapters. Both finetunes insert their new blocks between the original ones, so an adapter trained on Anima patched different blocks from the third one on. Measured from the checkpoints:
invokeai/backend/anima/block_layout.pyrecords the tables. A LoRA's DiT keys and an LLLite adapter's bindings now move to the blocks of the model they were trained on: Anima adapters on 2.9B and 3.8B, and 2.9B adapters on 3.8B. The adapter's depth is read off the highest block it addresses.Also included:
new-model-integration.mdx, from the gaps this change hit:Related Issues / Discussions
invokeai/app/invocations/anima_text_encoder.py. In v7 the node lives atinvokeai/app/invocations/text_encoder/anima_text_encoder.py, and this PR changes it as well: the new Qwen3.5 input, and the conditioning fields it stores. feat(anima): add prompt weighting support #9506 needs a v7 port either way, and that port picks up this file.QA Instructions
Automated
uv run pytest -n 8 tests/backend/model_manager tests/backend/architectures tests/backend/anima tests/backend/qwen3_5 tests/backend/util tests/backend/patches tests/app/invocations tests/app/services/shared tests/app/services/model_records tests/test_package_data.py: 6254 passed, 1 failed. The failure,test_16_channel_vae_loader.py::test_an_ldm_layout_file_is_converted_and_loses_no_tensor, reads an LFS fixture that was not pulled in the test worktree; it is unrelated.ruff check .andruff format --check .: clean.lint:tsc,lint:oxc,format:checkandarchitecture:checkpass. Vitest oversrc/features/generation,src/workbench,src/features/modelsandsrc/features/workflow: 7277 passed, 2 failed. Both failures are in image-map (clusterStats,indexProgress) and come from German-locale digit grouping on the test machine; unrelated.openapi.json/schema.ts, the capabilities fixture, and the graph-coverage snapshot. The docs build passes, andcheck-docs-datashows no diff.After review
Startup import: the prompt node imports the Qwen3.5 encoder where it runs, not at module level. That spares every app start transformers'
qwen3_5modeling module, about 0.25 s. A subprocess test fails on the previous module-level import.Migration on real databases: the full chain ran on copies of three populated databases, using
init_dbas the app does, and every model record read back afterwards.anima_qwen3.anima_qwen3.FP8 Storage and the connector's timestep path (Anima-3.8B, 832×1216, 30 steps, three seeds; every run fully resident and bit-reproducible):
The exception changed nothing measurable and cost 152 MiB (159M parameters), so it is gone. FP8 Storage changes Anima-3.8B's composition through its DiT blocks, not the connector; the user guide now says so.
LLLite through the block layout (
anima-lllite-any-test-like-v2, trained on Anima). Control: a grayscale render. Prompt: different content. The metric is the edge correlation between output and control:By index, the finetunes all but ignored the adapter; on the kept blocks they follow it as Anima base does. The LoRA path shares the tables and is unit-tested (key moves, untouched LLM adapter and text encoder keys, cached LoRA not modified); no Anima LoRA was available locally for an end-to-end run.
Numerical checks against the reference implementation, on the real weights
End to end: RTX 4090, fresh root, models installed in place through the API, 832×1216, 30 steps, Euler, seed 1234.
Speed, single runs only (not a benchmark): ~1.7 it/s for 2.9B against ~1.3 it/s for 3.8B, which also runs the connector every step.
Browser (built webv2 against the same server):
qwen3_5_encoder.Review
Remaining risks and limitations:
res_multistepand a beta schedule, which Anima's scheduler set does not offer.Compatibility / Rollout
2026_10_01_add_anima_variant(depends on2026_09_26_add_workflow_revision):variantis now a required field on Anima main records. Without the migration, every installed Anima model would fail validation on read and vanish from the model list. The migration reads each checkpoint's header, so a 3.8B installed on an earlier v7 build becomesanima_qwen35. A record whose file cannot be read getsanima_qwen3.qwen3_5_encoder, variantsanima_qwen3,anima_qwen35andqwen3_5_4b.anima_model_loader1.5.0 andanima_text_encoder1.5.0. The new inputs and outputs are optional, so existing workflows keep validating.invokeai.backend.qwen3_5is added topackage-data. A newtests/test_package_data.pychecks that every vendored*.json/*.json.gzunderinvokeai/backendships in the wheel.Checklist
What's Newcopy (if doing a release after this PR)