Skip to content

Version 7.0 Alpha Upgrade - #9613

Merged
lstein merged 3651 commits into
mainfrom
upstream-merge
Oct 2, 2026
Merged

lstein merged 3651 commits into
mainfrom
upstream-merge

Conversation

@joshistoast

@joshistoast joshistoast commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator
v7

InvokeAI 7 Alpha

320+ PRs 2,600+ commits 5,600+ files alpha

Important

Heads up before you upgrade: v7 needs Python 3.12, the new UI is the default (the old one is still there with --web-legacy), and the database migrations only go one way. Invoke backs up your database first. Details are in Compatibility / Rollout.

Summary

Okay. This is InvokeAI 7.

It's been a lot of work. I could try to list everything that changed, but you'd be scrolling until next Tuesday. So instead, here are the three ideas that got this whole thing started.

🗂️ Projects, so you can wander off and come back

I wanted to be able to start something new without the nagging feeling I was about to lose the thing I was working on before. So now everything lives in a project: your canvas, your settings, your workflows, your queue history, and a board of its own for everything it makes.

Keep a few open as tabs. Close one and go do something else. Come back next week from the new Launchpad and it's right where you left it.

It autosaves! Edits go to your browser first and then sync to the server. That means losing wifi, a crashing tab, or the same project open in two windows won't mean the end for your recent work. If the client/server versions end up disagreeing, you'll get both. Workflows belong to their project now, with their own undo history, and you can export the whole project as an .invk file- and carry it over to another machine, or just for safe keeping.

⚡ One page shell

Previously, each set of curated features were divided effectively into separate pages with a load screen in between, but that's history. v7 is a single workbench built out of widgets: Generate, Canvas, Gallery, Preview, Workflows, Video, Upscale, Image Map, etc. Drag them where you like, flip between layout presets, pop one out into its own floating window, or just hit Mod+K and type what you want.

And we wanted it fast, and staying fast. Widgets don't load until you open them. Every page has a budget for how much it downloads and how many requests it makes, and CI will flatly refuse a PR that goes over. Nothing gets heavier unless someone decides it should. We're a little smug about this one, and when you try it I think you'll see why.

🎨 A canvas you can actually paint on

The old canvas was built on Konva, and it was great at what it was built for: inpainting. We wrote our own engine from scratch because we wanted real, unrestrained, fast image editing, with generation built right in.

So: layer groups (nest away), 16 blend modes, non-destructive adjustments, and a layers panel that still holds up at 2,000+ layers, and you can drive it entirely from the keyboard. There's a pressure-sensitive brush, marquee and lasso selections, gradients, shapes, a text tool that takes your own fonts, a magnifying color picker, and click-to-select objects. You also get a proper history panel you can scrub back through, PSD export for when your art needs to go meet Photoshop, and much more in the future to come.

Inpainting, regional guidance, control layers and staging are all still there, living in the same document as everything else.


Those three are why we started. Then, well... it kept going. There's new video generation models (MiniMax H3, LTX-2.5), a rebuilt gallery with semantic search and a zoomable Image Map of everything you've ever made, a modernized workflow editor with loops and batch generators, a reworked execution engine, a pile of new models and quantization formats, and hundreds of fixes. The rest is below. Go poke at it! 🎉

🔍 The rest of the tour

🎬 Video, now with sound

There's a whole Video panel now, with its own prompt, a built-in Video layout (Alt+3) and a Launchpad tile to get you there. It reshapes itself around whichever model you pick, and if Invoke is greyed out it tells you why.

  • Two new model families. MiniMax H3 makes video with audio. Its Ref2VA mode takes up to three video and nine image references (each gets a <Video 1>/<Picture 1> badge you can use in your prompt), and turbo LoRAs get it done in 4–8 steps. LTX-2.5 does stereo audio too, plus two-stage upscale-and-refine to 1536p, audio→video and video→audio, first/last-frame keyframes, interpolation and extend.
  • Wan 2.2 (T2V, I2V, TI2V-5B, Lightning) gets the same panel treatment.
  • Conditioning media that's actually nice to use. Trim clips with live-frame thumbnails, reorder references from the keyboard, preview a trimmed window right in Preview, or drop in audio on its own.
  • Your videos remember how they were made. Generation metadata is embedded in the MP4, so you can download a clip, upload it to a different install and still recall it. There's also a new external video recall API for remixing from scripts.
  • Uploads just work. .mov, HEVC and ProRes get converted to H.264 MP4 on the way in, and audio files become waveform videos. Gallery thumbnails skip fades and title cards to find a frame that's worth looking at.
  • Bundled workflows cover every LTX-2.5 mode and the MiniMax H3 recipes.
🖼️ A gallery that knows what's in it
  • Rebuilt from the ground up. It has reliable ordering and pagination, a starred strip above the grid, drag-to-board, live updates when media arrives from another tab or script, and an "In progress" section that follows your generations across boards.
  • Semantic search. Type what you're looking for ("red car at night") or drop in an image to find similar ones. Results re-rank as you type.
  • The Image Map. Every image and video you've made, laid out by similarity on a zoomable map. It finds clusters on its own and names them ("portraits", "foggy forests"…). Click anything to jump to it in the gallery, or select a whole cluster at once. It stays quick on galleries of 170k+ images.
  • A gallery picker on every image slot, plus "Find in Gallery" on any thumbnail so you can see where something came from.
  • Intermediates manager in Settings. It shows how much space each project's temporary files take up and cleans them out safely, without touching anything an open project still needs.
  • Touch support all over: hold to drag, swipe between images, pinch to zoom.
🧩 Workflows, with loops
  • The node editor reached parity with legacy. Fields keep what you type, collections can be edited inline, and outdated nodes offer "Update node" or "Update all".
  • A library you can browse. It's a card grid with thumbnails and model badges. Each workflow shows what it needs, and you can install all the missing models in one click.
  • Workflows live in your project. Open a library template and you get your own copy to mess with. The shared library only changes when you say Save to library.
  • For / ForReturn loops. Loop over a collection while carrying state between iterations, stop early, nest loops, and resume after a restart. They come with new Concat, Zip and Cartesian collection nodes.
  • Media fields take files directly. Drop, upload or pick from the gallery, and choose a frame for video inputs.
  • Call Saved Workflow, If and batch generators all got a round of fixes. You can also export a workflow as a PNG.
🧠 Models, quantization & GPUs

The headline models are MiniMax H3 and LTX-2.5, both new in v7. On top of that, a lot of work went into making the models you already have smaller, faster and less likely to fall over:

Format What it buys you
FP8 Compute (opt-in, Ada+) Runs fp8 checkpoints on the tensor cores
FP8 Storage Turned on at install for fp8 files, and scaled fp8 stays packed. FLUX.2 Klein 4B goes from 7.4 → 3.9 GB with identical output
int8_convrot One shared scheme for FLUX.1, FLUX.2 Klein, Z-Image, Krea 2, Ideogram 4, Qwen3 encoders, PiD, MiniMax H3 and LTX-2.5
nvfp4 / MXFP8 Comfy nvfp4 checkpoints stay packed, and MXFP8 files load
GGUF Qwen3-VL A quantized text encoder for Krea 2 and Ideogram 4
  • Smarter VRAM handling. Model loads no longer starve each other, freshly loaded models don't get evicted straight away, the budget counts memory the allocator can give back, and installing a pile of models no longer loads entire checkpoints into RAM.
  • Multi-GPU. Cancel stops mid-step and frees VRAM, and prompt tools and image indexing hop onto whichever GPU is idle.
  • ROCm. Moved to torch 2.13 + ROCm 7.2, with fixes for black images and long-clip crashes.
  • Tiled FLUX.1 VAE. Decoding stays around 1.2 GB however big the image is.
  • Works offline. Tokenizers and encoder configs ship with Invoke instead of being fetched from Hugging Face.
  • Engine internals. Each architecture now declares itself in one file, which the frontend reads, so the two can't drift apart. The execution engine was refactored so that loops, conditionals and nested workflows share one scheduler. The frontend contract hasn't changed.
✨ And a bunch more
  • Dynamic prompts, wildcards (with search and autocomplete) and prompt templates
  • A redesigned Generate panel, scrubbable number fields, and seed modes (random, fixed, increment, decrement)
  • An Upscale widget, now with FLUX.1
  • Paste media anywhere, an install queue panel, and a Diagnostics panel with JSON export
  • Translations, a high-contrast mode, searchable settings, feature hints, and a Catppuccin Mocha theme 🐱
  • Full tablet support

📈 By the numbers

320

pull requests

2,607

commits

5,684

files touched

+771k

lines added

88k

lines of new Python tests
%%{init: {"pie": {"textPosition": 0.75}} }%%
pie showData title Where the new lines went (thousands)
    "webv2 frontend" : 508
    "Tests" : 88
    "API, services & invocations" : 87
    "Inference & model management" : 29
    "Build, tooling & other" : 32
    "Docs" : 3
Loading
xychart-beta
    title "PRs merged per week"
    x-axis ["Jun 29", "Jul 6", "Jul 13", "Jul 20", "Jul 27", "Aug 3", "Aug 10", "Aug 17", "Aug 24", "Aug 31", "Sep 7", "Sep 14", "Sep 21", "Sep 28"]
    y-axis "PRs" 0 --> 60
    bar [5, 0, 4, 6, 8, 14, 21, 37, 36, 29, 40, 40, 56, 24]
Loading
timeline
    title Three months of v7
    June : v7 begins (#2)
         : Gallery progress and intermediates
    July : i18n, reference images
         : The new canvas lands (#9)
         : Upscale widget, command palette
         : State architecture rebuilt
         : Dynamic prompts and wildcards
    August : Launchpad home screen, .invk projects
           : Image Map and semantic search
           : Floating widget windows
           : MiniMax H3 video and the Video panel
           : Generate panel redesign
    September : Canvas overhaul with layers panel and editor panes
              : Execution engine refactor
              : fp8, nvfp4 and MXFP8 quantization
              : LTX-2.5 video with audio
              : Project-owned workflows, intermediates manager
              : webv2 becomes the default
Loading

Compatibility / Rollout

Warning

There's no downgrade path. Migrations are one-way, and older builds refuse a database that has migrations they don't know about. Invoke writes <db>_backup_<timestamp>.db before migrating; restore that file to go back.

Area What changes What to do
Python >=3.12,<3.13 (3.11 dropped) Upgrade Python. Remove any py3.11 required status checks
Default UI The new UI is served by default. --web-legacy serves the old UI, and --webv2 remains as an alias Nothing, unless you want the old UI
Package paths invokeai.frontend.web → invokeai.frontend.webv1. OpenAPI and typegen moved to invokeai/frontend/api Update anything that imports the old path
Database 21 new migrations: projects and boards, image index, fonts, wildcards, intermediates, queue receipts, workflow revisions… Each migration now runs inside its own transaction Automatic, after a backup
Migration 33 Upstream's migration_33 (image subfolder moves) moved to migration_34, and the fork's projects table took 33. Repair migrations cover databases that already ran upstream's 33 Nothing. Covered by the v6.14 upgrade test above
Config Only additions: db_synchronous, fp8_compute, image_index_*, fonts_*. force_tiled_decode now also applies to FLUX.1 Optional
Dependencies Adds umap-learn, scikit-learn and numba. The ROCm extra moves to torch 2.13 + ROCm 7.2 uv sync
API Nothing removed. New /projects, /image_map, /intermediates, /fonts, /wildcards, /recall/video and /api/v2/models/capabilities endpoints. Workflow records gain revision Regenerate API clients
Projects Document schema 3, canvas schema 3–4. The server answers 412 to clients too old to open a project Nothing
Node authors Built-in nodes moved into per-architecture packages, and the old import paths still work through forwarding shims Nothing yet. Move to the new paths when convenient

Merge strategy: please use a merge commit, not a squash, so the 2.6k commits keep their history and authorship.

gitGraph
    commit id: "v6.14"
    branch v7
    checkout v7
    commit id: "webv2 + canvas"
    checkout main
    commit id: "upstream work"
    checkout v7
    merge main id: "syncs ×13"
    commit id: "video, projects, Image Map"
    checkout main
    merge v7 id: "InvokeAI 7 🎉"
Loading

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

💜 Credits

Remarkably, v7 came from four people, 320 PRs and more late nights and all-nighters than we'll admit to.


@joshistoast

@lstein

@Pfannkuchensack

@JPPhoto

Huge thanks to everyone who tested early builds and told us what was broken.

JPPhoto and others added 30 commits September 28, 2026 02:00
Keep Call Saved Workflow roots eligible during child waits.
… counter

- The adapter-wide PDH counter cannot be used: measured with an 8 GiB holder, the holding process read 15.66 GiB and
  another read 0.176 GiB for the same instance at the same moment. It also shares the driver's basis, so subtracting
  torch's live figure invented ~2 GiB of foreign residency after every load/unload cycle
- The ceiling is min(total, max(budget, our CurrentUsage)): a budget below our own residency asks this process to
  trim itself, which the cache does for the reservation. A foreign holder that leaves the budget alone is documented
  as undetectable
- One PDH query per counter, per-item CStatus checked, usage clamped rather than dropped with the budget
- Interpolate decode peaks in bytes so a smaller image is never priced above a larger one
Moving items onto Uncategorized now uses the same drop path as named boards.
User-facing copy, counts and the layer save action now say Uploads.
Arrow-key navigation and scroll-to-session follow the new starred, in-progress, listing order.
Modified Enter/Space no longer stops at the tile or starts a keyboard drag.
Selected tiles, including in-progress ones, get a shared inset accent ring.
Unselected line-tab triggers use the ghost-button hover tint.
Invalid form controls and drop zones tint red on hover instead of repainting a neutral border.
Label and value stay legible over the fill.
The form draft lives in an external store read per section; generate values are exposed as a store with narrow selectors.
The rail becomes route links in visual order: Workspace, open projects, Manage; entries show a leave arrow on hover.
Set a thumbnail from the gallery or an upload, or remove it; the mock backend serves the thumbnail routes.
… add FLUX.1

The widget accepted only SD1.5 and SDXL, and its compiler threw without a tile
ControlNet, so an architecture without one could not reach it. A capability table
now says what each architecture has; the compiler builds the shared prefix once
and dispatches to a per-architecture builder that `satisfies` the same keys.

FLUX.1 gets a builder: both VAE ends tile, and the frame is fitted to the grid
`flux_denoise` accepts, because Spandrel rounds to 8 and the node demands 16.
Structure, the tile ControlNet picker and the negative prompt are hidden where
nothing consumes them; T5, CLIP and VAE are required where the loader needs them.
Update the reconnect viewport seed and cover a lone off-screen call-workflow node.
Folders are identified by their scaling/shift constants and single files by an explicit base, then the backbone their name or install source states; adds FLUX.1-diffusers and SD3 VAE configs, and an override now decides within a latent family instead of being overruled.
VAELoader loads SD3 single files in either layout, FLUX.1 folders outside float16, and flat VAE folders requested as SubModelType.VAE.
Reidentify passes the stored install source, so name evidence survives a re-probe.
Restore workflow editor viewport after preview
Resolve selected-branch iteration paths in If runtime handling and document direct output_collection consumers.
test: keep the nvfp4 loader forwards off torch's CPU bf16 GEMM

@lstein lstein left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Amazing work.

lstein and others added 19 commits October 1, 2026 14:09
docs: reorganize for v7 with prefix-ordered sidebar and new Users Guide pages
…-gaps

Moves the quantized formats to Model Formats and Families, model identification under Advanced Topics, and the Ideogram 4 and ERNIE-Image additions to their new per-family pages.
Points the Ideogram 4 requirements link at its page and sorts Upscaling after Intermediates.
fix(webv2): darken success and warning toast fills for AA contrast
…indows-memory

Carries the branch's low-VRAM docs and regenerated API contracts to their moved locations under configuration/Optimization and frontend/api.
Carries the branch's regenerated API contracts to their moved location under frontend/api.
docs: document model, quantization and upscale changes from recent PRs
fix(z-image): hold the schedule shift at the end of its fit
Keeps the corrected force_tiled_decode description next to auto_tiled_decode and regenerates the API contracts.
Folds the automatic-tiling docs into the existing Tiled VAE decode and encode section and fixes the moved page's system-requirements link.
…-2.13

Keeps the CUDA or Intel XPU requirement and drops the CUDA 12.x mention from the FP8 Storage page, and regenerates the API contracts.
ROCm on Windows: keep VRAM resident and bound attention and VAE memory
…-2.13

Regenerates the API contracts with both noise_dtype and the ROCm memory settings.
chore(deps): move the cpu and cuda extras to torch 2.13.0 (cu130)
@lstein
lstein merged commit 8f7bd21 into invoke-ai:main Oct 2, 2026
14 checks passed
@lstein
lstein deleted the upstream-merge branch October 2, 2026 00:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

api backend PRs that change backend files CI-CD Continuous integration / Continuous delivery docker docs PRs that change docs frontend PRs that change frontend files frontend-deps PRs that change frontend dependencies invocations PRs that change invocations python PRs that change python files Root services PRs that change app services

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants