Skip to content

firmware: module code in PSRAM, a 41% faster blit, a crash record, and a console that outlives a hung pattern - #466

Merged
engmung merged 18 commits into
mainfrom
dev
Oct 1, 2026
Merged

engmung merged 18 commits into
mainfrom
dev

Conversation

@engmung

@engmung engmung commented Oct 1, 2026

Copy link
Copy Markdown
Owner

The firmware core, after a review asked whether anything in it could be raised a level rather than polished. Five things came out of it, each measured or reproduced on a board before it was committed. Two boards were on the bench: a fresh one on the default image (.217) and one on the Audio edition with six of its owner's modules (.246).

What changes

A module's code runs from PSRAM (88f9e5a, 617f058). "The S3 cannot execute from PSRAM" is true only of the pointer the heap returns: the instruction and data buses share one MMU table, so the same page is fetchable at that address + 0x06000000. The loader writes and relocates through the heap's pointer and calls through the alias; the code block owns whole cache lines (the S3 ROM's write-back erratum), is read back through the instruction bus before the first call, and a unit on which that fails falls back to internal RAM until reboot. PSRAM-first is the default (PF_MODULE_CODE_POLICY). Module files and the ABI are unchanged.

before after
Audio: module with 23 KB / 32 KB of code refused loads
Audio: internal heap with a 10 KB-code module resident 21-23 KB 27.8 KB
.246, Bus Window Rain (8.2 KB of code) refused one load in four, 528 B short loads every time, 29 KB free instead of 22
frame time, catalogue and community modules - -0.1..+1.2 %
frame time, 10-32 KB of code all executed every frame - +2.6..5.3 %

The blit is 41 % faster (a33d53b, ef9cc14, 20c9af2). Three short passes with the 8-bit case written in, the same DMA words (check_blit.py, 70 M compared). presentUs 5,960 -> 3,496 µs on both boards, so every pattern is 2.45 ms a frame shorter (Origin 12.1 -> 9.6 ms); 3.8 ms with white balance or gamma tuned. 1 KB less IRAM than the loop it replaces.

A crashed board says where (ed6863a..5cdb683). The SDK already wrote a core dump on every panic and nothing read it. Status gains a crash object: task, cause, PC, backtrace and build hash from the dump, plus a breadcrumb (pattern, phase, where the module's code was) kept through the reset, so an address inside a module comes out as an offset the .pfm's symbols resolve. Report-only. On the bench a null store in a module's draw() came back as pattern, draw, pc +0x4e, caller +0xb6.

A pattern that never returns no longer takes the console with it (43e4ae5..5551ebd). A request waiting on the render loop gives it up once the loop has not reached its frame boundary for 20 s and answers 503; status, the pages, /update and Reboot keep answering. Reproduced first: on the old image the pattern list got no reply and after it neither did status, until a reset over USB. No automatic reboot.

Seven ways input could take a panel down or leave it wrong (3464e07, 3fa2baf), six reproduced on a board before and gone after:

  • a multipart POST cut off after its first boundary: watchdog reset
  • curl -X POST /api/patterns or /update with no form: panic
  • a multipart boundary of 9,000 characters: stack overflow
  • a quote in a pattern's name: /api/patterns unparseable, the page could list nothing
  • a module's global constructor that calls the host: panic on every pick
  • an abandoned upload leaving storage latched busy, and its fields answering for later requests
  • lanes held off from day 25 by a signed compare in the hands-off timer (found by reading; not waited for)

check_parser.py (e677991) replays 13,953 cut-off, stalled and malformed requests through the real vendored parser and found two of the above on a PC first; check_crash.py walks the crash record through 583 boots. Both run in CI.

Checked

  • firmware/bundles/build.sh all: five compositions, marker scan clean; check_footprint.py re-pinned (IRAM -996 B, static DRAM +152..168 B net).
  • Every host check, locally and under g++ with -Wall -Wextra -Werror and ASan/UBSan.
  • Soaks on the final image: 160 and 120 switches with NVS writes and sleep cycles, no reset, no refusal, no false loop stall; earlier 400 + 200 + 160.

Not checked

  • The picture, by eye. The blit is bit-exact against its reference and nobody has looked at a panel since. The thing to look for is in src/hub75/VENDORED.md ("The blit must stay slower than the scan"): flicker in the lower rows on the SELECT screen. It holds for the shipped 16 MHz clock and not for a slower one.
  • Performance and MIDI images on hardware (they build; the feature routes' 503 path was not exercised).
  • A closed enclosure for hours; power-cycle boot restore (software reboots: 6 of 6).
  • The Performance edition with a dead uplink, which is what should set PF_LOOP_STALL_MS.

Left open

  • The vendored upload wait has no deadline (stock; pinned as KNOWN in check_parser.py).
  • PFLoopSync::run() is [[nodiscard]] but nothing fails a build that ignores it.
  • The shelf's Basics pack is still the 2026-08-15 build.

🤖 Generated with Claude Code

engmung and others added 18 commits October 1, 2026 23:34
A loaded module's .text needed one contiguous block of internal executable
RAM - the RAM the console, lwIP and every feature also live in - on the
premise that the S3 cannot execute from PSRAM. That is true only of the
address the heap hands back. The instruction bus (0x42000000) and the data
bus (0x3C000000) index one MMU table, so a PSRAM page reached at 0x3Dxxxxxx
is fetchable at that address + 0x06000000.

The loader now places code in PSRAM, writes and relocates it through the
heap's pointer, resolves every symbol in an executable section to the alias
(execAddress), and before the first call writes the data cache back,
invalidates the instruction cache for the alias and reads the code back
through it. The block is allocated as whole cache lines so the write-back
never touches a line the allocator or a neighbour is using (the S3 ROM's
write-back erratum); a load that does not read back, or a block with no
alias, demotes the unit to internal RAM until reboot instead of failing the
same way on every retry. PF_MODULE_CODE_POLICY picks the rule - internal
only, PSRAM as fallback, PSRAM first - and PSRAM first is what ships on the
S3. Module files and the ABI are unchanged.

Measured 2026-10-01 on two boards, alternating on one boot:
- catalogue and community modules within -0.1..+1.2% of their frame time
- 10/23/32 KB of code with every byte run every frame: +2.6..5.3%, and the
  self-check sum matched on every load
- Audio edition: 23 and 32 KB of code refused before, load now; internal
  heap with a module resident about 27.8 KB whatever its code size, where
  10 KB of code used to leave 21-23 KB
- a community module with 8 KB of code was refused one load in four for
  want of 528 bytes; it loads every time and leaves 29 KB instead of 22
- 400 + 200 + 160 alternating switches, with sleep cycles and NVS writes,
  no reset, no refusal

/api/status: load.code says where the code is, moduleMemory.codePolicy the
rule in force. Not measured: a closed enclosure over hours, the Performance
and MIDI images on hardware.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Four ways a panel could be taken down or left wrong by what it was given.
The first three were reproduced on a board before the change and are gone
after it; the fourth was found by reading.

- A multipart POST that stopped after its first boundary, to any route,
  left the form parser reading empty lines with nothing to wait for, and
  the Core-0 watchdog reset the board five seconds later. readLine() now
  says whether a line actually ended and both loops give a truncated form
  up. Giving up clears the fields it had collected, which otherwise
  answered for the arguments of later requests, and reports an already
  delivered file as aborted, which otherwise left /api/patterns "storage
  busy" with the panel paused (src/webserver/VENDORED.md, Fix 4).
- /api/patterns and /api/patterns/select wrote names into JSON unescaped,
  so one quote in a title made the pattern list unparseable. Names are
  escaped (PatternflowHttp::appendJsonText), and a sidecar name is decoded
  as the JSON string it is: an escaped quote no longer ends it, \uXXXX
  becomes UTF-8, and a name too long for its slot is cut between
  characters. The hotspot password is escaped in /api/hotspot.
- A module's global constructors ran before the entry point that hands it
  the host API, so a namespace-scope initialiser that allocates or logs
  crashed the board on every pick. Order is now entry, constructors, entry
  again (NAME and the labels may themselves be constructed), setup. The
  _ctor_probe module calls the host from its constructor.
- The lane hands-off timer used a signed compare to survive the 49-day
  wrap and thereby moved it to 24.9 days, after which a lane on a knob not
  touched since was held off for the next 24.9. It is a flag that ends
  once.

check_runtime.py now also exercises the sidecar-name decoder.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…arser

src/webserver/Parsing.cpp runs on pf-net, where a loop that waits for the
peer without sleeping starves IDLE0 and the Core-0 watchdog reboots the
board 5 s later. It has done that three times (VENDORED.md), the last with
one truncated multipart POST to any route, and each time the first thing to
run the code was a board.

check_parser.py compiles the real Parsing.cpp and WebServer.cpp against a
socket that replays a script (bytes, stalls, a FIN or silence) and a clock
that only moves when somebody sleeps. Every replay is held to two things:
the parser never looks at the socket 64 times in a row without reading a
byte or sleeping, and the request is over within 30 s of fake time. Cases:
eight well-formed requests (multipart with fields and files, urlencoded,
JSON, raw PUT, GET, DELETE); the multipart, raw and urlencoded bodies that
stop short, with the peer closed and with the peer silent; a closing
boundary without its CR/LF; stalls and a one-byte-every-3-ms trickle; then
every truncation point of every request ended both ways, and every
two-segment delivery - 13,953 replays, a few seconds under ASan/UBSan.

Against the parser as it is on dev this is red on 15 lines, starting with
"multipart: the first boundary, then the peer closes: SPUN". It is green
against Fix 4, which lands separately - so this check must not reach dev
before that fix does. Seeded back into the fixed file, each earlier
regression turns it red as well: the missing braces in _uploadReadByte,
readStringUntil, the raw body asking for a full buffer, a wait that stops
polling network maintenance.

Three things the parser does today are printed as KNOWN rather than failed,
and pinned so that fixing one turns the check red until the pin is removed:
the upload wait has no deadline (stock, left alone on purpose - it sleeps);
a raw body's query string is never parsed, so the handler sees the previous
request's arguments; and a POST that is not multipart, sent to a route with
an upload callback, reaches that callback as a raw body with upload() a
null reference.

The String and the socket are models. The String was compared, operation by
operation, with the real WString.cpp when it was written; the socket copies
WiFiClient.cpp's receive side and says where it stops vouching.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…column pair

The September kernel was compute-bound, not waiting on the DMA engine: 635
instructions per column pair in the 691 cycles a 5.9 ms frame works out to.
The excess was a run-time depth (25 instructions per plane word where 10 do),
x77/x150/x29 synthesised as shift/add chains by GCC 8.4 at -Os (about 20 a
pixel), and one 400-instruction loop body with more live values than
registers (48 stack reloads and 12 literal loads a column pair).

blitRGB888 now walks 32 columns at a time through pfPostRow (saturation and
LUTs), pfSpreadRow (on-time and plane spread) and pfStoreRow8 (plane words,
depth 8 written in). Same tables, same DMA words, same on-time sum:
check_blit.py passes unchanged under MSVC and under GCC with ASan/UBSan, and
the blit taken out of the linked firmware.elf and run one instruction at a
time matches the scalar reference word for word.

The loop it replaces stays, whole, as pfBlitRowPair for odd widths and depths
other than 8, and moves out of IRAM along with pfBuildSpread: the blit holds
1,184 B of IRAM where it held 2,440. The scratch between the passes is on the
stack (blitRGB888's frame goes 224 -> 592 B), so a blit split across the two
cores would share nothing.

Not measured on a panel. VENDORED.md also records why the blit has to stay
slower than the scan - 96 us a row pair - or wait for the flip; at 426
instructions it cannot go under 113.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
config.h ships gamma 1.0 and white balance 1/1/1, so the three LUTs pass 1a
reads hand every index back unchanged - at two instructions apiece, and with
three table pointers the loop has no registers for. pfPostRowRaw is the same
pass without them: 32 instructions a pixel instead of 41, 390 per column pair
instead of 426 (635 before the passes).

blitRGB888 asks the tables on every frame instead of remembering, because
they are the caller's and it rebuilds them in place when gamma or white
balance is tuned: 768 compares, against 36 instructions on each of 2,048
column pairs. A tuned panel takes the LUT path and costs what it cost one
commit ago.

Same DMA words either way: check_blit.py's trials use identity and random
LUTs, so it runs both forms, and both were run out of the linked image
against the scalar reference. 248 B more IRAM than the previous commit,
still 996 B less than dev.

The floor recorded in VENDORED.md moves with it: this blit cannot go under
104 us a row pair, and the scan takes 96.

Not measured on a panel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dary, no longer reboot the panel

Both were one request and a panic reset, found by check_parser.py on a PC
and then reproduced on a board.

- `curl -X POST /api/patterns` (or /update) with no form in it went down the
  vendored server's raw-body path, because canRaw() is true for any route
  with a body callback, and called the route's multipart upload callback -
  whose first line reads server().upload() through a null pointer. The raw
  path now excludes multipart routes; the request reaches the completion
  handler and is answered (400).
- _parseForm() sized a stack array from the boundary named in Content-Type.
  pf-net has an 8 KB stack. A boundary longer than the 70 characters one
  may have is refused, and the array is fixed.
- The raw path never parsed the query string, so PUT /update?size=N found
  no size and arg() answered with the previous request's. It is parsed now.
- /update's completion handler answered a request that carried no image
  from whatever the last upload left behind. It says so instead.

src/webserver/VENDORED.md, Fix 5. The host check's two pins for these
become requirements, a case for the oversized boundary is added, and the
script no longer rewrites a line of the parser for MSVC. One pin remains:
the upload wait has no deadline (stock).

Verified on a board: the three requests are answered, uptime continues,
curl -F and raw PUT uploads still install.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
heap_caps_aligned_alloc() was called from nowhere else in the image, and
the allocator's aligned path is IRAM: asking for it cost 1.5 KB of the
internal RAM this placement exists to give back (.iram0.text 71,439 before
the PSRAM work, 72,975 with it, measured from the linked sections).

The code block is now an ordinary PSRAM allocation with one cache line of
slack, and the code starts at the first line boundary inside it - every
line it occupies still lies wholly within the allocation, which is what
the write-back needs. LoadedSection keeps the heap's pointer to free.

With the blit kernel's smaller IRAM footprint .iram0.text is 70,443: about
1 KB under where this work started.

Verified on a board: self-check sums match, the constructor probe reaches
the host, 150 switches with NVS writes and no reset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…la, and the loop task's stack is in status

Measured on two boards, A-B-B-A between images: presentUs 5,960 -> 3,496 us
(3,802 with white balance tuned off the identity), the same on the default
and Audio images, and every pattern's frame shorter by the same 2.45 ms.
VENDORED.md had 'not measured'; it has the numbers now.

From the review of the kernel:
- The line the blit must stay slower than is the scan's row time - buffers
  per row x words per row / pixel clock - not the 96 us that holds for the
  shipped 16 MHz. At the slower clock core_display.h describes for EMC the
  scan is the slower of the two and the blit would write rows still on the
  panel. Said as a formula in VENDORED.md and where i2sspeed is set.
- blit_test.cpp only tried all-identity and all-random LUTs, and only
  widths that are multiples of the 32-column chunk. It now tries one
  channel off, a single entry off at 0, 128 and 255, and widths 2, 34, 98
  and 130 - a decision that asked one table, or a broken short chunk,
  passed before.
- /api/status gains loopStackMin. The blit takes about 640 B of the loop
  task's stack inside whatever a pattern's draw() holds, and nothing
  reported that task's high-water mark. Read on both boards across every
  module installed: 4,868 B free at the lowest.
- Footprint re-pinned: IRAM -996 B in every edition (the generic loop left
  IRAM), static DRAM -136..-152.

Not checked: the picture on a panel, by eye.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pattern breadcrumb in status

resetReason said "panic" and nothing else. The SDK has been set up to write
an ELF core dump to the coredump partition on every panic all along
(CONFIG_ESP_COREDUMP_ENABLE_TO_FLASH, partitions/app3M_fat9M_16MB.csv) and
nothing in the firmware read it. This is the reporting half of RESIL-1 /
NET-4 / MODULE-2 from the 2026-10-01 core review; it changes nothing the
device does - no supervisor, no reboot, the boot latch untouched.

src/core_crash.h, new:

- The dump. PFCrash::begin(), first thing in setup(), reads the summary once:
  task, EXCCAUSE, faulting address, PC, up to 16 backtrace frames and 16 hex
  digits of the crashed image's ELF hash (the same hash `build` publishes).
  esp_core_dump_get_summary() in IDF 4.4.7 maps as many bytes as the first
  word says and reads an ELF header 20 bytes in without checking either, so
  the size is bounded and the ELF magic looked for at that offset first, and
  the image is CRC-checked before it is parsed. The parse itself leaves a
  mark in RTC memory: a boot that finds the mark does not parse again, so a
  dump that kills the parser cannot keep a board from starting.

- The breadcrumb. 64 bytes of RTC no-init memory, magic + FNV checksum: the
  pattern's slug, the phase (loading, constructors, setup, update, draw,
  idle) and the address and size of the module's executable section. The
  phase is one word outside the checksum, stored before and after each call
  into pattern code - two stores per call, in the loader for a module and in
  pattern_registry.h for a preset. On a boot that follows a panic or a
  watchdog it is copied into the record.

- The two are not always one event. A dump outlives the breadcrumb (flash vs
  RTC memory), and a death does not always write one: the RTC watchdog never
  reaches the panic handler, and a dump too big for 64 KB is refused before
  the old one is erased. The breadcrumb keeps the checksum word of the dump
  the previous boot saw; only a dump that differs after a panic is
  `fromThisReset`, and only then is an address inside the module's code
  published as an offset ("+0x4e") instead of a heap address no ELF decodes.

/api/status gains `crash` (absent when there is nothing to report), text
fields through appendText. DELETE /api/crash erases the dump and drops the
record; 404 when there was none. The same record is printed at boot as
[CRASH] lines. docs/rest-api.md 1.5, "The crash record".

Cost: 8 bytes of internal RAM (two pointers, by nm), 64 bytes of RTC memory,
the record (~180 B) in PSRAM only when there is one.

Known side effect, written at the variable: .rtc_noinit also holds the SDK's
RTC clock bookkeeping, and this struct links ahead of it, so those two words
move 64 bytes. The first boot into this image after a warm reset (an update
over the air) reads them from memory nothing wrote, and the wall clock is
wrong on that boot until SNTP sets it. Derived from the SDK's code, not seen
on a panel.

tests/modules/_crash_probe is the bench case: a module that writes through a
null pointer from draw() when a knob moves.

Not run on hardware. Default and audio compile; check_boundaries, check_runtime,
check_module_elf, check_abi_freeze, check_sources pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The previous commit said eight, from nm on this file's own two pointers. A
baseline build of dev says otherwise: the default image's heap start moves
from 0x3fcaab78 to 0x3fcaac60, and 224 of those 232 bytes are four strings
in .dram0.data (__c$7214, 7216, 7218, 7221 - 43, 43, 91 and 47 bytes). They
are the error messages of esp_core_dump_get_summary(): the SDK's core dump
code logs through DRAM_STR so it can print with the flash cache off, and
that applies to the one function in it that only runs at boot.

core_crash.h now states the measured number and where it comes from. The
rest of the measured delta on the default build, same toolchain, dev at
7884ee1 against this branch: firmware.bin +4,768 B (1,091,968 -> 1,096,736),
IRAM unchanged (.iram0.text 71,439 both), .rtc_noinit 16 -> 80.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed from, and lives in DRAM

Three things a review of the crash record found, before it reached a panel.

The breadcrumb recorded sections[i].memory as the module's code range. dev
now puts a module's code in PSRAM and calls it through the instruction-bus
alias, so every PC inside a module is 0x43xxxxxx while `memory` is the
0x3Dxxxxxx the loader wrote it through: no frame of any crash would have
fallen inside the range, and the "+0x.." offset the range exists to produce
would never have been printed. It records `exec` now, which equals `memory`
for code in internal RAM, so one line covers both placements.

The breadcrumb was in RTC no-init memory. .rtc_noinit is not empty: the SDK
keeps the wall clock's bookkeeping there (s_rtc_last_ticks,
s_esp_rtc_time_us), initialised only on a power-on, and a 64-byte struct
linked ahead of them moves them. An update over the air is a warm reset, so
the first boot into the image with the struct - and the first boot back into
any older one - reads those two words from memory nothing wrote, and the libc
clock is garbage until SNTP sets it. It is `__NOINIT_ATTR` now: DRAM .noinit
is empty in every image, survives every reset the record acts on, and the
magic and FNV seal already reject a stale or shifted block. By nm on the
default build: PFCrash::trail at 0x3fc9b0f8 in .noinit (64 B), .rtc_noinit
back to the SDK's 16.

A preset's setup() at boot was not marked, so a preset that died there came
back as "idle" with no pattern named. The sketch's boot loop now names each
preset and puts the same pair of phase stores around its setup() that the
loader puts around a module's.

docs/rest-api.md: the code.base example is an instruction-bus address; cause
29 with vaddr 0 and panic_abort on top is abort(), an assert, the stack
check or the task watchdog; causes from 64 up are the SDK's own; after a
task_wdt or int_wdt the backtrace is the interrupt's stack; "setup" with
loopTask is the boot restore or a preset; a dump from before this firmware
is reported with fromThisReset false until DELETE /api/crash. The 1.5
history line gets its phrase. The crash object's text goes through
PatternflowHttp::appendJsonText.

Cost on the default build against a build of dev at 20c9af2, with the host
check's commit on top: internal RAM 296 B (heap start 0x3fcaa7d8 ->
0x3fcaa900: .noinit +64, .dram0.data +220, .dram0.bss +16), IRAM unchanged,
firmware.bin +4,976 B (1,095,424 -> 1,100,400). The audio edition moves by
the same 296 B.

Not run on hardware. Default and audio compile.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eadcrumb and dump

What a boot makes of the reset before it cannot be exercised on a panel on
demand - an RTC watchdog, a dump too large to write, a dump torn by a power
cut - and a mistake in it crashes nothing: it reports last week's backtrace
as today's, or pins a death on a pattern that was not running.

check_crash.py compiles src/core_crash.h unchanged against a scripted
esp_reset_reason, coredump partition and dump parser (tests/crash_test.cpp)
and walks 583 boots: every reset reason x breadcrumb intact, noise, broken
seal, wrong magic, phase out of range x no dump, the old one, a new one, a
first one, a torn one. Also a power-on with an old dump, a panic whose dump
was refused, the guard that keeps a dump which kills the parser from being
parsed again, sizes and headers the parser must never be handed (the
partition stub asserts no read leaves it), slug extraction, DELETE, and
inModule() at base, base+size-1, base+size, with code in internal RAM and
at the 0x43xxxxxx a PC carries for code in PSRAM. No seam was needed in the
header: the hardware includes are empty files and the test supplies them.

Sixteen hand-made mutations of core_crash.h each fail it. Passes with MSVC
and with g++ -Wall -Wextra -Werror under ASan/UBSan; CI runs it.

CHANGELOG: the crash record's entry, which says that it only reports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A null store in a module's draw(), running from PSRAM: status reported the
pattern, the phase and module-relative addresses that match the module's
symbols; the dump was read in 62 ms; DELETE /api/crash cleared it and a
second one answered 404; a dump left by an earlier image showed as
fromThisReset false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The loop task is on no watchdog and a module's draw() runs on it, so a
pattern that never returns holds the panel on its last frame for good. The
first console request that needed the loop then waited in
PFLoopSync::runRaw without limit - and the /patterns page opens with one -
while the server answers one connection at a time, so /api/status, /update
and Reboot, which need nothing from the loop, went silent behind it. A
frozen panel could be neither diagnosed nor restarted without the plug.
Found by reading (RESIL-BUG-1 in the 2026-10 core review); no panel has
been made to do it yet.

- The wait ends on evidence. A caller gives up only when the loop has not
  reached service() for PF_LOOP_STALL_MS (20 s), measured from the loop's
  own stamp and not from how long the caller has waited: a storage
  transaction waiting out a module's setup() through runWhen() is tested
  every frame and is not a hang. Twenty is above the 14 s a feature's
  fetch holds the loop when the name lookup times out. The boot restore,
  the longest stall there is, runs before the server is serviced at all.
- Giving up is taking the request back. Caller and service() both move
  pendingFn with one atomic exchange, so a posted request goes to exactly
  one of them and the loop cannot run a body on a stack frame that has
  returned. A request the loop has taken is waited out, however old the
  stamp.
- run() returns whether the body ran, and is [[nodiscard]]. Every caller
  answers a refusal instead of using what its body never set: the pattern
  list, select, upload, delete, format and the knob settings reply 503
  "render loop is not answering" and change nothing; the clock, mqtt, show
  and weather handlers do the same. The thumbnail cache belongs to the
  loop, so a delete that cannot reach it removes the file and leaves the
  loop a note to empty the cache when it next runs, and the disk worker
  writes nothing until it has.
- /api/status carries loopAgeMs, loopStalled and loopSyncGaveUp, read off
  the stamp without the loop's help.

Report and survive, nothing more: no watchdog, no supervisor, and nothing
restarts the panel by itself.

runtime_test.cpp covers the withdrawal, a transaction that outlasts the
limit while the loop keeps arriving, a taken call waited out, the age never
being read from the future, and the hand-off race itself - caller and loop
let at the request in the same instant, the aim corrected by whoever won
the last one, both outcomes required. Seeded back in, a load-then-store on
either side fails it on every run (g++ -O2, with and without sanitizers,
two cores and twenty). thumbs_test.cpp covers the deferred drop.

toolchain/tests/modules/_hang_probe is a module whose draw() stops
returning eight seconds after it is picked, for the bench. With K1 or K2
turned during the countdown the stop ends after thirty or twenty seconds
and repeats: thirty shows a loop waking to requests that gave up on it,
twenty puts the wake within a slice of the moment a waiting caller gives
up. It builds; it has not been on a board.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With the render loop stopped, /status went on reading "awake" over the
frame rate of the last frame before the stop: frameUs is only ever written
by a frame, and nothing clears it. That page is the screenshot people send
when they ask for help, and it described a healthy panel.

It now reads loopStalled: the Panel row says "frozen for <time>" in red,
ahead of asleep and paused, the frame rate and frame time go blank the way
they do for a dark panel, and the note under them says what still works -
Reboot on the Wi-Fi page - and that the panel comes back on the same
pattern. A status without the field (older firmware behind a cached page)
reads as before.

Checked by running the page's script against stub statuses; not yet seen
on a board.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…can and cannot say

From the review of the bounded loop sync, and from running it on a board.

- Twenty seconds is clear of one feature lookup with the internet down and
  not of two in the same iteration (weather and an MQTT broker named by
  host: 28-30 s), nor of a TLS handshake. The comment and the REST doc said
  otherwise. The limit stays; the words change: the Status page says 'not
  answering for', and says what else it can be.
- Two limits written down: the console's routes are registered by the loop
  on the first link, so a loop that hangs before then has no console to
  outlive it; and routes that only queue work for the loop still answer ok
  on such a panel. /update needs nothing from the loop only while it is
  armed without the UPDATE screen.
- The probe's 20 s mode checks both orders on a board, not the tie; the
  race case in runtime_test gets about 200 rounds under MSVC and 4,000
  under g++, so CI is its gate. Both said so wrongly or not at all.
- A local in the race test shadowed one added on dev.

Reproduced on a board with _hang_probe: on dev the pattern list got no
reply and after it neither did status, until a reset over USB. On this
branch the list answered 503 in 18 s, sixty more were refused at once with
the heap unchanged, /status, /update and / still loaded, and Reboot brought
the panel back; a 30 s hold ended with the list answering 200 again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…op sync

Static DRAM +304 B in every edition: 296 for the crash record (224 are the
SDK's own strings behind esp_core_dump_get_summary, 64 the breadcrumb in
.noinit) and 8 for the loop-sync stamp and counter. IRAM unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 1, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
pattern-flow_origin Building Building Preview Oct 1, 2026 10:47pm UTC

@engmung
engmung merged commit 1485831 into main Oct 1, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant