[pull] main from daijro:main - #20
Merged
Merged
Conversation
…GL, WebRTC, media, timing, launcher (#779) * feat(humanize): replay recorded human mouse movements (Cursory) humanize=True used to walk a Bezier curve through two random knots and emit a point every 10 ms. Both halves are tells: an analytic curve sampled at a fixed rate has velocity and jerk profiles that separate cleanly from a hand's, and every movement accelerated through the same easing function. Juggler now picks one of Cursory's 2357 recorded human movements whose direction, distance and wander suit the move, morphs it onto the requested endpoints and replays it with the recording's own timing. The generator is cursory-js (a bit-exact TypeScript port of Vinyzu/cursory) vendored under additions/juggler/input/cursory/; it is LGPLv3-or-later, not MPL-2.0, and ships its LICENSE and NOTICE inside juggler.jar. MouseTrajectories.hpp and ChromeUtils.camouGetMouseTrajectory are removed. sendTrajectoryAcked takes per-step pauses, drops points on the pixel the last dispatch left the cursor on (a zero-displacement move is never acked), and the humanize guards are updated for the new path shape. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(patches): shared helpers for binary resolution, a private Xvfb and Marionette resolve_binary() honours the runner's CAMOUFOX_EXECUTABLE_PATH before falling back to a Linux objdir (search-service-init and touchscreen-digitizer ignored it and ran the newest objdir, which after a macOS cross build is an arm64 Mach-O), hidden_display() gives a guard its own Xvfb so nothing ever opens on the user's display, and a minimal chrome-context Marionette client lets guards inspect browser UI state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(juggler): synthesized input carries what real mouse and keyboard input carries - pointerType was "" for every Playwright mouse event: juggler dispatched with MOZ_SOURCE_UNKNOWN. It now passes MOZ_SOURCE_MOUSE (#776). - keyboard.type() never pressed Shift: a shifted character now arrives bracketed by ShiftLeft keydown/keyup (location 1) with shiftKey set. - After the pointer was parked off content, pointerover/enter re-entered with buttons=1 and pressure 0.5; the tracked position is now forgotten on park. - A Windows identity gets contextmenu after mouseup with buttons=0, as Windows does; GTK/macOS keep it on press. - Wheel events are sent as line deltas (DOMMouseScroll.detail 3 per notch instead of the pixel count). - The browser rect is measured after the APZ flush await, so a chrome height change during the wait cannot put a y==0 dispatch one row above content. - ci/run_sundial.py moves and clicks the mouse so input vectors have data. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(juggler): evaluate() no longer grants user activation Upstream Playwright runs every evaluate() as handling user input and notifies a user-gesture activation. Init scripts go through that path at load, so every page started with navigator.userActivation.hasBeenActive === true, autoplay allowed and popups permitted before any input. Activation now only comes from juggler's trusted input events, as in a stock browser. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(juggler): stop hiding scrollbars in headless The headless agent sheet set scrollbar-width: none !important, which a page reads back from getComputedStyle and from overflow:scroll gutters. Scrollbar appearance is left to the platform look-and-feel (the launcher sets ui.useOverlayScrollbars per claimed OS). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ui): no visible automation cues in the browser window - Every Playwright context was a public container, so the URL bar showed a "JUGGLER <id>" label and a container colour. Contexts are now non-public identities (tabbrowser renders public identities only); startup cleanup still removes persisted leftovers. - showcursor defaulted to true, drawing a red dot that followed the mouse. It is now opt-in. tests/patches/visible-automation-cues.py checks both on a private Xvfb. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fonts): web fonts, local(), per-character fallback and native bundles - FontFace / @font-face were answered from the font allowlist by the FontFace's own family name, so every url() web font failed with NS_ERROR_FAILURE and never rendered, local() of an allowed font failed, and a miss rejected with an XPCOM code instead of NetworkError (#759). Stock FontFace/FontFaceImpl are restored; local() is filtered by the RESOLVED family in gfxUserFontSet. - GlobalFontFallback forced the cmap scan, which skips families whose charmap is not loaded yet, so any character outside Gecko's script-based common-fallback table rendered as the primary family's .notdef (U+1E9E on macOS). The platform fallback chooses again, and its choice is held to the mask. - For a native macOS/Windows identity the bundled font sets are not activated: a bundled face of a family the system also has (Papyrus, Helvetica) won the lookup with different metrics. On Windows the enumerator still keeps Twemoji Mozilla, the emoji font stock Firefox ships (flag emoji drew nothing). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fonts): CSS2 system fonts and system-ui follow the claimed OS - The host's own OS gets no system-ui override (macOS resolved system-ui to Helvetica instead of -apple-system). - CSS2 system font keywords use per-keyword faces and sizes; a Linux identity reports the Ubuntu desktop font; Windows form controls (-moz-button/field/list) answer "MS Shell Dlg 2" as Windows does, not Segoe UI. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(fonts): per-OS font model and fontconfig parity fonts.json is now generated from the bundle by scripts/gen-fonts-json.py (fc-scan + aliases + scan-time families, intersected with the per-OS manifest in scripts/data/font-manifests.json) so every reportable family is renderable; scripts/verify-fonts.py checks that invariant, the generics and the reject globs. font-groups.json lets the draw keep co-shipped families together. Linux fontconfig: stock metric aliases (Arial -> Liberation Sans, ...), 49-sansserif, urw-base35 and the non-Latin rule files in stock conf.d order, generics resolving like a stock Ubuntu (Noto Sans / Noto Serif / DejaVu Sans Mono / Z003), hintslight so advances are not pinned to whole pixels, and weak <prefer> lists instead of strongly-bound generic pins so lang can promote a script face. Windows fontconfig: GDI substitution aliases, MS Shell Dlg 2, cursive/fantasy generics, duplicate-face rejects and Sitka / Segoe UI Variable optical-size families. NOTE: generated against a ~3.9 GB target font bundle that is not part of this change (one file is over GitHub's 100 MB limit); see docs/FONTS.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(locale): localize browser strings with the spoofed locale; stop rewriting explicit locales - With locale="fr-FR", Intl/number/date went French but input.validationMessage and XML parse errors stayed English, a mix no real Firefox produces. Official language packs are now baked in as packaged locales (scripts/fetch-langpacks.py, scripts/inject-locales.py, called by package.py, fetched on demand) and the launcher selects the UI locale through intl.locale.requested. A langpack add-on cannot do this: the parent pre-creates those string bundles first. - locale-spoofing.patch overrode Language/Script/Region on every intl::Locale, so new Intl.DisplayNames(['en'],{type:'region'}).of('DE') returned the spoofed region's name and Intl.Locale('ja-Jpan-JP').minimize() returned the spoofed tag. Only the OS/default locale is spoofed now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(media): enumerate, capture and label the identity's media devices coherently The fake media engine now enumerates the identity's microphones, cameras and speakers (labels and group ids from new mediaDevices:*Labels/*Groups config keys), MediaManager uses it whenever mediaDevices:enabled, and stock exposure rules apply: before a grant one device per input kind, no outputs, no labels; after a grant OS-style labels, distinct deviceIds, shared groupIds. So enumerateDevices(), getUserMedia() tracks and getSettings() ids agree, and a claimed camera captures instead of throwing NotFoundError. Fixes the content-process crash on an identity with a camera and no microphone (InsertElementAt on an empty array). docs/MEDIA-DEVICES.md; guard tests/patches/media-devices-coherence.py. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(navigator): globalPrivacyControl agrees between window and workers (#760) The main-thread Navigator getter ignored the config key that WorkerNavigator::GlobalPrivacyControl honours, so a page read false in the window and true in a worker. Both read the key the same way now; the launcher also mirrors it into privacy.globalprivacycontrol.enabled so the Sec-GPC header agrees. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(timezone): apply the launch-level timezone from the first read (#773) The timezone config key was only applied lazily from a navigator getter, so Intl and Date reported the host zone until a page happened to touch navigator. It is now applied eagerly in every process (nsJSContext::EnsureStatics) and per realm when a new inner window is created, entering that window's realm rather than whichever one triggered the navigation. window.setTimezone() still takes precedence per context. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(screen): the CSS color media feature follows the spoofed colorDepth screen.colorDepth was spoofed at the WebIDL level only, so on a 10-bit panel a 24-bit identity reported 24 with (color: 10), a pair Gecko cannot produce. Gecko_MediaFeatures_GetColorDepth now resolves the depth in the same order as nsScreen::PixelDepth. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(webgl): pass live state through instead of answering it from the table getParameter answered everything from the sampled table, so state a page had just set read back wrong (lineWidth(5) read 1, VIEWPORT/SCISSOR_BOX stayed 300x150 on a 64x64 canvas), extension parameters were null (anisotropy, draw buffers), COMPRESSED_TEXTURE_FORMATS was null instead of [], and getContextAttributes() ignored the attributes requested ({antialias:false} still reported 4 samples). Identity and limits still come from the table; live state and context attributes are real. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(webrtc): ICE gathering completes behind a proxy (#774) With Playwright's per-context proxy and media.peerconnection.ice.proxy_only_if_behind_proxy, ICE failed before gathering started and iceGatheringState stayed "new" forever, where stock Firefox completes with host candidates. When that happens around the fabricated candidates the new -> gathering -> complete state walk is replayed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(patches): refresh window-setter-seal.patch offsets Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(windows): embed Firefox's application manifest in camoufox.exe config/rules.mk embeds <program>.manifest and browser/app only ships firefox.exe.manifest, so --with-app-name=camoufox produced an exe with no manifest. Without the Windows 10 supportedOS GUID the process and its children run as a pre-Windows-10 application and Gecko's Windows-10-gated paths switch off (MediaCapabilities.decodingInfo powerEfficient false for H.264/VP9 where stock is true). The new patch adds a byte-for-byte copy as camoufox.exe.manifest; the old rename hunk in windows-theming-bug-modified.patch is dropped. Guard: tests/patches/windows-exe-manifest.py. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(settings): stock values for page-observable prefs; launcher prefs at startup Page-observable defaults that no stock Firefox has, restored: - the forced built-in dark theme (it also removed the 1 px nav-bar separator, and was applied ~1 s after startup, resizing the viewport) and ui.systemUsesDarkTheme (prefers-color-scheme disagreed with the desktop); - focus rings off, autoplay allowed, popup blocker off; - gfx.color_management.mode=0 (Playwright's test pref: ICC-tagged images were drawn unconverted, readable from a canvas pixel); - ui.use_standins_for_native_colors (non-native system colours); - GMP updates off (Widevine/OpenH264 never available); - storage.estimate() quota derived from the raw disk instead of the stock cap. The HardwareAcceleration:false enterprise policy is removed: it locked software WebRender with no hardware video decoding on every OS (guard tests/patches/hardware-acceleration-policy.py). The minimal-theme chrome.css is emptied: its ~55 px chrome made outerHeight - innerHeight impossible. Playwright's non-persistent launch writes no user.js, so launcher prefs only arrived through juggler after startup and anything Gecko reads while starting raced (on Windows the UI locale lost 3 of 4 launches). camoufox.cfg now applies the launcher's CAMOU_PREFS_1..N env chunks as default prefs at startup (guard tests/patches/startup-prefs.py). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(branding): chrome://branding assets match stock Firefox chrome://branding/content/ is content-accessible. The wordmark SVGs had different intrinsic sizes (336x48 / 172x48 vs 300x67) and document.ico, document_pdf.svg and the private-browsing about logos were missing, all measurable from a page with an <img>. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pythonlib): identity draws that match real machines and stay stable Launcher-side fixes found by comparing camoufox against stock Firefox 152.0.4 on Linux, Windows 11 and macOS hosts: - DNT / GPC: BrowserForge draws doNotTrack "1" on most Firefox samples, but a stock Firefox 152 reports "unspecified" and globalPrivacyControl false; the stock defaults are used unless the caller sets them, and both are applied as prefs so the API, the worker and the DNT / Sec-GPC headers agree (#760). - Timezone and geolocation: the timezone is passed to the browser, and a configured position sets permissions.default.geo so permissions.query agrees with the auto-grant (#769, #773). - hardwareConcurrency: the reported count is the fingerprint's and the browser is pinned to that many cores (cpu_affinity.py, Linux/Windows), so worker timing agrees with it; otherwise the host count snapped into the core counts real machines ship with (never 2, Firefox's resistFingerprinting value). - Fonts: the OS base is always present in full, OS-version variants are drawn all-or-nothing, co-shipped groups stay together, Cascadia is never claimed off Windows, a native macOS/Windows identity claims only the real OS base, and gfx.font_rendering.fallback.async is off on Linux so per-character fallback does not depend on cmap-load timing. - Speech voices: a per-OS installed-voice model (voice-manifests.json) with the voiceURI formats each backend really produces (voice-uris.json); no default voice where stock has none. - WebGL: extensions a release Firefox never exposes are filtered, but OVR_multiview2 stays for Windows D3D11 renderers, which expose it. - Media devices: a seeded draw of common per-OS devices with OS-style labels. - Windows scrollbars follow the drawn Windows version (overlay on 11). - Glyph-advance perturbation (fonts:spacing_seed) defaults to off: it moved every measureText width off the value the same font gives on a real machine. - Launcher prefs are also exported as CAMOU_PREFS_1..N so camoufox.cfg applies them at startup, and the browser UI locale follows the spoofed locale. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pythonlib): per-identity salt for seeded draws; core pinning under concurrency Found in review of the previous commits: - identity_seed() hashed only the UA, platform, screen size and core count. Those take a handful of values per OS, so over 500 launches the seed took 12-30 distinct values and every install drew its fonts, voices, GPU, media devices and canvas/audio noise seeds from that same short list. The seed now mixes in identity_salt(): derived from what the caller pinned the identity with (a Fingerprint, a preset dict, a config naming the UA) so relaunching that identity reproduces every draw, and random otherwise. A pinned preset now reproduces its noise seeds too; seeds the caller sets are kept. - Concurrent AsyncNewBrowser launches on one driver interleaved pin/restore: one browser inherited the other's mask and the driver could stay pinned. pin -> launch -> restore is serialized per driver. - Every pinned browser landed on cores 0..N-1; pins now take N adjacent cores from a random start. - A pinnable host with 1-3 cores reported 1, 2 or 3 (2 is the resistFingerprinting value); the table floor of 4 applies as on other hosts. - launch_options() callers that launch the browser themselves (launch_server, direct use) kept the drawn core count although nothing pins the browser; only Camoufox/AsyncCamoufox pass pin_cpu_cores=True now, everyone else reports the host's snapped count. - PLAUSIBLE_CORE_COUNTS gains 18, 22, 28 and 32, all recorded in the -v150 corpus. - The Windows voice list was drawn before the locale was resolved, so an fr-FR identity got en-US voices; it is drawn after locale/geoip now. - macOS "Alex" gets its com.apple.speech.synthesis.voice identifier. - CAMOU_PREFS env chunks are ASCII-only JSON (Windows getenv goes through the ANSI code page). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fonts): canvas accepts CSS2 system-font keywords; local() works on macOS - GetSpoofedSystemFontForRFP's per-OS branches returned before marking the result a system font. ComputeSystemFont copies that flag into FontFamilyList::is_system_font, and without it the canvas font setter could not serialize the value: ctx.font = 'caption' (or icon, menu, message-box, small-caption, status-bar) was silently ignored and read back '10px sans-serif' where stock reads back the keyword. - CoreTextFontList::LookupLocalFont builds a CTFontEntry with no family name, and local() sources are held to the spoofed font list by the resolved family, so on macOS every local() face (Helvetica, Menlo, Arial...) failed with NetworkError, installed and allowed or not. The entry now carries the family CoreText resolved. A blocked lookup's entry is released instead of leaked. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(webgl): only device limits come from the spoofed table getParameter still answered ~100 state pnames from the table, so state the page had just changed read back wrong: UNPACK_FLIP_Y_WEBGL / PREMULTIPLY_ALPHA / COLORSPACE_CONVERSION after pixelStorei, FRAGMENT_SHADER_DERIVATIVE_HINT after hint(), DRAW_BUFFERi after drawBuffers(), RED/ALPHA/DEPTH/STENCIL_BITS and IMPLEMENTATION_COLOR_READ_* for the bound framebuffer, and COMPRESSED_TEXTURE_ FORMATS after enabling an extension. UNMASKED_VENDOR/RENDERER_WEBGL came back without the extension enabled, where stock returns null with INVALID_ENUM. The table now answers only the MAX_*/ALIASED_*/SUBPIXEL_BITS limits, WebGL 2 limits on WebGL 2 contexts only, and extension limits (anisotropy, draw buffers, OVR multiview) only once that extension is enabled; everything else is the real context's answer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(webrtc): fabricate candidates only where a real gather would have them With webrtc:ipv4/ipv6 set (every geoip launch), new RTCPeerConnection() with no iceServers produced a srflx carrying the spoofed IP, and after end-of-candidates a second host set with a different mDNS name. getStats() exposed that srflx as id 'camou-srflx' and rewrote every candidate address, including .local host names and the remote peer's candidates. A srflx is now fabricated only when the page configured an ICE server, host candidates only when none reached the page (sharing the real UDP host's port otherwise), the synthetic stats id has the shape real candidate ids have (8 hex digits, fixed per connection), and only this side's non-mDNS addresses are rewritten. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(media): honour mediaDevices:enabled=false; page fake:true behaves as stock - MaskConfig::GetBool returns std::optional<bool>, and the checks tested its presence: "mediaDevices:enabled": false still enabled the fake devices. - media.navigator.permission.fake=true was page-readable: a page's own getUserMedia({video: true, fake: true}) prompted and never resolved, where stock resolves at once with its generic fake device. The pref is off again; the identity's devices count as real hardware in the capturing checks instead (prompt, sharing indicator, post-grant labels), and a page's fake:true request gets stock's generic devices rather than the identity's. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(storage): per-context values are session state, and misses are cached - Values lived on the user pref branch, which a persistent profile writes to prefs.js: relaunching with a different timezone (or navigator values) kept reporting the previous session's in the page, iframes and workers. They now live on the default branch, which is never saved, and reads ignore user values an older build left behind. - A read of an unset key did a synchronous IPC to the parent every time, and in a launch without per-context values every read is unset: navigator.hardwareConcurrency, screen.* and (color) media queries measured ~20x slower than stock. A miss is now cached per key until a pref change or a local put clears it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(timezone): no per-realm override for the process-wide zone; cache DateTimeInfo With a launch-level timezone every new document and worker got a per-realm override of the zone the process already reported. Setting one releases all JIT code in the runtime (hot code after adding an iframe ran ~4x slower), and the realm rebuilt its DateTimeInfo on every call (getHours() ~40x slower than stock). The override is applied only when the zone differs from the one JS::SetTimeZoneOverride applied process-wide, and a realm keeps its DateTimeInfo until its override changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(juggler): wheel scrolls in native notches; Shift leads the key it modifies - A wheel notch now reaches the page as its own 3-line event carrying one native tick (new WHEEL_EVENT_NATIVE_NOTCHES option in patches/wheel-native-ticks.patch), so wheelDelta is -120 per notch as with a physical wheel; it was -396, and a multi-notch scroll arrived as one event. Several notches are spaced a few tens of ms apart. - Auto-Shift pressed Shift 0 ms before the character's keydown and released it 0 ms after its keyup; it now leads and trails by a drawn human-scale delay, and a failing keydown no longer leaves Shift latched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(settings): stock cookie partitioning, preconnect, popup notification; about dialog CSS - network.cookie.cookieBehavior 4 (Playwright's) -> Firefox's default 5. With 4 a cross-site iframe (captcha and anti-bot widgets are exactly that) sees document.hasStorageAccess() true and its first-party cookies and localStorage, where stock partitions them. Playwright set 4 so storageState need not carry thirdPartyCookie^ permissions. - network.http.speculative-parallel-limit 0 turned <link rel=preconnect> into a no-op, visible in Resource Timing. - privacy.popups.showBrowserMessage false: stock shows a notification bar for a blocked popup, which shrinks the viewport and fires resize. - chrome://branding/content/aboutDialog.css is page-loadable and was empty; it is the official branding's now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(patches): stock-parity probes for the leaks found in review One launch (plus two persistent relaunches) checks page-observable invariants stock Firefox 152 holds: canvas CSS2 system-font keywords, WebGL state readback and UNMASKED_RENDERER without the extension, no srflx without ICE servers and no 'camou' stats id, getUserMedia({fake: true}), cross-site storage partitioning, wheel notches, per-read cost of (color)/hardwareConcurrency and of local Date getters under a launch timezone, and a persistent profile's timezone after relaunch (page and worker). Run against the build before these fixes it fails on every one of them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(humanize): judge cadence by distinct values and spread, not a share of the gap count The page clock is clamped to 1 ms and Cursory's recorded steps mostly sit between 12 and 20 ms, so the number of distinct gap values cannot grow with the number of gaps. Requiring len(gaps) // 4 made the guard fail on visibly uneven runs whenever event delivery was steady (3 of 4 runs once the per-read sync IPC jitter was gone). A fixed 10 ms cadence yields about three values within a few ms of each other, which the new rule (>= 6 values, >= 20 ms spread) still fails. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(juggler): restore the popup, wheel and history contracts The Playwright suite went red on this branch, and five of its six shards ran out their 40-minute budget before reporting, so ~470 sync tests were never run at all. Three causes, all ours: - window.open() from page.evaluate() returned null. The popup blocker being ON (Firefox's default) and evaluate() no longer granting user activation are each defensible alone; together they block every gesture-less popup. ~35 tests, each burning 30s x 4 attempts x 2 worlds, which is what exhausted the shards. The blocker goes back to Playwright's and geckodriver's value. Reading it costs a detector a popup window the user sees, so it is not a check an anti-bot script in the page runs -- unlike navigator.userActivation.hasBeenActive, which is one property read, and which is why the activation half stays. - mouse.wheel(0, 100) delivered deltaY 114 (or 132, depending on the host's font metrics) in deltaMode 1. Quantising into native wheel notches is what a physical wheel does, but it changes the number the caller asked for, so it now rides behind humanize= with the rest of the humanized input. Default is the exact requested delta in deltaMode 0. - page.go_back() did nothing after history.pushState(). canGoBack is the BACK BUTTON's answer: under browser.navigation.requireUserInteraction it reports false when every entry behind this one was pushed without the user touching the page, which is now every entry, because evaluate() grants no activation. goBack() itself does not skip those entries and neither does history.back(), so ask canGoBackIgnoringUserInteraction, as Marionette does. Two keyboard tests are skiplisted rather than fixed: auto-Shift means typing "!" emits the Shift a US keyboard requires, and upstream asserts the character's three events with shiftKey false throughout. The character's own key/code/keyCode are unchanged; what upstream asserts is the absence of a Shift no real typist could omit. Full suite against the fixed build: 2223 passed, 6 failed -- the two keyboard tests above, and four client-certificate tests that fail only on this machine (Node/OpenSSL rejects the fixture server) and pass in CI. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pythonlib): stop overriding the corpus on core counts; pinning is opt-in Two findings from auditing the sweep's fixes against one bar: a difference is a leak only if a page's JavaScript can actually read it. hardwareConcurrency 2 was excluded from PLAUSIBLE_CORE_COUNTS because "2 is what Firefox reports under resistFingerprinting". That has not been true for years: RuntimeService::ClampedHardwareConcurrency hardcodes 4, and 8 on macOS, both of which are already in the table. The exclusion protected against nothing and cost every genuinely dual-core machine -- 20% of the macOS presets in the recorded corpus, 4.2% of Linux draws. It also made the small-host tail worse: a 3-core host reported 4, which cannot be pinned, so a page measured 3 while being told 4. At 2 the pin succeeds. pin_cpu_cores now defaults to False. What it buys is defence against a page timing N parallel workers; what it costs is a browser-wide CPU cap, a per-driver launch lock, and nothing at all on macOS. Unpinned, the host's own snapped count is reported, so reported and measurable still agree -- the identity just loses one drawn value. Callers who want the draw kept can still ask for it. The WebGL sampler keeps rejecting software rasterisers, and its docstring now says so: it described the opposite of what the code does. llvmpipe as the presented GPU is a live check on a string every fingerprint script reads, which is worth ~1.5% of corpus fidelity. 251 pythonlib tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pythonlib): keep 2 out of the core table -- an Apple M1 is never dual-core Reverts the PLAUSIBLE_CORE_COUNTS half of fd501e5. The conclusion there was wrong, even though the fact that prompted it was right. Right: "2 is what Firefox reports under resistFingerprinting" is false, and has been for years. RuntimeService::ClampedHardwareConcurrency hardcodes 4, and 8 on macOS, both already in the table. Wrong: concluding from that, and from the corpus recording 2 on 20% of macOS presets, that 2 should be drawable. 85% of macOS identities draw "Apple M1, or similar" as the WebGL renderer, and no Apple Silicon part has ever had fewer than 8 cores. A page reading navigator.hardwareConcurrency and UNMASKED_RENDERER_WEBGL together -- two property reads, both already in every fingerprint payload -- would see a machine that does not exist. The corpus frequency is not counter-evidence. Those rows report 2 more often than 4 on macOS (11 vs 3), which no real hardware population does: the corpus is scraped from live traffic, so it carries privacy-hardened browsers, 2-vCPU VMs and other people's bots. The corpus settles what real machines report where the field is hardware; hardwareConcurrency is a number a browser can be made to say. The comment now records the true reason, so the next reader does not undo this by discovering the RFP claim is false -- which is exactly how it came undone. pin_cpu_cores stays opt-in, and the WebGL software-rasteriser rejection is unchanged. 251 pythonlib tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pythonlib): a preset naming an unknown GPU must not fail the launch The pythonlib job failed on a Windows preset whose GPU is "ANGLE (Unknown, Adreno (TM) 650 Direct3D11 vs_5_0 ps_5_0)" -- a phone GPU, in the Windows pool. sample_webgl raised, and launch_options() with it. The crash is not new here: upstream/main has the same branch, which looks the preset's vendor/renderer up in webgl_data.db to get the parameters that belong to it. What is new is a test that draws a RANDOM preset, so it surfaced as a 1-in-11 flake instead of a deterministic failure. 39 of the 435 bundled presets name a GPU that is not among the 33 in the database, so ~9% of preset launches have always raised -- and a caller passing their own preset dict had no way to know which pairs are supported. The named GPU cannot be kept: parameters, extension list and shader precisions all have to come from one real recorded device, and there is none for an unknown renderer. So the fallback draws a GPU that fits the screen and REPLACES the pair. Replacing it needs the two keys popped first, because merge_into does not overwrite what the preset already set -- without that the page reads "Adreno (TM) 650" with a desktop GPU's parameters behind it, which is a louder mismatch than the one being fixed. Measured on the bundled presets: 30 of 312 now swap, e.g. a macOS preset claiming a 2008 Radeon HD 3200 becomes Apple M1. Guarded by a parametrised test over EVERY bundled preset in all three pools, asserting the pair that survives is one the database knows. It fails without the fix. 254 pythonlib tests pass. The rows themselves are corpus contamination and worth cleaning separately: the presets are scraped from live traffic, so the Windows pool carries an Android GPU and the macOS pool carries GPUs no Mac has shipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(patches): check both wheel modes, not just the notched one The stock-parity guard asserted that mouse.wheel(0, 300) arrives as three notched events. That is now the humanize=True behaviour, not the default, so the guard failed on the build it was meant to certify -- it encoded one side of a decision that has two sides. It now checks both: with humanize off, one event carrying the delta the caller asked for (deltaMode 0, deltaY 300); with humanize on, three events whose wheelDeltaY is a multiple of 120, as a physical wheel produces. A second short launch covers the humanized half, in the style of the timezone relaunch probe. Verified against the local build: PASS on every probe, with the humanized scroll arriving as 3 events of wheelDeltaY -120. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(patches): the query-cost probe measured the JIT, not the IPC query-cost failed in CI with "hardwareConcurrency 19 ms vs userAgent 0 ms". The 0 ms is the tell: the loop reads a getter and discards the result, so the JIT elided the BASELINE loop entirely on that runner. With the baseline at zero the check `costHwc > 5 * costUA + 15` collapses to a flat 15 ms allowance for 20000 reads -- 0.75 us each -- while a healthy read of a value that lives in the config costs ~1 us. It was timing whether the JIT dropped the loop. Every read is now accumulated into a sink that is returned, so the loop cannot be optimised away. (userAgent still measures ~0 because the string is cached, hence the second change.) The allowance on the two config-read checks goes to 40 ms. The state they guard against is a sync IPC per read, measured at ~12 us each when it regressed, i.e. ~240 ms over this loop; a healthy read is ~20 ms. 40 sits an order of magnitude under the defect and clear of a slow runner. The timezone-cost check keeps its 15 ms allowance and is commented to say why: its healthy numbers are ~2 ms vs ~1 ms and its regression is ~20 ms over the same loop, so widening it to 40 would step over the very thing it exists to catch. Verified against the local build: PASS, costHwc 7-13 ms, costLocalDate 1-2 ms. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(pythonlib): generate fingerprints with fpgen instead of BrowserForge BrowserForge (and the Apify fingerprint-suite data behind it) is replaced by fpgen -- scrapfly/fingerprint-generator, Apache-2.0, same author as Camoufox, trained on Scrapfly's live traffic. The reason is coverage. Measured over 120-200 Firefox draws per OS: BrowserForge / Apify fpgen distinct GPUs 2-3 per OS 6-9 sampled, 999 in the value space (webgl_data.db has 33) fonts 4-22 names 539-802 on macOS, 87 distinct sets on Windows media devices always empty for Firefox counts (labels are not collectable -- see below) voices absent real lists with voiceURIs system fonts/colours absent per keyword, per OS WebGL params absent params, extensions, shaderPrecisionFormats, contextAttributes audio absent 1238 distinct hashes This commit is the swap alone: fpgen supplies exactly what BrowserForge did -- navigator, screen, window geometry, Accept-Encoding -- through a new fpgen.yml mapping. The richer fields are NOT wired up yet; Camoufox still draws WebGL from webgl_data.db and fonts/voices from its own catalogues. That is the follow-up, and the one with the real diversity win in it. Notes for callers: - `camoufox.fingerprints.Screen` replaces `browserforge.fingerprints.Screen`, same four bounds. It becomes an fpgen predicate rather than a filter, and an unsatisfiable bound (a 1x1 Xvfb) falls back to an unbounded draw, as BrowserForge did by silently dropping the constraint. - `fingerprint=` now takes an fpgen dict, not a browserforge Fingerprint. - `from_browserforge()` is now `from_fpgen()`. - fpgen fetches ~7 MB of model data from GitHub on first use, so a fully offline install cannot generate a fingerprint. Presets and caller-supplied configs are unaffected. - doNotTrack is deliberately unmapped: Camoufox owns DNT, and a drawn value was being stripped anyway. - fpgen's own draws still need the coherence fixes -- conditioned on an Apple M1 renderer it still returns hardwareConcurrency 2 in 12% of draws. Changing the source does not remove the need for fix_hardware_concurrency and friends. 254 pythonlib tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(pythonlib): every identity passes a whole-identity coherence check Camoufox assembles an identity from pools sampled independently -- navigator and screen from the generator, GPU from webgl_data.db, fonts and voices from their own catalogues. Nothing compared them, so a machine that never existed could be built out of parts that are each fine on their own. Cleaning the pools cannot fix that: the incoherence is created at composition. Measured before this, per 200 generated identities: - 14% of macOS identities drew "Radeon R9 200 Series, or similar", a desktop PC card, or "Intel(R) HD Graphics 400", a Braswell Atom IGP. Both are in webgl_data.db's macOS column at 3.7% and 7.4%; neither shipped in a Mac. - 7% paired Apple Silicon with colorDepth 24. Deep colour is the macOS default, measured 30 on the Mac mini. - 2% of Windows identities reported maxTouchPoints 256. - 18 of the 312 bundled presets carried a GPU or a screen no desktop has, including a 736x414 iPhone viewport and a portrait 1440x2560. coherence.py states each invariant as a rule with the measurement behind it, repairs what has a determined correct value, and reports what does not. It runs on every identity whatever built it -- generated, preset, or caller-supplied. Where a value can be replaced rather than repaired it is dropped BEFORE the pool that would defer to it (a preset's own GPU pair wins over sampling, so an impossible pair is dropped and sampling draws a coherent one), and webgl_data.db is filtered by the same predicate before sampling, so the check and the draw cannot disagree about what a Mac may claim. After: 0 violations across 600 generated identities and all 312 presets. Guarded by tests/test_coherence.py, which checks each rule against the value that motivated it, asserts the three machines captured on 2026-09-17 pass as themselves, and walks every bundled preset. test_preset_screens_are_never_lifted asserted that a preset IS a real device and its screen must never be rewritten. That premise does not survive the data, so the test now distinguishes a small desktop panel, which is still kept, from a phone viewport, which is repaired. 273 pythonlib tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pythonlib): filter incoherent identities out of the shipped data too The coherence layer filters when an identity is drawn. This filters the data it is drawn from, so an impossible row never reaches a build: belt and braces, and the two use the same rules. scripts/clean-fingerprint-data.py checks each preset the way a launch converts it, and each webgl_data.db pair against the OS weights it is offered under. Presets are DROPPED rather than repaired -- repairing would write an invented value ("what core count does a 2-core Apple M1 really have?") into a file whose whole purpose is being real. GPU rows are kept with the impossible OS weight zeroed, because "Radeon R9 200 Series" is a genuine Linux and Windows card that simply never shipped in a Mac. Removed, with --write: - 38 of 435 presets: 27 with a GPU their OS cannot report, 7 pairing Apple Silicon with a core count Apple never shipped (2, 18), 4 with a colour depth their GPU contradicts, 3 with a phone viewport (736x414, 960x540, portrait 1440x2560), 1 claiming 40 touch points, 1 whose renderer is "Mozilla". - 2 impossible macOS weights in webgl_data.db (Intel HD Graphics 400 at 7.4%, Radeon R9 200 Series at 3.7%). macOS loses the most: 97 presets -> 63. Windows 255 -> 251, Linux 83 -> 83. tests/test_shipped_data.py asserts the files stay clean, so a refresh that reintroduces a bad row fails CI instead of shipping, and that no OS pool is emptied by the filter -- one GPU per OS would be its own tell. Two rules were loosened after checking them against real hardware rather than against the pools: maxTouchPoints is now a range (0..10) instead of a list of values seen in scraped data, because fpgen's Windows pool never offers the 5 that win-i9 actually reports; and APPLE_SILICON_CORES gained 11, which is the M3 Pro's 6P+5E. Over-filtering costs realism as surely as under-filtering. A new device-pixel-ratio rule snaps scraped artefacts (1.818, 1.09) to the nearest real scaling step, and rejects fractional DPR on macOS, which has none. The three captured machines (dpr 1 / 2.5 / 2) pass as themselves. The writer reproduces each file's own formatting byte-for-byte on a no-op run, so the diff of a real run is the dropped rows and nothing else. Verified: every surviving preset is unchanged and the metadata is untouched. 278 pythonlib tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: point CLAUDE.md at the coherence layer and the data cleaner Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: a pythonlib data tool must not cost an hour of compiling The build cache is keyed on a hash of the inputs that can change compiled output, and scripts/ is hashed wholesale because it holds the build machinery (patch.py, copy-additions.sh, package.py). It also holds tools that only rewrite the PYTHON package's data files, and those pay the same price: adding scripts/clean-fingerprint-data.py, which edits pythonlib JSON, invalidated a 665 MB cached browser and bought a full rebuild of a browser whose sources had not moved. Measured on this branch -- the only browser-input file changed across the last five pushes was that script. NON_NATIVE_SCRIPTS names the exceptions. Nothing is excluded for looking unrelated: excluding a script a build DOES run is the dangerous direction, since the cache would then serve a browser built from different sources while every suite downstream passed against it. So ci/tests/test_ci.py checks each entry against the files a build enters through (Makefile, multibuild.py, patch.py, package.py, copy-additions.sh, _mixin.py) and fails if one is reachable, and a second test fails if an entry no longer exists -- a stale exclusion stops excluding anything, quietly. Verified against the real cache: a clean checkout of 77edaeb hashes to fa62d4be1bdff96eee6e9018552255c4, which is the browser cached for this pull request, and a clean checkout of HEAD with this change hashes to the same value. So the next run restores that browser instead of compiling one. This does not change `browser_changed`, which asks a different question (build versus fetch the published release, against the pull request's base) and is correctly true for every push to a branch that has touched the browser once. 169 ci tests, 278 pythonlib tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pythonlib): keep the window's inner dimensions real, and its chrome honest Three patch guards failed after the fpgen switch -- humanize-edge-deadlock timed out at the full 600s, humanize-mouse-trajectory saw only its first endpoint, and mouse-boundary-sweep lost 15 of 25 ring points. All three are geometry, and both causes were mine. The same guards pass on the pre-fpgen pythonlib with the same binary, which is how the two were separated from a browser fault. **Inner dimensions must stay real.** Playwright's setViewportSize resizes the window and then waits for the page to report the size it asked for. A spoofed innerWidth/innerHeight never changes, so that wait never returns -- the trap no_viewport already exists for (#666), reached here through an explicit set_viewport_size() call. BrowserForge hid it by accident: its Firefox samples carry innerWidth/innerHeight as 0, and _cast_to_properties skips falsy values, so they were never spoofed. fpgen reports the real numbers, and mapping them turned an unmapped field into a hang. They are no longer mapped, which also means the content area a page measures is the one it actually has. **The claimed chrome height cannot be smaller than the real one.** The window is sized from window.outerHeight, so the content area that can receive input is outerHeight minus the browser's own 86px of chrome. An identity claiming `outerHeight - innerHeight` below that claims a viewport taller than the window can hold, and the difference is dead: mouse events dispatched into those rows reach nothing. Measured -- a drawn pair of outer 801 / inner 717 (chrome 84) left the bottom 2px unreachable, which is exactly what the boundary sweep saw, always on the bottom two rows. browserforge's pairs were never below 86; fpgen's sometimes are. New coherence rule, repaired by growing the window where the screen has room and shrinking the viewport where it does not. humanize-mouse-trajectory also had a latent bug of its own: it moved to a flat (1100, 650), which silently tested nothing whenever the drawn window was smaller -- the move landed outside the content area and the run failed reporting only the start point. One drawn window was 924x1364, narrower than that x. It now derives the destination from the viewport it actually got. 22/22 patch guards pass locally (5m37s); 279 pythonlib tests pass, including a regression test that no generated config carries window.innerWidth/innerHeight. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fonts): draw a real OS-version base, at measured real-world rates The font draw modelled a machine as "always-present core + a flat 30-78% sample of everything else". Neither half held up: * the sample was UNIFORM, so every addition was equally likely. Office (on ~60% of real Windows machines) and the Pan-European supplemental pack (~2.8%) were drawn at the same rate, and `kind`, `prob`, `requiresLocale` and `sizes` from scripts/data/font-manifests.json were thrown away when font-groups.json was written by hand. * `_ESSENTIAL_FONTS_MACOS` was a SUPERSET of one base (553 entries), which silently forced 131 Sonoma-only families onto every macOS identity. It has to be the INTERSECTION of that OS's bases, or adding a second base does nothing. A machine now draws one OS-version base, whole and never subsetted, then each addition unit independently at its own probability. The bases are measured from real sources rather than inherited: Ubuntu 24.04 / 26.04 official desktop ISO manifest -> the shipped .debs -> fc-scan. Firefox enumerates via fontconfig on Linux, NOT nameID 1: the same file reports "Noto Sans MeeteiMayek" in nameID 1 and "Noto Sans Meetei Mayek" in fontconfig. The old base named 26 families Ubuntu does not ship and missed 41 it does. Windows 11 26200.9457 a real box. Office is installed there, so OS-native files are separated by WinSxS hardlink (an OS font has one, an Office font does not): 143 of 340 files. Scanning English-only drops the 17 localized CJK names, so all language ids are kept. macOS 26.6.2 / 27.0 a real Mac, verified clean. They differ by three families: 27 drops Noto Sans Brahmi and Noto Sans CanAborig and adds Noto Sans Sunuwar. macOS Sonoma NOT re-verifiable (that machine has since been upgraded). Six family names were stored mojibaked (Shift-JIS read as UTF-16BE) and are repaired here. Windows 10 is dropped (end of support Oct 2025), so win11 carries weight 1.0 and the Win11 families are part of the base; the NATIVE path now subtracts them on a Windows 10 host instead of adding them on a Windows 11 one. Apple ships 189 families as optional Font Book downloads rather than enabled by default. The fpgen corpus independently puts all 57 reportable ones at exactly 81.8%, so they are an `apple-optional` unit, not base. Dot-prefixed macOS families are stripped: CoreText excludes them from enumeration. Adobe's fonts are removed from the bundle (185 files, 240 MB): Adobe permits no redistribution, and the corpus puts every Adobe family on under 2% of real machines, so they bought no realism. The `adobe-cc` unit goes with them, because reporting a family the bundle cannot render is a reverse leak. font-groups.json and the new font-bases.json are now GENERATED by scripts/gen-font-groups.py instead of maintained by hand, which is what keeps the draw's probabilities equal to the manifest's. tests/test_font_distribution.py asserts the properties verify-fonts.py cannot: complete-base containment, base weights, per-unit probability, bundle atomicity, a-la-carte piecemeal sizing, locale gating, and variation. Checked by mutation: reintroducing the uniform sample fails 3 tests, collapsing variation fails 14, truncating a base fails 1. The font bundle itself is deliberately not in this commit: it is 3.7 GB and one file is 184 MB, over GitHub's push limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(fonts): store each face once, fetch the bundle as a release asset The bundle was three per-OS directories, so a face used by more than one OS was stored more than once: 2955 files, 3.96 GB, of which 1.63 GB (41%) was byte-identical copies. That was not a decision, it was the absence of one -- the DIRECTORY was the only selection mechanism, so fontconfig could be scoped to an OS only by giving that OS its own full copy of everything it needed. Each face is now stored once, in a directory named for the set of OSes that use it (L, M, W, LM, LW, MW, LMW). An OS reads the four groups its letter appears in; bundle/fonts/groups.json records the mapping and utils._generate_fontconfig hands fontconfig exactly those directories. 1513 faces, 2.16 GB lin -> L+LM+LMW+LW (792) mac -> LM+LMW+M+MW (779) win -> LMW+LW+MW+W (809) package before after linux 3.96 GB 2.16 GB (-46%) macos 2.21 GB 1.31 GB (-41%) windows 2.69 GB 1.79 GB (-34%) pythonlib/camoufox/fonts.json is BYTE-IDENTICAL across the change: the same families are renderable and reportable on all three OSes. The saving is pure redundancy. The group directory is also a better gate than what it replaces. The 455 <rejectfont> globs in fontconfig/windows/fonts.conf are gone: a face Windows must not see is simply not in a group Windows reads (verified: 0 of the 243 formerly rejected basenames appear in any Windows group). Those globs were unsound anyway -- 114 bundled filenames contain [ ] (e.g. ReemKufi[wght].ttf), which fontconfig parses as a character class, so they silently matched nothing. macOS (CoreText) and Windows (DirectWrite) cannot read subfolders, so package.py flattens the groups for those targets and the allowlist gates them, as before. The bundle itself leaves git. It is 2.16 GB extracted and 843 MB as .tar.xz; GitHub rejects files over 100 MiB, and an xz archive cannot be delta-compressed, so committing it -- split into ~9 parts -- would append the whole archive to history on every font change, paid by everyone who clones the repo. It is now a release asset in its own tag namespace (font-bundle-*, excluded from build.yml so it does not trigger a browser build), fetched on demand: make fetch-fonts download + verify make fonts-extract unpack make fonts-check verify only make fonts-clean drop the unpack scripts/data/font-bundle.json pins the asset name, size and sha256, so a commit still names exactly one bundle and a truncated download fails loudly instead of producing a browser that reports fonts it cannot render. This is the trust model the build already uses for the Firefox source (`make fetch` pulls a ~500 MB tarball from archive.mozilla.org), so a clone was never buildable offline. Only the compressed archive is kept; bundle/fonts/ is extracted on demand and is gitignored. Round-trip verified byte-identical. package-* now depend on fonts-extract, and gen-fonts-json.py / verify-fonts.py fail with the fix ("run make fonts-extract") rather than a stack trace. 642 font blobs (974 MB) leave the index. bundle/fonts/000_README.txt moves to bundle/FONTS-README.txt and cleanfonts.sh to scripts/, since bundle/fonts/ is now ignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci(fonts): fetch the font bundle before packaging The bundle stopped being repo content in the previous commit, so a tagged build would package a browser with NO fonts at all while pythonlib/camoufox/fonts.json still reports 335-567 families per OS -- every one of them a family the browser cannot render, which is the reverse leak this whole series exists to remove. `make fonts-extract` downloads and unpacks it (verified against the sha256 in scripts/data/font-bundle.json), and verify-fonts.py then asserts the bundle and the manifest actually agree before anything is packaged. A build that would have shipped a mismatched font set now fails in CI instead. Net disk in CI goes DOWN: 3.96 GB of fonts from the checkout becomes a 0.84 GB archive plus 2.16 GB extracted. The tests workflow is untouched: it reads the generated JSON, not the bundle. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fonts): drop the stale README the v1 archive still carries bundle/fonts/000_README.txt and cleanfonts.sh were inside the bundle when the v1 archive was built, and they moved to bundle/FONTS-README.txt and scripts/ in the same change that took the bundle out of git. Extracting therefore re-created a stale copy of a file git now owns -- invisible to git (bundle/fonts/ is ignored) but contradicting the tracked README, and it would have been folded back in if a later bundle were rebuilt from an extracted tree. Prune both after extraction. A future archive will not contain them, at which point this is a no-op. Extraction verified reproducible: two consecutive `make fonts-extract` runs produce a byte-identical 1514-file tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fonts): stage the groups, not per-OS dirs that no longer exist scripts/stage-fonts.sh copied bundle/fonts/{linux,macos,windows} into an unpackaged build's dist/bin. Neither half of that is true any more: the bundle stores each face once under a group directory (L, M, W, LM, LW, MW, LMW) and is a release asset rather than repo content, so on a fresh checkout the source paths do not exist and bundle/fonts/ itself does not either. Under `set -e` the cp fails, which fails `make stage-fonts`, which fails the "Package the binary for the test jobs" step in tests.yml -- so the build job, and with it the required gate, would have gone red on every pull request. Locally it broke the patch guards and build-tester the same way. Stage the group directories verbatim instead, plus groups.json, which is what utils._generate_fontconfig reads to decide which of them a claimed OS may see. Staging the groups rather than a flattened copy is what makes the per-OS gate work against an unpackaged build exactly as it does in a package. groups.json is copied last, so an interrupted copy leaves no marker and the next run stages again instead of trusting a partial tree; the group dirs are copied by name so fetch-fonts.py's .bundle-sha256 bookkeeping file stays out of a browser's font directory. `stage-fonts` now depends on `fonts-extract`, since a fresh checkout has nothing to stage from. To keep that cheap enough to run before every launch, fetch-fonts.py stamps the unpacked tree with the archive's sha256 and `--extract` returns immediately when it matches: 30s to 0.03s, and it no longer needs the 843 MB archive to still be on disk. A bundle bump changes the sha256, so a stale tree is replaced rather than trusted. The artifact check in tests.yml asserted fonts/linux; it now asserts fonts/groups.json and fonts/LMW. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(fonts): prove no identity can reach another OS's faces verify-fonts.py handed every OS the bundle ROOT as its fontconfig <dir>. fontconfig scans <dir> recursively, so all three OSes were being tested against all 1513 faces -- which is why each reported an identical "fc-list publishes 1307 families". The per-OS group gate, the thing that replaced the 455 Windows reject globs, was therefore never checked by anything. Hand each OS the four group directories it actually reads, as utils._generate_fontconfig does, and then assert the converse of the existing invariant: ask fontconfig for every file it can reach under that conf and require all of them to sit inside that OS's own groups. The existing checks only prove an OS can render what it reports; nothing proved it cannot reach what it must not, and no reported name would reveal it -- a Windows-only or macOS-only face on the search path is a glyph-fallback candidate, so an emoji or CJK glyph could resolve to Segoe UI Emoji or PingFang on a machine claiming Linux. The gate holds: 810 / 780 / 793 faces reachable for win / mac / lin, all inside their own groups, and the per-OS family counts are now honestly different (529 / 872 / 583 published against 335 / 567 / 358 reported). Mutation-tested -- putting the root back makes all three fail with 1513 reachable faces. Also in build.yml: install dependencies before fetching the bundle, so the 843 MB download uses aria2c's parallel connections instead of silently falling back to single-connection curl, and add fontconfig explicitly since the verify step resolves every reportable family through fc-list/fc-match. The over-100-MiB report is no longer a warning: it is the settled reason the bundle is a release asset, not an outstanding problem. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(fonts): describe the bundle that actually ships docs/FONTS.md still described a world several changes back: one bundle per OS at bundle/fonts/{linux,windows,macos}, a "target bundle" not yet in git, a Windows 10 base, Sonoma as the only macOS base, a uniform 30-78% draw, the 455 reject globs, Adobe CC among the additions, and a Linux package that duplicates fonts. Every one of those is now wrong, which makes the document worse than no document. Rewritten against the shipped data, with the figures read out of the files rather than recalled: the release-asset workflow and the sha256 pin, the group layout and which four groups each OS reads, both invariants (reported ⊆ renderable, and nothing outside an OS's groups reachable) and which one each package type relies on, the per-OS base weights and the per-unit bundle / à-la-carte probabilities as tables, and the resulting draw sizes. The namespace trap that produced two wrong font lists during this work -- Windows/macOS enumerate nameID 1, Linux enumerates through fontconfig, and system_profiler/WPF report nameID 16 -- is written down so it is not rediscovered a third time. Known residue is now stated rather than implied: the Linux base over-claims against the corpus because the bundle ships 30-metric-aliases, the macOS base weights are estimates where the addition probabilities are measured, nothing has been checked in a running browser, and font-bundle.json points at a release on this fork, which has to be re-uploaded and re-pinned before upstream can build. Also: per-context-patches.md described the font staging as OS subdirectories and named createRuntimeFontconfig(), which does not exist anywhere in the tree (it is utils._generate_fontconfig); gen-fonts-json.py's header still said it scanned bundle/fonts/<os> and listed Adobe CC; licences.py pointed at the README's old path inside the bundle. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fingerprints): pin fpgen's model, and close the taskbar gap it exposed fpgen supplies every synthetic fingerprint, and it does not ship its model -- it downloads one on first import, and again whenever the files are over five weeks old. Four things are wrong with that fetch, all in fpgen/pkgman.py: * TLS verification is disabled on BOTH the API call and the download (verify=False), so anyone on the path can serve the model; * the archive is never checksummed, and goes straight into extractall() with no path-traversal guard; * the GitHub API is called unauthenticated, on a rate limit shared by every job on the runner's IP; * it takes the FIRST release the API lists. model-4/2025 and model-2/2026 carry an identical created_at (2025-03-22, inherited from the tag's commit), so the sort ties and resolves to the lower id -- model-4/2025. The consequence is that every Camoufox generates from an April-2025 corpus and cannot be talked into anything newer: newest Firefox 137, newest GPU an RTX 40, no RDNA4. The 2026 model has been sitting unreachable for seven months. scripts/pin-fpgen-model.py installs the model named by scripts/data/fpgen-model.json before anything imports fpgen, with verification on and the sha256 checked, and refuses any member that escapes the data directory. Finding fpgen's data dir must not import fpgen -- importing is what triggers the download -- so it reads the module origin via find_spec without executing it. Wired into all seven CI jobs that install pythonlib. Pinning to model-2/2026 then failed tests/test_launch_geometry.py about 30% of the time, on `availHeight < height`. That turned out to be our bug, not the model's: fix_screen_no_taskbar only fired when avail equalled screen on BOTH axes, so `availWidth < width, availHeight == height` -- a Windows taskbar docked left or right -- passed straight through and the identity claimed no vertical chrome at all. Rare shape, but the 2026 corpus produces it in ~30% of draws once conditioned on a small display, against under 1% unconditioned. Trigger on the height alone: a side dock is far rarer than a Mac menu bar, a bottom taskbar or a Linux panel, so the vertical delta is worth more than the few genuine side-docked machines it overwrites. 316 passed, 2 skipped with the 2026 model pinned. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fonts): host the font bundle upstream, not on a contributor's fork scripts/data/font-bundle.json pointed at a release in JWriter20/camoufox, so merging this would have left daijro/camoufox fetching a required build input from a personal fork -- a build that breaks if that fork is renamed, made private, or has the release deleted, by someone with no obligation to keep it. The identical asset is now published at daijro/camoufox under the same font-bundle-v1 tag, so only `repo` and `url` move here; `size` and `sha256` are byte-for-byte what they were, which is the point -- the pin proves the bytes did not change when the host did. The upstream release is deliberately NOT marked latest: v152.0.4-beta.30 holds that badge, and a build input must not displace the browser download people actually come for. `font-bundle-*` tags are already excluded from build.yml, so publishing it triggered no browser build. Verified against the new host from scratch: local archive moved aside, `make fetc…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )