Repository navigation
Conversation
The custom OpenSSL build exists so blasthttp can reach servers that only speak deprecated ciphers, but nothing was reaching them. `SslConnector:: builder` installs its own cipher list, `DEFAULT:!aNULL:!eNULL:!MD5:!3DES: !DES:!RC4:!IDEA:!SEED:...`, and we never replaced it, so RC4, DES, 3DES and SEED never made it into the ClientHello. A server speaking only one of them was unreachable unless the caller passed `cipher_string` by hand. `set_security_level(0)` looked like it covered this and does not. The security level decides how weak a negotiated cipher may be; the cipher list decides which ones are offered at all. Set `ALL` when the caller names nothing, on both the pooled path and the `connect_stream` path used by `raw_connect`, `resolve_ip` and `request_target`, which build their SSL contexts separately. Null-encryption suites stay out of the default. They remain available through an explicit `cipher_string`, but negotiating one by accident would return a connection that looks like TLS and encrypts nothing. tests/legacy_default.rs covers this end to end: each test stands up a real TLS server pinned to one legacy cipher or protocol version and connects with a client that sets no TLS options at all. All five cipher cases failed before this change. Side effect worth knowing: the ClientHello grows from 31 cipher suites to 105, so the client's JA3/JA4 fingerprint changes. README dropped its claims about export ciphers and SSLv3. OpenSSL removed export ciphers in 1.1.0, and SSLv3 is disabled in our build despite `enable-ssl3` being passed, so neither has ever worked.
SSLv3 has never been reachable in any build, despite the README advertising it and the build script asking for it. Three separate things were in the way. The build. `scripts/build-openssl.sh` passed `enable-ssl3`, but OpenSSL's Configure carries a disable cascade that reads "if ssl3-method is off, turn ssl3 off too", and ssl3-method is off by default. So the flag was undone during configure. `configdata.pm` recorded both `enable-ssl3` and `no-ssl3`, and every shipped build came out with OPENSSL_NO_SSL3 defined and no SSLv3_method symbols in libssl.a. Passing `enable-ssl3-method` alongside it fixes this. The SSL options. `SslConnector::builder` sets NO_SSLV3 and nothing cleared it, so even a build with SSLv3 compiled in would refuse to negotiate it. Cleared on both the pooled path and `connect_stream`. The spelling. `parse_tls_version` understood 1.0 through 1.3 and nothing else, so there was no way to name SSLv3 at all. It now takes `3.0`, `ssl3` and `sslv3`, and the error message mentions it. An SSLv3-only server is now reachable without asking, matching how the legacy ciphers behave after the previous commit, and can also be pinned explicitly. The test server clears NO_SSLV3 too so a test can pin it. Separately, the build cache was keyed only on the target, so changing what the build asks for was silently ignored and a stale install stayed in place. It now keys on the version and feature flags as well, which is what makes this fix reach anyone with an existing checkout. Upgrading triggers one rebuild. Measured: JA4 is unchanged, `supported_versions` gains 0300, and google, github, cloudflare, amazon, reddit, wikipedia, microsoft and apple all still return 200.
First piece of the 1.x benchmark work. Nothing in the impersonation plan is verifiable without a way to see what we actually put on the wire, and doing that against public fingerprinting services does not scale: they rate-limit, several that the ecosystem still cites are dead, and one has been hijacked and now serves a gambling site. So measure it locally. The test server installs a client_hello callback, which fires after the hello is parsed and before a cipher is chosen, and records the offer. That is the only point where the full ClientHello is still visible. The safe openssl wrapper covers ciphers and versions but has no accessor for the extension list, which is most of a fingerprint, so that part goes through openssl-sys. Both new dev-dependencies were already in the tree as transitive dependencies of openssl. JA4 rather than JA3, for a reason that decides the shape of the stealth work: JA4 sorts its lists and strips GREASE, and extension order plus GREASE are exactly the two things OpenSSL cannot control. They cost us nothing on this metric. JA3 is also a poor target in its own right, since Chrome permutes extensions per connection and one Chrome build yields a different JA3 on every handshake. The implementation is checked against ground truth rather than against itself: ja4_conformance.rs computes the JA4 of a captured Chrome 150 hello from lexiforest/curl-impersonate's signature corpus (MIT) and asserts the published value, which four independent projects agree on across three fingerprinting services. All three segments match. fingerprint_baseline.rs pins where we stand: t13i9911h2 locally, against Chrome's t13d1516h2. It also records the gaps as assertions that will flip as each is closed, so no GREASE anywhere, none of the six browser-only extensions, and every protocol version from SSLv3 up on offer. Shared test helpers move to tests/support/. Cargo compiles every top-level file in tests/ as its own test binary, so tls_server.rs was already building as a target with no tests in it, and a second helper next to it failed outright once it referenced a sibling. One thing worth knowing that fell out of this: JA4's count fields are two digits, so our 105 cipher suites render as 99. Past that point JA4 cannot tell us apart from any other client with an unreasonable cipher list.
Second piece of the 1.x benchmark work. The capture side answers what we send; this answers what happens to us. A scan has to tell three things apart that all arrive looking like a response: the origin answered, a protection product answered instead, or the product answered and handed us a session anyway. Collapsing those is what turns a protected host into an apparently dead one, which is the false negative the whole effort is about. Classification keys on headers, cookies and status, never on a vendor's name in the body. A first pass did the latter and scored datadome.co, imperva.com and kasada.io as challenges from their own marketing copy; that case is now a test. Two things only showed up by running it against the real thing, and both came from a fixture I had assumed rather than captured: - www.akamai.com's 403 carries no `server` header at all. It identifies itself with akamai-grn, x-akam-sw-version and an ak_p server-timing entry, so requiring `server: AkamaiGHost` matched nothing. - that page entity-encodes its own punctuation, so the body holds `errors.edgesuite.net` and a search for the dotted hostname never fires. Matching the bare label survives it. The live benchmark is opt-in (`--ignored`) and prints a scorecard rather than asserting, since "we are not blocked" would fail for reasons outside this repository. It also cross-checks our local JA4 against a remote observer, which surfaced a known divergence worth recording: the JA4 header is fixed width and the spec caps counts at 99, our legacy list is 105 suites, and tls.peet.ws prints 105 where we print 99. Both hash segments agree exactly, so only the rendering differs. A stealth profile at 15 suites removes it. Current scorecard: 1/4 through. cloudflare.com present, akamai.com blocked, both scrapingcourse challenge pages unreachable as expected since that tier wants JavaScript.
The scorecard reported "1/4 got through" when the real bypass count was zero. Two ways that number lied. The one success was cloudflare.com, which is a control. Its homepage challenges nobody, so every client reaches it and passing says nothing about us. Counting it made a 0% bypass rate read as 25%. It also put the JavaScript tier in the denominator. Cloudflare's managed challenge and the scrapingcourse pages want the browser to execute their code and report back, which no HTTP client does, curl_cffi included. So the score could never approach 4/4 however good the fingerprint work got, and a permanently unreachable denominator makes the metric useless for tracking progress. Now targets carry a tier. Controls are verified and not scored, and a control that fails aborts the run, since a blocked control means the network path or the site changed and nothing else in the run can be trusted. The JavaScript tier is reported so a change would be visible but never counted against us. Only the passive-fingerprint tier is scored, because it is the only one this work can move. Reads 0/1 on the passive tier today. That is the number stealth mode has to change, and curl_cffi already scores 1/1 on it.
Starts closing the gap the benchmark measures. A profile is data rather than code, because browsers ship every few weeks and anything needing a rewrite per release will not be kept current. Chrome 131 rather than something newer, deliberately. Chrome 133 and later sign with ML-DSA, codepoints OpenSSL 3.3.2 does not know, so their signature_algorithms list cannot be reproduced and their JA4 is out of reach. The cipher and extension sets are identical from Chrome 124 through 150, so only the signature algorithms are dated. Both TLS builders now go through one `apply_tls_settings`, which also closes a standing hazard: they were copy-pasted siblings, so anything applied to only one silently gave raw_connect, resolve_ip and request_target requests a different fingerprint from ordinary requests. Where this leaves the JA4, against Chrome 131's t13d1516h2_8daaf6152771_02713d6af862: was t13d9911h2 105 ciphers, 11 extensions now t13d1614h2 16 ciphers, 14 extensions Signature algorithms match exactly. Every cipher matches except one, and four extensions are now sent that were not: 0005, 0012, 4469 and fe0d. Three gaps remain, and all three are things OpenSSL keeps for itself. Probed rather than assumed: it refuses a custom extension for 0x0005, 0x001b and 0xff01, and accepts 0x0012, 0x4469 and 0xfe0d. status_request was solved without a patch by setting it per-connection, which is the only place OpenSSL's own API for it works. The other two need build changes, along with suppressing the TLS_EMPTY_RENEGOTIATION_INFO_SCSV that inflates the cipher count by one; suppressing that and sending renegotiation_info are the same change, since OpenSSL sends the SCSV precisely when it is not sending the extension. Two traps found by measurement, both worth knowing: The `openssl` crate stores a custom extension's payload in ex_data keyed on the payload TYPE, so every extension returning Vec<u8> shares one slot and overwrites the others. Three registered that way send one. Distinct newtypes give each its own slot. The harness was under-reporting. SSL_client_hello_get1_extensions_present only returns extensions the server's OpenSSL has a definition for, so ALPS and ECH were invisible to it, which is to say it was blind to exactly the extensions being added. Capture now parses the ClientHello off the wire, which also preserves extension order and GREASE for later. Local and remote readings now agree exactly.
akamai.com now answers 200 with a real session where it returned a hard 403 before. That is the measurement this work exists to move: the passive-fingerprint tier goes 0/1 to 1/1. The OpenSSL patch does most of the TLS work. Stock OpenSSL signals renegotiation-info support with the SCSV pseudo-cipher where every browser uses the extension, which costs a match twice over: a cipher nothing else sends, and a missing extension everything else sends. Switching it fixes both, and the cipher segment of our JA4 now equals Chrome 131's exactly. The patch is conditioned on the TLS floor rather than applied unconditionally, because the SCSV exists to protect pre-RFC5746 servers that mishandle unfamiliar extensions. A client whose floor is TLS 1.2 is modern and sends the extension; one willing to speak older protocols keeps stock behaviour. The compatibility path is unchanged, and a test asserts it. Build patches are now part of the cached build recipe, so editing one invalidates the build. Without that a patch change would silently do nothing, which is how the SSLv3 flag went unnoticed for months. Three of Chrome's extensions are deliberately NOT sent, each tried and each withdrawn for breaking real sites: ALPS (0x4469): google.com negotiates it, then expects the settings exchange inside HTTP/2 that we do not implement. TLS completes and the request then fails. ECH (0xfe0d): a hand-built GREASE blob is not shaped precisely enough. reddit.com rejects it at the handshake. Registering the server-reply contexts fixed cloudflare.com but not reddit, so the payload is the problem rather than the registration. compress_certificate (0x001b): our build is OPENSSL_NO_COMP_ALG and cannot decompress a certificate. Enabling one means a compression library across six cross-compiled wheel targets. That last point generalises, and is the finding worth carrying forward: advertising a protocol feature is a promise to implement it, and a server that takes you up on it gets a client that cannot follow through. So the JA4 does not match Chrome exactly, and did not need to. Thirteen extensions against Chrome's sixteen was enough to turn Akamai around. Exact parity was never the bar; plausibility was, and a client that connects beats one that matches a hash and cannot. Verified on akamai, cloudflare, google, github, wikipedia, reddit and microsoft: all reachable with the profile.
src/client/hyper.rs made no http2_* calls at all, so every SETTINGS value and the connection window were hyper-util defaults, which match no browser. Four of them are reachable through hyper-util and now follow the profile: before 2:0;4:2097152;5:16384;6:16384|5177345|0|m,s,a,p after 2:0;4:6291456;6:262144|15663105|0|m,s,a,p chrome 1:65536;2:0;4:6291456;6:262144|15663105|0|m,a,s,p INITIAL_WINDOW_SIZE, MAX_HEADER_LIST_SIZE and the WINDOW_UPDATE now match Chrome exactly. Omitting MAX_FRAME_SIZE is part of the match rather than an oversight: our default announced 16384 and Chrome announces nothing. Adaptive window is off because it resizes the connection window as traffic flows, emitting WINDOW_UPDATE frames no browser sends. Two fields are left alone, both fork-gated rather than forgotten. HEADER_TABLE_SIZE exists in hyper's own config but hyper-util does not expose it. The pseudo-header order is hardcoded in h2's Iter::next(), and reaching it means forking h2, hyper and hyper-util, since each layer translates a fixed option set rather than passing options through. Worth noting that curl_cffi gets Chrome's order for free: it is nghttp2's natural order, and Rust's h2 simply chose differently. Verified reachable with the profile: akamai, cloudflare, google, github, reddit.
"We are blocked" does not say which part of the imitation failed, and rebuilding between experiments makes finding out slow enough that you guess instead. BLASTHTTP_BISECT=tls,headers,http2 turns off any combination of the three layers at runtime. It paid for itself immediately. PerimeterX blocks blasthttp+chrome while letting both plain blasthttp and curl_cffi through, which I had assumed meant our TLS fingerprint was not good enough. It is not the TLS. On priceline.com, three passes each: plain (all off) 3/3 tls only 3/3 headers only 0/3 http2 only 3/3 tls + http2, no headers 3/3 full profile 0/3 Narrowing into the headers, it is the User-Agent on its own. Adding only sec-ch-ua, only sec-fetch-*, or only accept/accept-language/priority to an otherwise plain client changes nothing; adding only a browser User-Agent drops it to zero. And it is any browser, not Chrome specifically: blasthttp/0.10.1 (default) 3/3 Chrome 131 Windows 0/3 Chrome 150 macOS 0/3 (byte-identical to curl_cffi's) Firefox 147 0/3 NotARealBrowser/1.0 3/3 curl_cffi sends that same Chrome 150 macOS string and passes 3/3. So the User-Agent is not what is scored; it selects which scoring applies. Claim to be a browser and the client gets checked against that claim. Claim anything else, including nonsense, and it does not. That model explains results that looked contradictory. Akamai and DataDome reward the profile because they fingerprint passively and ours is close enough. PerimeterX verifies the claim, curl_cffi survives the verification and we do not, so against it our partial imitation is worse than being honest. It also generalises the udemy.com case from the research: not "browser headers over script TLS" but "any browser claim invites verification". Two consequences. A browser profile should not become a global default, since it makes us worse against anything that verifies. And the case for closing the remaining gaps, GREASE, extension order, ALPS, ECH, post-quantum key share, is now specific rather than aesthetic: they are what the verification is looking at.
Phase 1 of the profile plan, and a live bug on its own. Nothing that
chooses a connection profile can work until a failure says what went
wrong, and until now none of them did.
The reason was being destroyed on the way out. The connector hands
hyper-util a Box<dyn Error>; hyper-util wraps it in an error whose
Display is the literal string "client error (Connect)" for every
connect-time failure; nothing ever calls .source() to get back down to
the openssl error underneath. So the substring test that was supposed to
tell TLS failures apart could never match, and a cipher mismatch, a
rejected certificate, a DNS failure and a refused connection all arrived
identically as ErrorKind::Connection.
That also meant TLS failures were being retried, since Connection is
retryable and Tls is not, contradicting test_tls_error_is_not_retryable.
Now the connector records the reason before boxing it away, into a map
keyed by host and port. Keyed per host because a cached client serves
many hosts at once, and CertSlot already shows what a single shared slot
does under concurrency. The map lives on HyperClient rather than on
CachedClient: a retry with different TLS settings builds a *different*
cached client and still needs to read why the previous attempt failed,
so hanging it off the cached client would partition the map by exactly
the thing it exists to inform.
TlsFailure is derived from OpenSSL reason codes rather than message
text, and distinguishes the cases that lead somewhere different:
NoSharedCipher and UnsupportedProtocol mean a wider offer might work,
CertificateVerify means it never will, Reset and Timeout mean we learned
nothing about our ciphers.
Alert counts as "try wider" too, and that is the case that matters.
Running it against an RC4-only server offered AES returns
`ssl/tls alert handshake failure`, not `no shared cipher`: the latter is
what OpenSSL raises when *we* work it out locally, while a real server
just sends a fatal alert and declines to explain. Excluding alerts would
have meant never widening the offer in the single most common case. The
first version of this did exclude them; three test servers found it.
The stack is scanned for a code we recognise rather than read at index
zero, because OpenSSL pushes several errors per failure and a generic
"handshake failure" often sits on top of the one that explains it.
Python gets a TransportError (subclassing RuntimeError, so existing
`except RuntimeError` still works) carrying `kind`, `tls_failure` and
`retryable` as attributes. Six call sites were doing
PyRuntimeError::new_err(e.message) and dropping everything else.
Measured end to end against local servers:
cipher mismatch TLS handshake failed: peer sent an alert: ssl/tls
alert handshake failure
version mismatch TLS handshake failed: no protocol version in common
nothing listening request failed: client error (Connect)
The third is the one that matters as much as the other two: a refused
connection must not claim to be a TLS problem, or the ladder would widen
its cipher offer at a host that was never listening.
Phase 2. The classifier already existed under tests/, regression-tested
against real captured responses, and took exactly the shape of a
response. It was just unreachable from anything that ships.
Now on Response::protection(), computed rather than stored so it stays
right for a response built by hand, and exposed to Python as a
(outcome, vendor) pair.
The distinction worth having is challenge versus blocked. A blocked
request might succeed with different connection settings; a challenge
never will, because answering it means running the page's JavaScript.
That is the signal to hand a URL to a real browser instead of retrying,
and BBOT should not have to find it by grepping our response body for
cf-mitigated.
Verified from Python against the wheel:
akamai.com 403 ('blocked', 'akamai')
scrapingcourse cf-challenge 403 ('challenge', 'cloudflare')
google.com 200 ('ok', '-')
And the error path from the same build, which Phase 1 made possible:
refused connection TransportError kind=connection tls_failure=None
retryable=True
cipher mismatch kind=tls tls_failure=alert retryable=False
version mismatch kind=tls tls_failure=unsupported_protocol
retryable=False
Together those are the whole contract a consumer needs: whether we got a
response, whether something intercepted it, whether retrying could help,
and whether a browser is the only way through.
Phase 3. BrowserProfile becomes ConnectionProfile, and there are now three: compatibility, modern, chrome131. The trap this had to avoid: several behaviours were gated on *whether a profile existed* rather than on anything in it. The padding and encrypt_then_mac flips, add_browser_extensions, the OCSP status_request and http2_adaptive_window were all "if a profile is set". Naming today's default behaviour as a profile would have silently switched every one of them on for the configuration that is supposed to be unchanged. They are now explicit fields: `browser_extensions` on the TLS half, `http2` as an Option, headers as a possibly-empty list. So compatibility is a transcription, not a cleanup, and every None in it is load-bearing. tests/legacy_default.rs and test_compat_path_keeps_the_scsv are its specification, and all 314 tests pass unchanged. Confirmed on the wire too: no profile and --profile compatibility both produce t13d10512h2_c86627d460ee_c50a3655fff1 with 105 suites, bit for bit. resolved_profile() is now infallible. Naming no profile means the default rather than "no profile", which collapses five Option branches, and an unknown name is an error instead of silently resolving to whatever the default happened to be. modern is new: 11 suites instead of 105, TLS 1.2 and 1.3 only, no browser claim. One correction, because the measurement contradicted the argument I wrote it on. modern was justified as "unremarkable in real traffic". It is not. Checked against tlsfingerprint.io, which indexes around 17 billion observed connections, all three profiles come back never seen, modern included. Narrowing the cipher list fixes the JA4 and does nothing for the rest of the hello, which stays distinctive for reasons OpenSSL gives no way to change: no GREASE anywhere, and extensions in OpenSSL's fixed order rather than a browser's shuffled one. The case for modern therefore rests on the JA4 going from t13d10512h2 to t13d1113h2, and on stopping the SSLv3 and TLS 1.0 offer, not on blending in. Blending in is not available on this TLS stack, and the profile docs now say so rather than claiming otherwise.
Phase 4. A hop that fails now retries with a different profile, per hop rather than per request, because a redirect can land on a differently protected host. Two axes with separate triggers, which is why this is a decision function and not a chain. A handshake failure that suggests a wider offer widens to compatibility; a refusal changes what we claim to be. Conflating them would answer a certificate rejection by offering weaker crypto. Which direction a refusal moves in is empirical, not a rule keyed on vendor. The evidence points both ways: Akamai, Kasada and Cloudflare deployments reward a browser claim, PerimeterX punishes it on every target tested, and two deployments of the same product routinely disagree. So try the other side once and let the result decide. Three things the tests and the wire caught. next_rung cycled. modern refused moves to chrome131, chrome131 refused moves back to modern, forever. The termination test was written before the function and failed immediately. Fixed by passing what has already been tried, so termination is a property of the function rather than of the cap at the call site. MAX_RUNGS stays as a belt to that braces. The ladder fired on any 403. Most 403s are ordinary authorization failures, not bot blocks, and retrying each one with different TLS settings doubles the request count of any scan that touches one. It now moves only on positive evidence that a product intervened, meaning a named vendor. The cost is missing a product we cannot fingerprint; the alternative cost is much larger and falls on every scan. And one such miss, found immediately: wizzair.com answers 405 with `x-amzn-waf-action: captcha`, and the classifier only knew `challenge`, so a genuine interception read as an ordinary refusal. AWS WAF is now a vendor the classifier knows and the header's presence is the signal rather than one of its values. A profile is remembered per host only after it has actually got through, so a scan does not re-walk the ladder on every request. Only on success: a failure could be the host having a bad minute, and recording those would let one flake pin a host to a worse profile for the rest of a run. A pinned config, meaning an explicit profile, cipher string or TLS version, never moves at all. On the wire, with the default still compatibility: akamai.com 403 -> 200 walks compatibility, modern, chrome131 wizzair.com 405 -> 302 after the AWS WAF fix priceline.com 200 unchanged; never claims a browser, so never zillow.com 200 walks into the PerimeterX trap
Phase 5, and the breaking change the rest was building toward.
A request that names no profile now sends `modern`: 11 cipher suites,
TLS 1.2 and 1.3. It used to send 105 suites and every version back to
SSLv3, to every host on the internet, so that the small number needing
that stayed reachable. That is now paid for only where it is wanted: a
server too old for `modern` is still reached, by the ladder widening to
`compatibility` on the second handshake.
Response carries what the ladder tried and what it amounts to:
attempts [(profile, outcome, status)] in order
protection (outcome, vendor) from the classifier
conclusion reached | blocked | needs_browser | needs_legacy_tls |
unreachable
needs_browser is the one worth acting on. It means a JavaScript
challenge, which no HTTP client passes, so retrying with different
settings cannot help and the URL should go to a real browser. Blocked is
the opposite: a different approach might work. A consumer should not have
to tell those apart by grepping our response body for cf-mitigated.
TransportError carries the same verdict plus kind, tls_failure and
retryable, so a failure that never produced a response is just as legible
as one that did.
Verified from the 1.0 wheel:
akamai.com 200 reached modern refused, chrome131 reached
priceline.com 200 reached modern reached
cf challenge 403 needs_browser cloudflare modern and chrome131 challenged
google.com 200 reached modern reached
Two things the default change turned up.
The test server accepted one connection then exited, which was fine while
a failed handshake ended the story. It does not any more: the ladder
retries, so every legacy-cipher test needs two connections, and with a
single-shot server the second gets ECONNREFUSED and the test fails for a
reason unrelated to what it tests. It now serves until shutdown.
And asking for SSLv3 without naming a profile failed outright with "no
ciphers available", because the default offers nothing that exists in
SSLv3, and naming a version pins the configuration, which is exactly what
stops the ladder rescuing it. A pre-TLS1.2 floor now selects
`compatibility` on its own.
The ten legacy_default.rs tests pass unchanged and now mean something
stronger: they assert RC4, 3DES, SEED, Camellia, anonymous DH and SSLv3
are reachable through the ladder rather than on the first hello.
fingerprint_baseline.rs is re-pinned, with compatibility's old values
kept under their own test so the transcription stays honest.
The Python API is what BBOT consumes and it had no pytest coverage at all: not profiles, not conclusion, not attempts, not protection, not TransportError. That mattered more than the usual reason, because the interesting half of this release is whether a *failure* is legible enough to act on, and that is exactly the path nothing exercised. Twelve tests against a local server, so they are deterministic and need no network. They cover which profile a request gets and what it puts on the wire, header precedence, the reached / blocked / needs_browser distinction, a bare 403 naming no vendor, and the attributes on a transport failure. Writing them found a real gap. validate_profile() was added in the profile work and never called, so an unknown name still resolved silently to the default: precisely the behaviour I had described as fixed. resolved_profile falls back on purpose, because it runs deep inside the connector where a failure has nowhere to go, so the check belongs at the edge next to validate_proxy. A typo is now a message rather than a profile the caller did not ask for. 174 pytest and 326 cargo tests.
Every request in an opening burst walked the ladder independently, because the per-host memory is only written once something succeeds and nothing had succeeded yet. Measured against akamai.com: twelve concurrent requests, twenty-four handshake attempts. The overshoot is bounded by concurrency rather than scan size, so a thousand-path scan at fifty concurrent pays about fifty extra handshakes, not a thousand. They all land in the same burst though, against a host that is by definition already watching, which is the worst possible moment to look like a pile of odd clients. One request now discovers and the rest wait for the answer, via a per-host semaphore with the usual double-check after acquiring: the leader often finishes while a waiter is queued, which is the whole point of having queued. Bounded on purpose. A waiter gives up after the request timeout and walks the ladder itself rather than inheriting a stalled leader; a wasted handshake is a far smaller problem than a request that never returns. The semaphore is dropped once a host's profile is known, so the map holds hosts under discovery rather than every host ever seen. The permit is released the moment the profile is known rather than at the end of the hop, and that detail is most of the win. Held to the end, waiters also sat through redirect handling and cookie bookkeeping they had no use for: 13 attempts but 4.0s. Released early, the same 13 attempts in 1.6s. before 24 attempts 1.4s after, released late 13 attempts 4.0s after, released early 13 attempts 1.6s floor (one attempt each) 12 attempts Half the handshakes for 0.2s, and nothing at all on a host that does not need the ladder: google.com is 12 attempts in 1.0s either way.
Reverts 614a3c3, which added a per-host semaphore so that in an opening burst one request walks the ladder and the rest wait for its answer. It worked. Twelve concurrent requests to akamai.com went from 24 handshake attempts to 13, at 1.6s against 1.4s. But a waiter can only be released once the leader has classified its response, and classification needs the body, so the wait is a whole request long. That is a head-of-line stall, and not blocking a fast request behind a slow one is a property request_batch_stream exists to provide and has a test guarding it. The fix broke that test. Bounding the wait to 250ms put the test back and removed the entire saving, 13 attempts back to 24, because no bound short enough to keep the latency property is long enough for the leader to have learned anything. The two goals are in direct conflict and latency wins: the extra handshakes are bounded by concurrency, happen only on the first burst against a host that needs a shift, and cost nobody a response. The test the fix came with is kept, rewritten to record the overshoot and what was tried, so the cost stays visible instead of being rediscovered later.
Three places where the 1.0 work stopped short of the paths that do not go through the pooled client. raw_connect, and requests using resolve_ip or request_target, now widen their cipher offer and retry when a handshake fails in a way that says the peer could not negotiate. They bypass the pooled client, so they bypassed the ladder too, and when the default narrowed from 105 suites to 11 they quietly lost the legacy reach the custom OpenSSL build exists to give them. A virtualhost sweep is exactly where an old appliance turns up and exactly where it would have gone unseen. They stop at that rung. The other one reacts to a refusal by changing what the client claims to be, and that would be wrong here: these are the callers who asked for exact control over one request, so re-sending a probe dressed differently answers a question nobody asked. A raw connection has no response to classify at all. A named profile, cipher string or TLS version still pins everything, and there is a test for that, because widening past a pin would make a TLS enumeration sweep report every server as speaking everything. connect_stream returns the profile it settled on so dispatch_direct builds its request under the same one, rather than letting the headers drift away from the TLS underneath them. download() takes profile, cipher_string, min_tls_version, max_tls_version and redirect_cookies. It hardcoded no profile with no way to pass one, so a file behind a host that only answers a browser was unreachable through download while the same URL through request was fine. blasthttp.mock forwards every kwarg to the real client on a passthrough request instead of naming six and dropping the rest. Ignoring them is right when the URL is intercepted, since nothing is dialled, but a passed-through request is real and was being given different TLS, a different profile and no timeout from what the caller asked for. BBOT runs its local-target suite through this path. That last test then found a fourth thing: a request that named a profile was writing it into the per-host memory, although reads already skip the memory when the caller pinned something. So one deliberate profile="chrome" request converted the host for every later request, which is the worst direction for it to drift in, since claiming to be a browser is what invites a detector to check the claim. It also made the result depend on the order two unrelated requests ran in. Writes now follow the same rule as reads.
Second attempt at the burst cost, after 614a3c3 was backed out in d1d0180. The problem is unchanged: at the start of a batch the per-host memory is empty, so every concurrent request to one host separately finds the default refused and separately shifts. Six requests, twelve handshakes, where seven would do. The first attempt made the others wait on a per-host lock. It cut the handshakes and was slower anyway, because a waiter cannot be released until the leader has classified its response, classification needs the body, and so the wait is a whole request long. It also broke the rule that a slow request never holds up faster ones, which is the entire promise of send_batch_stream. This does not wait. A request whose host is already being discovered gives up its slot, lets a request for a different host have it, and goes once there is an answer. Nothing idles, so the saving costs no time. The condition is the whole design and it comes from the test that killed the lock: six requests, one host, concurrency six. Step aside there and you are standing in the street, because every other request wants the same host. So a request only steps aside when there is undispatched work for a different host. When there is not, it goes and discovers the host again, which is the cheaper mistake. That makes this strictly better or equal to having no scheduler, never worse, which is what the lock failed to be. Measured against www.akamai.com, six concurrent requests: alone 12 handshakes (unchanged, correctly) mixed with another host 7 handshakes Seven is the floor: one request discovers, five are told. Details worth keeping: The decision and the commit happen under one lock. Deciding and then committing separately lets two requests both look, both see nobody discovering, and both go. The gate is a watch rather than a Notify, because the answer can land between a request reading the state and awaiting on it, and a watch carries the value so that request sees it instead of sleeping until its timeout. It sends with send_replace, since send refuses when nothing is subscribed and leaves the value alone, which is the common case: the leader usually finishes before anyone has stepped aside. A unit test caught that one. The leader's guard marks the host answered on drop rather than on success, so a request that fails, times out or panics releases whoever stepped aside for it. There is a ceiling on how many requests may be standing aside at once, set to the concurrency limit. A request standing aside holds no permit, so the stream driver spawns another in its place; without a ceiling a batch alternating between two slow hosts would spawn task after task, each handing its permit straight back, until the whole batch was resident. send_batch steps aside before taking a permit. The stream driver hands the permit over with the request, so there it is given back for the wait and retaken after.
Load testing the scheduler found it doing nothing at all on any batch
larger than its concurrency limit, which is every batch worth
scheduling.
The ceiling on how many requests may stand aside at once was set to the
concurrency limit for both callers. That is right for the stream
driver, which takes a permit before it spawns: a request that stands
aside hands the permit straight back, the driver spawns another, and
without a bound a batch alternating between two slow hosts would spawn
the lot. It is wrong for send_batch, which spawns every task up front
regardless. There standing aside creates nothing, so the ceiling
protects nothing, and the requests it turns away go and duplicate the
discovery instead.
Measured over a grid of batch shapes, before and after, counting
handshakes above the floor of one per request plus one per host:
hosts reqs conc ceiling lifted
2 100 100 98 0
2 100 50 63 0
2 400 50 62 0
2 400 200 199 0
4 200 200 196 0
10 500 100 139 0
20 2000 50 62 0
Wall time is unchanged in five of those seven. The two that move are
the ones a single wave deep, where the batch is no bigger than its own
concurrency limit: there the discovery cannot overlap with anything
else and the batch pays one extra round trip, 0.22s to 0.32s against a
0.1s server. Multi-wave batches absorb it completely.
Also corrects something I had wrong in the earlier commit message. The
saving is not free. A request standing aside frees its slot, but the
condition only asks whether work for another host exists, not whether
there is enough of it to fill every slot being handed back. When there
is not, the slots idle and the batch pays that round trip. It never
costs time without saving anything, and it never reorders results, but
it is a trade rather than a freebie.
A soak of 10,000 requests alternating both paths: nothing lost, no
hangs, descriptors flat, per-round time flat, memory identical to the
build without the scheduler. The first round now costs one extra
handshake per host rather than a hundred.
One measurement note for anyone repeating this. The first attempt
showed a 3.5x slowdown that was entirely the test harness: the change
bunches connections tightly enough to overflow a listening socket's
default backlog of 128, and the SYN retransmit reads as a stall in the
client. With a real backlog the same case is 0.45s against 0.43s.
# Conflicts: # Cargo.toml # src/client/hyper.rs # tests/tls_server.rs
The classifier reads a 64 KiB window from the top of the body, and sliced a str straight at that byte. Slicing at a byte that is not a character boundary panics rather than erring, and max_body_size cuts the body at a byte count, so a capped read can leave a half-written character sitting exactly there. It took the whole request down with it. This only became reachable once dev's max_body_size work met the classifier, which is why neither branch saw it alone. Walk back to the nearest boundary before slicing. Also reformat test_mock.py, which ruff 0.15.10 (what CI pins) wanted.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
0.10 sent one connection shape at every host: 105 cipher suites and every protocol version back to SSLv3. That reaches almost anything, and it is also unlike any browser in existence. A default that nothing else on the internet sends is a fingerprint in itself.
1.0 replaces that with three profiles.
compatibilityis a faithful transcription of the 0.10 default, kept byte-identical and held to it by the existing testsmodernis the new default: 11 suites, TLS 1.2 or 1.3chrome131imitates Chrome's TLS, HTTP/2 settings and headers together, because imitating one of the three and not the others is worse than imitating noneBreaking: a request naming no profile sends
modern. A server too old for that is still reached, by the ladder falling back tocompatibilityon the second handshake.Anti-bot handling is the main focus
This is the part the release is really about. Reaching a server was never the hard problem. Being allowed to stay is.
The client now notices what happened to a request and changes shape in response. A handshake failure suggesting the peer could not negotiate widens the offer. A refusal from a recognised protection product changes what the client claims to be. This happens per redirect hop, since a redirect can land on a differently protected host, and it does not spend the caller's retry budget, which is a separate question with a separate answer.
The ladder only moves on positive evidence that a product intervened. That restraint matters more than it sounds: most 403s are ordinary authorization failures, and retrying every one of them with different TLS would double the request count of any scan that touches one. A bare 403 from nginx is not evidence of anything.
What it found is reported rather than hidden:
Response.protectiongives(outcome, vendor), where outcome is ok, present, challenge, blocked or errorResponse.conclusionreduces that to one verdict: reached, blocked, needs_browser, needs_legacy_tls or unreachable.needs_browsermeans a JavaScript challenge, which no HTTP client passes, so that URL should go to a real browser instead of being retried foreverResponse.attemptslists what the ladder tried, in orderA profile that gets through is remembered per host, so a scan does not re-walk the ladder on every request. Naming a
profile,cipher_string,min_tls_versionormax_tls_versionpins the configuration and the ladder stays put.BLASTHTTP_BISECT=tls,headers,http2turns the layers off independently, for working out which one a detector is actually reacting to.TLS failures now say why
The reason used to be destroyed in transit. hyper-util's error displays as the literal "client error (Connect)" for every connect-time failure, so a cipher mismatch, a rejected certificate, a DNS failure and a refused connection were indistinguishable. Failures now raise
TransportErrorcarryingkind,tls_failure,retryable,conclusionandattempts. It subclassesRuntimeError, so existing handlers keep working.That also fixes a consequence: TLS failures were being retried, because they were misclassified as connection errors, which contradicts the rule saying they should not be.
Batch behavior
Batches no longer have every request in an opening burst work out the same host's profile separately. A request whose host is already being discovered stands aside and lets a request for a different host take its slot, then goes once there is an answer.
Standing aside only happens when there is undispatched work elsewhere, so a batch that is all one host sends everything immediately and discovers the profile more than once, which is the cheaper mistake. A slow request never holds up faster ones behind it. Across a grid of batch shapes the duplicate handshakes go to zero in every one: 2000 URLs over 20 hosts drops 62 handshakes, 500 URLs over 10 hosts drops 139, and wall time is unchanged.
Smaller things
raw_connect, and requests usingresolve_iporrequest_target, widen their cipher offer and retry on a handshake the peer could not negotiate. These bypass the pooled client, so they bypassed the ladder too, and when the default narrowed from 105 suites to 11 they quietly lost the legacy reach the custom OpenSSL build exists to provide. They stop at that rung, since the other one reacts to a refusal by changing what the client claims to be, and these are the callers who asked for exact controldownload()takesprofile,cipher_string,min_tls_version,max_tls_versionandredirect_cookies. It hardcoded no profile with no way to pass one, so a file behind a host that only answers a browser was unreachable throughdownloadwhile the same URL throughrequestwas fineblasthttp.mockforwards every kwarg to the real client on a passthrough. It named six and dropped the rest, so a URL excluded from mocking was dialled with different TLS, a different profile and no timeout from what the caller asked forprofile="chrome"call turned every later request to that host into a browser claim and made the result depend on the order two unrelated requests happened to run incompatibility, rather than failing with "no ciphers available" because the default offers nothing that exists in SSLv3Synced with dev
This branch is merged up to dev as of 9e682db, which brings the proxy tunnel TLS fix, the HTTP/2 empty path fix, the cookie store fixes, the body decode work and the dependabot bumps.
Two conflicts needed judgment rather than a pick. dev collapsed the three pooled client variants into one when it moved the target handshake inside the proxy tunnel, so the dispatch match had nothing left to match on; dev's single client is kept along with 1.0's error classifier, which reads back why a handshake failed instead of guessing from the error text. And 1.0 had moved the TLS test server under
tests/support/, so dev's new h2 acceptor was ported there andh2_path.rsandproxy_tls.rswere repointed at it.The sync also turned up a crash that neither branch could hit alone. The antibot classifier scans a 64 KiB window from the top of the body and sliced a
strstraight at that byte. Slicing at a byte that is not a character boundary panics rather than erring, andmax_body_sizecuts bodies at a byte count, so a capped read could leave a half-written character sitting exactly there and take the whole request down. Fixed by walking back to the nearest boundary, with a regression test that reproduces the original panic.Before merging
The
compatibilityprofile iscipher_list: "ALL", which includes the anonymous suites. Those carry no certificate, and OpenSSL ignores verify-peer when no certificate arrives, so a caller with verification on who lands oncompatibility, whether by naming it or by the ladder falling back to it, gets a handshake that reports success while checking nothing. #119 fixes this for 0.10 and the fix needs porting into the profile here, either by landing #119 in dev first and syncing again or by carrying it straight over.