Skip to content

BlastHTTP 1.0 - Connection Profile - #128

Open
liquidsec wants to merge 22 commits into
devfrom
1.0-connection-profiles
Open

liquidsec wants to merge 22 commits into
devfrom
1.0-connection-profiles

Conversation

@liquidsec

@liquidsec liquidsec commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

What changes

0.10 sent one connection shape at every host: 105 cipher suites and every protocol version back to SSLv3. That reaches almost anything, and it is also unlike any browser in existence. A default that nothing else on the internet sends is a fingerprint in itself.

1.0 replaces that with three profiles.

  • compatibility is a faithful transcription of the 0.10 default, kept byte-identical and held to it by the existing tests
  • modern is the new default: 11 suites, TLS 1.2 or 1.3
  • chrome131 imitates Chrome's TLS, HTTP/2 settings and headers together, because imitating one of the three and not the others is worse than imitating none

Breaking: a request naming no profile sends modern. A server too old for that is still reached, by the ladder falling back to compatibility on the second handshake.

Anti-bot handling is the main focus

This is the part the release is really about. Reaching a server was never the hard problem. Being allowed to stay is.

The client now notices what happened to a request and changes shape in response. A handshake failure suggesting the peer could not negotiate widens the offer. A refusal from a recognised protection product changes what the client claims to be. This happens per redirect hop, since a redirect can land on a differently protected host, and it does not spend the caller's retry budget, which is a separate question with a separate answer.

The ladder only moves on positive evidence that a product intervened. That restraint matters more than it sounds: most 403s are ordinary authorization failures, and retrying every one of them with different TLS would double the request count of any scan that touches one. A bare 403 from nginx is not evidence of anything.

What it found is reported rather than hidden:

  • Response.protection gives (outcome, vendor), where outcome is ok, present, challenge, blocked or error
  • Response.conclusion reduces that to one verdict: reached, blocked, needs_browser, needs_legacy_tls or unreachable. needs_browser means a JavaScript challenge, which no HTTP client passes, so that URL should go to a real browser instead of being retried forever
  • Response.attempts lists what the ladder tried, in order

A profile that gets through is remembered per host, so a scan does not re-walk the ladder on every request. Naming a profile, cipher_string, min_tls_version or max_tls_version pins the configuration and the ladder stays put.

BLASTHTTP_BISECT=tls,headers,http2 turns the layers off independently, for working out which one a detector is actually reacting to.

TLS failures now say why

The reason used to be destroyed in transit. hyper-util's error displays as the literal "client error (Connect)" for every connect-time failure, so a cipher mismatch, a rejected certificate, a DNS failure and a refused connection were indistinguishable. Failures now raise TransportError carrying kind, tls_failure, retryable, conclusion and attempts. It subclasses RuntimeError, so existing handlers keep working.

That also fixes a consequence: TLS failures were being retried, because they were misclassified as connection errors, which contradicts the rule saying they should not be.

Batch behavior

Batches no longer have every request in an opening burst work out the same host's profile separately. A request whose host is already being discovered stands aside and lets a request for a different host take its slot, then goes once there is an answer.

Standing aside only happens when there is undispatched work elsewhere, so a batch that is all one host sends everything immediately and discovers the profile more than once, which is the cheaper mistake. A slow request never holds up faster ones behind it. Across a grid of batch shapes the duplicate handshakes go to zero in every one: 2000 URLs over 20 hosts drops 62 handshakes, 500 URLs over 10 hosts drops 139, and wall time is unchanged.

Smaller things

  • raw_connect, and requests using resolve_ip or request_target, widen their cipher offer and retry on a handshake the peer could not negotiate. These bypass the pooled client, so they bypassed the ladder too, and when the default narrowed from 105 suites to 11 they quietly lost the legacy reach the custom OpenSSL build exists to provide. They stop at that rung, since the other one reacts to a refusal by changing what the client claims to be, and these are the callers who asked for exact control
  • download() takes profile, cipher_string, min_tls_version, max_tls_version and redirect_cookies. It hardcoded no profile with no way to pass one, so a file behind a host that only answers a browser was unreachable through download while the same URL through request was fine
  • blasthttp.mock forwards every kwarg to the real client on a passthrough. It named six and dropped the rest, so a URL excluded from mocking was dialled with different TLS, a different profile and no timeout from what the caller asked for
  • A request that names a profile no longer writes it into the per-host memory. Reads already skipped the memory when the caller pinned something, writes did not, so one deliberate profile="chrome" call turned every later request to that host into a browser claim and made the result depend on the order two unrelated requests happened to run in
  • Asking for a pre-TLS1.2 version without naming a profile selects compatibility, rather than failing with "no ciphers available" because the default offers nothing that exists in SSLv3
  • An unknown profile name is an error instead of a silent fallback

Synced with dev

This branch is merged up to dev as of 9e682db, which brings the proxy tunnel TLS fix, the HTTP/2 empty path fix, the cookie store fixes, the body decode work and the dependabot bumps.

Two conflicts needed judgment rather than a pick. dev collapsed the three pooled client variants into one when it moved the target handshake inside the proxy tunnel, so the dispatch match had nothing left to match on; dev's single client is kept along with 1.0's error classifier, which reads back why a handshake failed instead of guessing from the error text. And 1.0 had moved the TLS test server under tests/support/, so dev's new h2 acceptor was ported there and h2_path.rs and proxy_tls.rs were repointed at it.

The sync also turned up a crash that neither branch could hit alone. The antibot classifier scans a 64 KiB window from the top of the body and sliced a str straight at that byte. Slicing at a byte that is not a character boundary panics rather than erring, and max_body_size cuts bodies at a byte count, so a capped read could leave a half-written character sitting exactly there and take the whole request down. Fixed by walking back to the nearest boundary, with a regression test that reproduces the original panic.

Before merging

The compatibility profile is cipher_list: "ALL", which includes the anonymous suites. Those carry no certificate, and OpenSSL ignores verify-peer when no certificate arrives, so a caller with verification on who lands on compatibility, whether by naming it or by the ladder falling back to it, gets a handshake that reports success while checking nothing. #119 fixes this for 0.10 and the fix needs porting into the profile here, either by landing #119 in dev first and syncing again or by carrying it straight over.

The custom OpenSSL build exists so blasthttp can reach servers that only
speak deprecated ciphers, but nothing was reaching them. `SslConnector::
builder` installs its own cipher list, `DEFAULT:!aNULL:!eNULL:!MD5:!3DES:
!DES:!RC4:!IDEA:!SEED:...`, and we never replaced it, so RC4, DES, 3DES
and SEED never made it into the ClientHello. A server speaking only one
of them was unreachable unless the caller passed `cipher_string` by hand.

`set_security_level(0)` looked like it covered this and does not. The
security level decides how weak a negotiated cipher may be; the cipher
list decides which ones are offered at all.

Set `ALL` when the caller names nothing, on both the pooled path and the
`connect_stream` path used by `raw_connect`, `resolve_ip` and
`request_target`, which build their SSL contexts separately.

Null-encryption suites stay out of the default. They remain available
through an explicit `cipher_string`, but negotiating one by accident
would return a connection that looks like TLS and encrypts nothing.

tests/legacy_default.rs covers this end to end: each test stands up a
real TLS server pinned to one legacy cipher or protocol version and
connects with a client that sets no TLS options at all. All five cipher
cases failed before this change.

Side effect worth knowing: the ClientHello grows from 31 cipher suites
to 105, so the client's JA3/JA4 fingerprint changes.

README dropped its claims about export ciphers and SSLv3. OpenSSL
removed export ciphers in 1.1.0, and SSLv3 is disabled in our build
despite `enable-ssl3` being passed, so neither has ever worked.
SSLv3 has never been reachable in any build, despite the README
advertising it and the build script asking for it. Three separate things
were in the way.

The build. `scripts/build-openssl.sh` passed `enable-ssl3`, but
OpenSSL's Configure carries a disable cascade that reads "if ssl3-method
is off, turn ssl3 off too", and ssl3-method is off by default. So the
flag was undone during configure. `configdata.pm` recorded both
`enable-ssl3` and `no-ssl3`, and every shipped build came out with
OPENSSL_NO_SSL3 defined and no SSLv3_method symbols in libssl.a. Passing
`enable-ssl3-method` alongside it fixes this.

The SSL options. `SslConnector::builder` sets NO_SSLV3 and nothing
cleared it, so even a build with SSLv3 compiled in would refuse to
negotiate it. Cleared on both the pooled path and `connect_stream`.

The spelling. `parse_tls_version` understood 1.0 through 1.3 and nothing
else, so there was no way to name SSLv3 at all. It now takes `3.0`,
`ssl3` and `sslv3`, and the error message mentions it.

An SSLv3-only server is now reachable without asking, matching how the
legacy ciphers behave after the previous commit, and can also be pinned
explicitly. The test server clears NO_SSLV3 too so a test can pin it.

Separately, the build cache was keyed only on the target, so changing
what the build asks for was silently ignored and a stale install stayed
in place. It now keys on the version and feature flags as well, which is
what makes this fix reach anyone with an existing checkout. Upgrading
triggers one rebuild.

Measured: JA4 is unchanged, `supported_versions` gains 0300, and google,
github, cloudflare, amazon, reddit, wikipedia, microsoft and apple all
still return 200.
First piece of the 1.x benchmark work. Nothing in the impersonation plan
is verifiable without a way to see what we actually put on the wire, and
doing that against public fingerprinting services does not scale: they
rate-limit, several that the ecosystem still cites are dead, and one has
been hijacked and now serves a gambling site.

So measure it locally. The test server installs a client_hello callback,
which fires after the hello is parsed and before a cipher is chosen, and
records the offer. That is the only point where the full ClientHello is
still visible. The safe openssl wrapper covers ciphers and versions but
has no accessor for the extension list, which is most of a fingerprint,
so that part goes through openssl-sys. Both new dev-dependencies were
already in the tree as transitive dependencies of openssl.

JA4 rather than JA3, for a reason that decides the shape of the stealth
work: JA4 sorts its lists and strips GREASE, and extension order plus
GREASE are exactly the two things OpenSSL cannot control. They cost us
nothing on this metric. JA3 is also a poor target in its own right, since
Chrome permutes extensions per connection and one Chrome build yields a
different JA3 on every handshake.

The implementation is checked against ground truth rather than against
itself: ja4_conformance.rs computes the JA4 of a captured Chrome 150
hello from lexiforest/curl-impersonate's signature corpus (MIT) and
asserts the published value, which four independent projects agree on
across three fingerprinting services. All three segments match.

fingerprint_baseline.rs pins where we stand: t13i9911h2 locally, against
Chrome's t13d1516h2. It also records the gaps as assertions that will
flip as each is closed, so no GREASE anywhere, none of the six
browser-only extensions, and every protocol version from SSLv3 up on
offer.

Shared test helpers move to tests/support/. Cargo compiles every
top-level file in tests/ as its own test binary, so tls_server.rs was
already building as a target with no tests in it, and a second helper
next to it failed outright once it referenced a sibling.

One thing worth knowing that fell out of this: JA4's count fields are
two digits, so our 105 cipher suites render as 99. Past that point JA4
cannot tell us apart from any other client with an unreasonable cipher
list.
Second piece of the 1.x benchmark work. The capture side answers what we
send; this answers what happens to us.

A scan has to tell three things apart that all arrive looking like a
response: the origin answered, a protection product answered instead, or
the product answered and handed us a session anyway. Collapsing those is
what turns a protected host into an apparently dead one, which is the
false negative the whole effort is about.

Classification keys on headers, cookies and status, never on a vendor's
name in the body. A first pass did the latter and scored datadome.co,
imperva.com and kasada.io as challenges from their own marketing copy;
that case is now a test.

Two things only showed up by running it against the real thing, and both
came from a fixture I had assumed rather than captured:

  - www.akamai.com's 403 carries no `server` header at all. It
    identifies itself with akamai-grn, x-akam-sw-version and an ak_p
    server-timing entry, so requiring `server: AkamaiGHost` matched
    nothing.
  - that page entity-encodes its own punctuation, so the body holds
    `errors.edgesuite.net` and a search for the dotted hostname
    never fires. Matching the bare label survives it.

The live benchmark is opt-in (`--ignored`) and prints a scorecard rather
than asserting, since "we are not blocked" would fail for reasons
outside this repository. It also cross-checks our local JA4 against a
remote observer, which surfaced a known divergence worth recording: the
JA4 header is fixed width and the spec caps counts at 99, our legacy
list is 105 suites, and tls.peet.ws prints 105 where we print 99. Both
hash segments agree exactly, so only the rendering differs. A stealth
profile at 15 suites removes it.

Current scorecard: 1/4 through. cloudflare.com present, akamai.com
blocked, both scrapingcourse challenge pages unreachable as expected
since that tier wants JavaScript.
The scorecard reported "1/4 got through" when the real bypass count was
zero. Two ways that number lied.

The one success was cloudflare.com, which is a control. Its homepage
challenges nobody, so every client reaches it and passing says nothing
about us. Counting it made a 0% bypass rate read as 25%.

It also put the JavaScript tier in the denominator. Cloudflare's managed
challenge and the scrapingcourse pages want the browser to execute their
code and report back, which no HTTP client does, curl_cffi included. So
the score could never approach 4/4 however good the fingerprint work
got, and a permanently unreachable denominator makes the metric useless
for tracking progress.

Now targets carry a tier. Controls are verified and not scored, and a
control that fails aborts the run, since a blocked control means the
network path or the site changed and nothing else in the run can be
trusted. The JavaScript tier is reported so a change would be visible
but never counted against us. Only the passive-fingerprint tier is
scored, because it is the only one this work can move.

Reads 0/1 on the passive tier today. That is the number stealth mode has
to change, and curl_cffi already scores 1/1 on it.
Starts closing the gap the benchmark measures. A profile is data rather
than code, because browsers ship every few weeks and anything needing a
rewrite per release will not be kept current.

Chrome 131 rather than something newer, deliberately. Chrome 133 and
later sign with ML-DSA, codepoints OpenSSL 3.3.2 does not know, so their
signature_algorithms list cannot be reproduced and their JA4 is out of
reach. The cipher and extension sets are identical from Chrome 124
through 150, so only the signature algorithms are dated.

Both TLS builders now go through one `apply_tls_settings`, which also
closes a standing hazard: they were copy-pasted siblings, so anything
applied to only one silently gave raw_connect, resolve_ip and
request_target requests a different fingerprint from ordinary requests.

Where this leaves the JA4, against Chrome 131's
t13d1516h2_8daaf6152771_02713d6af862:

  was     t13d9911h2   105 ciphers, 11 extensions
  now     t13d1614h2    16 ciphers, 14 extensions

Signature algorithms match exactly. Every cipher matches except one, and
four extensions are now sent that were not: 0005, 0012, 4469 and fe0d.

Three gaps remain, and all three are things OpenSSL keeps for itself.
Probed rather than assumed: it refuses a custom extension for 0x0005,
0x001b and 0xff01, and accepts 0x0012, 0x4469 and 0xfe0d. status_request
was solved without a patch by setting it per-connection, which is the
only place OpenSSL's own API for it works. The other two need build
changes, along with suppressing the TLS_EMPTY_RENEGOTIATION_INFO_SCSV
that inflates the cipher count by one; suppressing that and sending
renegotiation_info are the same change, since OpenSSL sends the SCSV
precisely when it is not sending the extension.

Two traps found by measurement, both worth knowing:

The `openssl` crate stores a custom extension's payload in ex_data keyed
on the payload TYPE, so every extension returning Vec<u8> shares one
slot and overwrites the others. Three registered that way send one.
Distinct newtypes give each its own slot.

The harness was under-reporting. SSL_client_hello_get1_extensions_present
only returns extensions the server's OpenSSL has a definition for, so
ALPS and ECH were invisible to it, which is to say it was blind to
exactly the extensions being added. Capture now parses the ClientHello
off the wire, which also preserves extension order and GREASE for later.
Local and remote readings now agree exactly.
akamai.com now answers 200 with a real session where it returned a hard
403 before. That is the measurement this work exists to move: the
passive-fingerprint tier goes 0/1 to 1/1.

The OpenSSL patch does most of the TLS work. Stock OpenSSL signals
renegotiation-info support with the SCSV pseudo-cipher where every
browser uses the extension, which costs a match twice over: a cipher
nothing else sends, and a missing extension everything else sends.
Switching it fixes both, and the cipher segment of our JA4 now equals
Chrome 131's exactly.

The patch is conditioned on the TLS floor rather than applied
unconditionally, because the SCSV exists to protect pre-RFC5746 servers
that mishandle unfamiliar extensions. A client whose floor is TLS 1.2 is
modern and sends the extension; one willing to speak older protocols
keeps stock behaviour. The compatibility path is unchanged, and a test
asserts it.

Build patches are now part of the cached build recipe, so editing one
invalidates the build. Without that a patch change would silently do
nothing, which is how the SSLv3 flag went unnoticed for months.

Three of Chrome's extensions are deliberately NOT sent, each tried and
each withdrawn for breaking real sites:

  ALPS (0x4469): google.com negotiates it, then expects the settings
  exchange inside HTTP/2 that we do not implement. TLS completes and the
  request then fails.

  ECH (0xfe0d): a hand-built GREASE blob is not shaped precisely enough.
  reddit.com rejects it at the handshake. Registering the server-reply
  contexts fixed cloudflare.com but not reddit, so the payload is the
  problem rather than the registration.

  compress_certificate (0x001b): our build is OPENSSL_NO_COMP_ALG and
  cannot decompress a certificate. Enabling one means a compression
  library across six cross-compiled wheel targets.

That last point generalises, and is the finding worth carrying forward:
advertising a protocol feature is a promise to implement it, and a
server that takes you up on it gets a client that cannot follow through.

So the JA4 does not match Chrome exactly, and did not need to. Thirteen
extensions against Chrome's sixteen was enough to turn Akamai around.
Exact parity was never the bar; plausibility was, and a client that
connects beats one that matches a hash and cannot.

Verified on akamai, cloudflare, google, github, wikipedia, reddit and
microsoft: all reachable with the profile.
src/client/hyper.rs made no http2_* calls at all, so every SETTINGS value
and the connection window were hyper-util defaults, which match no
browser. Four of them are reachable through hyper-util and now follow the
profile:

  before  2:0;4:2097152;5:16384;6:16384|5177345|0|m,s,a,p
  after   2:0;4:6291456;6:262144|15663105|0|m,s,a,p
  chrome  1:65536;2:0;4:6291456;6:262144|15663105|0|m,a,s,p

INITIAL_WINDOW_SIZE, MAX_HEADER_LIST_SIZE and the WINDOW_UPDATE now match
Chrome exactly. Omitting MAX_FRAME_SIZE is part of the match rather than
an oversight: our default announced 16384 and Chrome announces nothing.
Adaptive window is off because it resizes the connection window as
traffic flows, emitting WINDOW_UPDATE frames no browser sends.

Two fields are left alone, both fork-gated rather than forgotten.
HEADER_TABLE_SIZE exists in hyper's own config but hyper-util does not
expose it. The pseudo-header order is hardcoded in h2's Iter::next(), and
reaching it means forking h2, hyper and hyper-util, since each layer
translates a fixed option set rather than passing options through. Worth
noting that curl_cffi gets Chrome's order for free: it is nghttp2's
natural order, and Rust's h2 simply chose differently.

Verified reachable with the profile: akamai, cloudflare, google, github,
reddit.
"We are blocked" does not say which part of the imitation failed, and
rebuilding between experiments makes finding out slow enough that you
guess instead. BLASTHTTP_BISECT=tls,headers,http2 turns off any
combination of the three layers at runtime.

It paid for itself immediately. PerimeterX blocks blasthttp+chrome while
letting both plain blasthttp and curl_cffi through, which I had assumed
meant our TLS fingerprint was not good enough. It is not the TLS.

On priceline.com, three passes each:

  plain (all off)            3/3
  tls only                   3/3
  headers only               0/3
  http2 only                 3/3
  tls + http2, no headers    3/3
  full profile               0/3

Narrowing into the headers, it is the User-Agent on its own. Adding only
sec-ch-ua, only sec-fetch-*, or only accept/accept-language/priority to
an otherwise plain client changes nothing; adding only a browser
User-Agent drops it to zero.

And it is any browser, not Chrome specifically:

  blasthttp/0.10.1 (default)   3/3
  Chrome 131 Windows           0/3
  Chrome 150 macOS             0/3   (byte-identical to curl_cffi's)
  Firefox 147                  0/3
  NotARealBrowser/1.0          3/3

curl_cffi sends that same Chrome 150 macOS string and passes 3/3. So the
User-Agent is not what is scored; it selects which scoring applies. Claim
to be a browser and the client gets checked against that claim. Claim
anything else, including nonsense, and it does not.

That model explains results that looked contradictory. Akamai and
DataDome reward the profile because they fingerprint passively and ours
is close enough. PerimeterX verifies the claim, curl_cffi survives the
verification and we do not, so against it our partial imitation is worse
than being honest. It also generalises the udemy.com case from the
research: not "browser headers over script TLS" but "any browser claim
invites verification".

Two consequences. A browser profile should not become a global default,
since it makes us worse against anything that verifies. And the case for
closing the remaining gaps, GREASE, extension order, ALPS, ECH,
post-quantum key share, is now specific rather than aesthetic: they are
what the verification is looking at.
Phase 1 of the profile plan, and a live bug on its own. Nothing that
chooses a connection profile can work until a failure says what went
wrong, and until now none of them did.

The reason was being destroyed on the way out. The connector hands
hyper-util a Box<dyn Error>; hyper-util wraps it in an error whose
Display is the literal string "client error (Connect)" for every
connect-time failure; nothing ever calls .source() to get back down to
the openssl error underneath. So the substring test that was supposed to
tell TLS failures apart could never match, and a cipher mismatch, a
rejected certificate, a DNS failure and a refused connection all arrived
identically as ErrorKind::Connection.

That also meant TLS failures were being retried, since Connection is
retryable and Tls is not, contradicting test_tls_error_is_not_retryable.

Now the connector records the reason before boxing it away, into a map
keyed by host and port. Keyed per host because a cached client serves
many hosts at once, and CertSlot already shows what a single shared slot
does under concurrency. The map lives on HyperClient rather than on
CachedClient: a retry with different TLS settings builds a *different*
cached client and still needs to read why the previous attempt failed,
so hanging it off the cached client would partition the map by exactly
the thing it exists to inform.

TlsFailure is derived from OpenSSL reason codes rather than message
text, and distinguishes the cases that lead somewhere different:
NoSharedCipher and UnsupportedProtocol mean a wider offer might work,
CertificateVerify means it never will, Reset and Timeout mean we learned
nothing about our ciphers.

Alert counts as "try wider" too, and that is the case that matters.
Running it against an RC4-only server offered AES returns
`ssl/tls alert handshake failure`, not `no shared cipher`: the latter is
what OpenSSL raises when *we* work it out locally, while a real server
just sends a fatal alert and declines to explain. Excluding alerts would
have meant never widening the offer in the single most common case. The
first version of this did exclude them; three test servers found it.

The stack is scanned for a code we recognise rather than read at index
zero, because OpenSSL pushes several errors per failure and a generic
"handshake failure" often sits on top of the one that explains it.

Python gets a TransportError (subclassing RuntimeError, so existing
`except RuntimeError` still works) carrying `kind`, `tls_failure` and
`retryable` as attributes. Six call sites were doing
PyRuntimeError::new_err(e.message) and dropping everything else.

Measured end to end against local servers:

  cipher mismatch   TLS handshake failed: peer sent an alert: ssl/tls
                    alert handshake failure
  version mismatch  TLS handshake failed: no protocol version in common
  nothing listening request failed: client error (Connect)

The third is the one that matters as much as the other two: a refused
connection must not claim to be a TLS problem, or the ladder would widen
its cipher offer at a host that was never listening.
Phase 2. The classifier already existed under tests/, regression-tested
against real captured responses, and took exactly the shape of a
response. It was just unreachable from anything that ships.

Now on Response::protection(), computed rather than stored so it stays
right for a response built by hand, and exposed to Python as a
(outcome, vendor) pair.

The distinction worth having is challenge versus blocked. A blocked
request might succeed with different connection settings; a challenge
never will, because answering it means running the page's JavaScript.
That is the signal to hand a URL to a real browser instead of retrying,
and BBOT should not have to find it by grepping our response body for
cf-mitigated.

Verified from Python against the wheel:

  akamai.com                  403  ('blocked', 'akamai')
  scrapingcourse cf-challenge 403  ('challenge', 'cloudflare')
  google.com                  200  ('ok', '-')

And the error path from the same build, which Phase 1 made possible:

  refused connection  TransportError kind=connection tls_failure=None
                      retryable=True
  cipher mismatch     kind=tls tls_failure=alert retryable=False
  version mismatch    kind=tls tls_failure=unsupported_protocol
                      retryable=False

Together those are the whole contract a consumer needs: whether we got a
response, whether something intercepted it, whether retrying could help,
and whether a browser is the only way through.
Phase 3. BrowserProfile becomes ConnectionProfile, and there are now
three: compatibility, modern, chrome131.

The trap this had to avoid: several behaviours were gated on *whether a
profile existed* rather than on anything in it. The padding and
encrypt_then_mac flips, add_browser_extensions, the OCSP status_request
and http2_adaptive_window were all "if a profile is set". Naming today's
default behaviour as a profile would have silently switched every one of
them on for the configuration that is supposed to be unchanged. They are
now explicit fields: `browser_extensions` on the TLS half, `http2` as an
Option, headers as a possibly-empty list.

So compatibility is a transcription, not a cleanup, and every None in it
is load-bearing. tests/legacy_default.rs and test_compat_path_keeps_the_scsv
are its specification, and all 314 tests pass unchanged. Confirmed on the
wire too: no profile and --profile compatibility both produce
t13d10512h2_c86627d460ee_c50a3655fff1 with 105 suites, bit for bit.

resolved_profile() is now infallible. Naming no profile means the default
rather than "no profile", which collapses five Option branches, and an
unknown name is an error instead of silently resolving to whatever the
default happened to be.

modern is new: 11 suites instead of 105, TLS 1.2 and 1.3 only, no browser
claim.

One correction, because the measurement contradicted the argument I wrote
it on. modern was justified as "unremarkable in real traffic". It is not.
Checked against tlsfingerprint.io, which indexes around 17 billion
observed connections, all three profiles come back never seen, modern
included. Narrowing the cipher list fixes the JA4 and does nothing for the
rest of the hello, which stays distinctive for reasons OpenSSL gives no
way to change: no GREASE anywhere, and extensions in OpenSSL's fixed order
rather than a browser's shuffled one.

The case for modern therefore rests on the JA4 going from t13d10512h2 to
t13d1113h2, and on stopping the SSLv3 and TLS 1.0 offer, not on blending
in. Blending in is not available on this TLS stack, and the profile docs
now say so rather than claiming otherwise.
Phase 4. A hop that fails now retries with a different profile, per hop
rather than per request, because a redirect can land on a differently
protected host.

Two axes with separate triggers, which is why this is a decision function
and not a chain. A handshake failure that suggests a wider offer widens
to compatibility; a refusal changes what we claim to be. Conflating them
would answer a certificate rejection by offering weaker crypto.

Which direction a refusal moves in is empirical, not a rule keyed on
vendor. The evidence points both ways: Akamai, Kasada and Cloudflare
deployments reward a browser claim, PerimeterX punishes it on every
target tested, and two deployments of the same product routinely
disagree. So try the other side once and let the result decide.

Three things the tests and the wire caught.

next_rung cycled. modern refused moves to chrome131, chrome131 refused
moves back to modern, forever. The termination test was written before
the function and failed immediately. Fixed by passing what has already
been tried, so termination is a property of the function rather than of
the cap at the call site. MAX_RUNGS stays as a belt to that braces.

The ladder fired on any 403. Most 403s are ordinary authorization
failures, not bot blocks, and retrying each one with different TLS
settings doubles the request count of any scan that touches one. It now
moves only on positive evidence that a product intervened, meaning a
named vendor. The cost is missing a product we cannot fingerprint; the
alternative cost is much larger and falls on every scan.

And one such miss, found immediately: wizzair.com answers 405 with
`x-amzn-waf-action: captcha`, and the classifier only knew `challenge`,
so a genuine interception read as an ordinary refusal. AWS WAF is now a
vendor the classifier knows and the header's presence is the signal
rather than one of its values.

A profile is remembered per host only after it has actually got through,
so a scan does not re-walk the ladder on every request. Only on success:
a failure could be the host having a bad minute, and recording those
would let one flake pin a host to a worse profile for the rest of a run.
A pinned config, meaning an explicit profile, cipher string or TLS
version, never moves at all.

On the wire, with the default still compatibility:

  akamai.com    403 -> 200   walks compatibility, modern, chrome131
  wizzair.com   405 -> 302   after the AWS WAF fix
  priceline.com 200         unchanged; never claims a browser, so never
  zillow.com    200         walks into the PerimeterX trap
Phase 5, and the breaking change the rest was building toward.

A request that names no profile now sends `modern`: 11 cipher suites,
TLS 1.2 and 1.3. It used to send 105 suites and every version back to
SSLv3, to every host on the internet, so that the small number needing
that stayed reachable. That is now paid for only where it is wanted: a
server too old for `modern` is still reached, by the ladder widening to
`compatibility` on the second handshake.

Response carries what the ladder tried and what it amounts to:

  attempts    [(profile, outcome, status)] in order
  protection  (outcome, vendor) from the classifier
  conclusion  reached | blocked | needs_browser | needs_legacy_tls |
              unreachable

needs_browser is the one worth acting on. It means a JavaScript
challenge, which no HTTP client passes, so retrying with different
settings cannot help and the URL should go to a real browser. Blocked is
the opposite: a different approach might work. A consumer should not have
to tell those apart by grepping our response body for cf-mitigated.

TransportError carries the same verdict plus kind, tls_failure and
retryable, so a failure that never produced a response is just as legible
as one that did.

Verified from the 1.0 wheel:

  akamai.com    200  reached                   modern refused, chrome131 reached
  priceline.com 200  reached                   modern reached
  cf challenge  403  needs_browser cloudflare  modern and chrome131 challenged
  google.com    200  reached                   modern reached

Two things the default change turned up.

The test server accepted one connection then exited, which was fine while
a failed handshake ended the story. It does not any more: the ladder
retries, so every legacy-cipher test needs two connections, and with a
single-shot server the second gets ECONNREFUSED and the test fails for a
reason unrelated to what it tests. It now serves until shutdown.

And asking for SSLv3 without naming a profile failed outright with "no
ciphers available", because the default offers nothing that exists in
SSLv3, and naming a version pins the configuration, which is exactly what
stops the ladder rescuing it. A pre-TLS1.2 floor now selects
`compatibility` on its own.

The ten legacy_default.rs tests pass unchanged and now mean something
stronger: they assert RC4, 3DES, SEED, Camellia, anonymous DH and SSLv3
are reachable through the ladder rather than on the first hello.
fingerprint_baseline.rs is re-pinned, with compatibility's old values
kept under their own test so the transcription stays honest.
The Python API is what BBOT consumes and it had no pytest coverage at
all: not profiles, not conclusion, not attempts, not protection, not
TransportError. That mattered more than the usual reason, because the
interesting half of this release is whether a *failure* is legible
enough to act on, and that is exactly the path nothing exercised.

Twelve tests against a local server, so they are deterministic and need
no network. They cover which profile a request gets and what it puts on
the wire, header precedence, the reached / blocked / needs_browser
distinction, a bare 403 naming no vendor, and the attributes on a
transport failure.

Writing them found a real gap. validate_profile() was added in the
profile work and never called, so an unknown name still resolved
silently to the default: precisely the behaviour I had described as
fixed. resolved_profile falls back on purpose, because it runs deep
inside the connector where a failure has nowhere to go, so the check
belongs at the edge next to validate_proxy. A typo is now a message
rather than a profile the caller did not ask for.

174 pytest and 326 cargo tests.
Every request in an opening burst walked the ladder independently,
because the per-host memory is only written once something succeeds and
nothing had succeeded yet. Measured against akamai.com: twelve
concurrent requests, twenty-four handshake attempts.

The overshoot is bounded by concurrency rather than scan size, so a
thousand-path scan at fifty concurrent pays about fifty extra
handshakes, not a thousand. They all land in the same burst though,
against a host that is by definition already watching, which is the
worst possible moment to look like a pile of odd clients.

One request now discovers and the rest wait for the answer, via a
per-host semaphore with the usual double-check after acquiring: the
leader often finishes while a waiter is queued, which is the whole point
of having queued.

Bounded on purpose. A waiter gives up after the request timeout and
walks the ladder itself rather than inheriting a stalled leader; a
wasted handshake is a far smaller problem than a request that never
returns. The semaphore is dropped once a host's profile is known, so the
map holds hosts under discovery rather than every host ever seen.

The permit is released the moment the profile is known rather than at
the end of the hop, and that detail is most of the win. Held to the end,
waiters also sat through redirect handling and cookie bookkeeping they
had no use for: 13 attempts but 4.0s. Released early, the same 13
attempts in 1.6s.

  before                     24 attempts  1.4s
  after, released late       13 attempts  4.0s
  after, released early      13 attempts  1.6s
  floor (one attempt each)   12 attempts

Half the handshakes for 0.2s, and nothing at all on a host that does not
need the ladder: google.com is 12 attempts in 1.0s either way.
Reverts 614a3c3, which added a per-host semaphore so that in an opening
burst one request walks the ladder and the rest wait for its answer.

It worked. Twelve concurrent requests to akamai.com went from 24
handshake attempts to 13, at 1.6s against 1.4s.

But a waiter can only be released once the leader has classified its
response, and classification needs the body, so the wait is a whole
request long. That is a head-of-line stall, and not blocking a fast
request behind a slow one is a property request_batch_stream exists to
provide and has a test guarding it. The fix broke that test.

Bounding the wait to 250ms put the test back and removed the entire
saving, 13 attempts back to 24, because no bound short enough to keep
the latency property is long enough for the leader to have learned
anything. The two goals are in direct conflict and latency wins: the
extra handshakes are bounded by concurrency, happen only on the first
burst against a host that needs a shift, and cost nobody a response.

The test the fix came with is kept, rewritten to record the overshoot
and what was tried, so the cost stays visible instead of being
rediscovered later.
Three places where the 1.0 work stopped short of the paths that do not
go through the pooled client.

raw_connect, and requests using resolve_ip or request_target, now widen
their cipher offer and retry when a handshake fails in a way that says
the peer could not negotiate. They bypass the pooled client, so they
bypassed the ladder too, and when the default narrowed from 105 suites
to 11 they quietly lost the legacy reach the custom OpenSSL build
exists to give them. A virtualhost sweep is exactly where an old
appliance turns up and exactly where it would have gone unseen.

They stop at that rung. The other one reacts to a refusal by changing
what the client claims to be, and that would be wrong here: these are
the callers who asked for exact control over one request, so re-sending
a probe dressed differently answers a question nobody asked. A raw
connection has no response to classify at all. A named profile, cipher
string or TLS version still pins everything, and there is a test for
that, because widening past a pin would make a TLS enumeration sweep
report every server as speaking everything.

connect_stream returns the profile it settled on so dispatch_direct
builds its request under the same one, rather than letting the headers
drift away from the TLS underneath them.

download() takes profile, cipher_string, min_tls_version,
max_tls_version and redirect_cookies. It hardcoded no profile with no
way to pass one, so a file behind a host that only answers a browser
was unreachable through download while the same URL through request
was fine.

blasthttp.mock forwards every kwarg to the real client on a passthrough
request instead of naming six and dropping the rest. Ignoring them is
right when the URL is intercepted, since nothing is dialled, but a
passed-through request is real and was being given different TLS, a
different profile and no timeout from what the caller asked for. BBOT
runs its local-target suite through this path.

That last test then found a fourth thing: a request that named a
profile was writing it into the per-host memory, although reads already
skip the memory when the caller pinned something. So one deliberate
profile="chrome" request converted the host for every later request,
which is the worst direction for it to drift in, since claiming to be a
browser is what invites a detector to check the claim. It also made the
result depend on the order two unrelated requests ran in. Writes now
follow the same rule as reads.
Second attempt at the burst cost, after 614a3c3 was backed out in
d1d0180. The problem is unchanged: at the start of a batch the per-host
memory is empty, so every concurrent request to one host separately
finds the default refused and separately shifts. Six requests, twelve
handshakes, where seven would do.

The first attempt made the others wait on a per-host lock. It cut the
handshakes and was slower anyway, because a waiter cannot be released
until the leader has classified its response, classification needs the
body, and so the wait is a whole request long. It also broke the rule
that a slow request never holds up faster ones, which is the entire
promise of send_batch_stream.

This does not wait. A request whose host is already being discovered
gives up its slot, lets a request for a different host have it, and
goes once there is an answer. Nothing idles, so the saving costs no
time.

The condition is the whole design and it comes from the test that
killed the lock: six requests, one host, concurrency six. Step aside
there and you are standing in the street, because every other request
wants the same host. So a request only steps aside when there is
undispatched work for a different host. When there is not, it goes and
discovers the host again, which is the cheaper mistake. That makes this
strictly better or equal to having no scheduler, never worse, which is
what the lock failed to be.

Measured against www.akamai.com, six concurrent requests:

  alone                    12 handshakes   (unchanged, correctly)
  mixed with another host   7 handshakes

Seven is the floor: one request discovers, five are told.

Details worth keeping:

The decision and the commit happen under one lock. Deciding and then
committing separately lets two requests both look, both see nobody
discovering, and both go.

The gate is a watch rather than a Notify, because the answer can land
between a request reading the state and awaiting on it, and a watch
carries the value so that request sees it instead of sleeping until its
timeout. It sends with send_replace, since send refuses when nothing is
subscribed and leaves the value alone, which is the common case: the
leader usually finishes before anyone has stepped aside. A unit test
caught that one.

The leader's guard marks the host answered on drop rather than on
success, so a request that fails, times out or panics releases whoever
stepped aside for it.

There is a ceiling on how many requests may be standing aside at once,
set to the concurrency limit. A request standing aside holds no permit,
so the stream driver spawns another in its place; without a ceiling a
batch alternating between two slow hosts would spawn task after task,
each handing its permit straight back, until the whole batch was
resident.

send_batch steps aside before taking a permit. The stream driver hands
the permit over with the request, so there it is given back for the
wait and retaken after.
Load testing the scheduler found it doing nothing at all on any batch
larger than its concurrency limit, which is every batch worth
scheduling.

The ceiling on how many requests may stand aside at once was set to the
concurrency limit for both callers. That is right for the stream
driver, which takes a permit before it spawns: a request that stands
aside hands the permit straight back, the driver spawns another, and
without a bound a batch alternating between two slow hosts would spawn
the lot. It is wrong for send_batch, which spawns every task up front
regardless. There standing aside creates nothing, so the ceiling
protects nothing, and the requests it turns away go and duplicate the
discovery instead.

Measured over a grid of batch shapes, before and after, counting
handshakes above the floor of one per request plus one per host:

  hosts  reqs  conc    ceiling   lifted
      2   100   100         98        0
      2   100    50         63        0
      2   400    50         62        0
      2   400   200        199        0
      4   200   200        196        0
     10   500   100        139        0
     20  2000    50         62        0

Wall time is unchanged in five of those seven. The two that move are
the ones a single wave deep, where the batch is no bigger than its own
concurrency limit: there the discovery cannot overlap with anything
else and the batch pays one extra round trip, 0.22s to 0.32s against a
0.1s server. Multi-wave batches absorb it completely.

Also corrects something I had wrong in the earlier commit message. The
saving is not free. A request standing aside frees its slot, but the
condition only asks whether work for another host exists, not whether
there is enough of it to fill every slot being handed back. When there
is not, the slots idle and the batch pays that round trip. It never
costs time without saving anything, and it never reorders results, but
it is a trade rather than a freebie.

A soak of 10,000 requests alternating both paths: nothing lost, no
hangs, descriptors flat, per-round time flat, memory identical to the
build without the scheduler. The first round now costs one extra
handshake per host rather than a hundred.

One measurement note for anyone repeating this. The first attempt
showed a 3.5x slowdown that was entirely the test harness: the change
bunches connections tightly enough to overflow a listening socket's
default backlog of 128, and the SYN retransmit reads as a stall in the
client. With a real backlog the same case is 0.45s against 0.43s.
# Conflicts:
#	Cargo.toml
#	src/client/hyper.rs
#	tests/tls_server.rs
The classifier reads a 64 KiB window from the top of the body, and sliced
a str straight at that byte. Slicing at a byte that is not a character
boundary panics rather than erring, and max_body_size cuts the body at a
byte count, so a capped read can leave a half-written character sitting
exactly there. It took the whole request down with it.

This only became reachable once dev's max_body_size work met the
classifier, which is why neither branch saw it alone. Walk back to the
nearest boundary before slicing.

Also reformat test_mock.py, which ruff 0.15.10 (what CI pins) wanted.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant