logToErased asserts that a log message is at most 30 characters of
plain text, but only in a debug build and only once the line actually
runs. A message on a rare path therefore ships fine and then panics on
whoever first reaches it: `serve --host 0.0.0.0` without
--advertise-host crashed on startup in every debug build, because
"advertising loopback for wildcard bind" is 38 characters.
Every message is a literal, so make the six wrappers take a comptime
msg and apply the same two rules through @compileError. The runtime
check stays as the backstop for the paths the compiler does not
analyse for the current target, and now reads the same constant.
Eleven messages were over the limit; shorten them. The detail already
lives in the kv pairs in each case. renderFailed takes its message as
comptime now, the only call site that passed a runtime one.
Note the check only covers code analysed for the target being built:
the two in Certificates.zig sit in an OS switch prong that Linux never
compiles, and were found by scanning the source rather than by the
compiler.
Some status-codes should never have a body except for a single trailing blank
line. If we don't handle these, then we end up with a dirty connection in our
connection pool:
1 - read the header, but not the body
2 - put the connection back in the pool
3 - try to read the header, but actually get the body from #1
WPT /fetch/api/basic/response-null-body.any.html exercises this path and is
flaky (because it depends whether the request goes back out on a keep-alive
connection)..but for a given run,you'll almost always get 1-3 failures.
This commit processes the request, but tells libcurl not to re-use the
connection.
1 - Centralized cache-awareness into Cache and pulled header details out of
SqliteCache and HttpClient
2 - Added support for expires header
3 - Support caching more status types (but not all, since HttpClient would need
to be aware of what caching a 3xx/206 means)
4 - Revalidate cares about "not specified" vs "no-store" vs "stale"
(e.g. expires=0 means "stale", not fallthrough the last-modified logic)
Currently, our waitForImport blocks the caller, but continues to process any
already-queued requests. This can result in new JavaScript running while v8
is linking modules and that JavaScript can itself import a module that is
part of the still-being-linked graph.
waitForImport now works like a syncRequest. While HttpClient will continue to
make progress on all transfers, all other transfers will gate behind the waiting
one (using the same infrastructure that exists for syncRequest).
This crash was seen on an unknown srape URL.
This feature is significant because it adds support for processing an HTTP
request via the worker. It requires parking the connection and then having the
worker notify the loop when the response is ready. A lot of this was already
in-place (e.g. worker -> loop notification) but not quite do this extent.
A multipart form POST followed by a 302 changed to GET and lost its body,
but retained Content-Type: multipart/form-data. Servers could then try to
parse an absent multipart body and return 400. This was reproduced on a
local redirect server and a storefront localization flow.
Delete Fetch's request-body header names when rewriting to GET. Preserve
method and body on 307/308, rewrite only POST on 301/302, and preserve GET
and HEAD on 303 rather than rewriting every request indiscriminately.
Test method/header transitions and header handling through the existing
CDP fulfilled-redirect path.
Compiled patterns are useful anywhere someone else writes the pattern:
the adblock lists today, agent tool arguments next. The wrapper moves
out of the adblock directory and gains an options struct (case, UTF-8
subjects) and a compile diagnostic the caller can log or show. The App
owns the one context every consumer compiles through, the blocker
included.
Requests issued to an origin while its first connection is still
handshaking each opened their own socket, up to --http-max-host-open,
because curl only learns from ALPN whether the origin multiplexes. With
pipewait they wait for that answer and share one h2 connection.
Fixture: 12 fetch() calls to a fresh cdnjs (h2) origin, release build.
new TCP+TLS connections 6 -> 1
in-page time to last resp ~285 ms -> ~105-135 ms
H1-only origins are unchanged in connection count; their first burst
waits one handshake before fanning out.
https://github.com/lightpanda-io/browser/pull/3447 made better use of the
GlobalScope to simplify various callsites. This changes HttpClient.Owner to
contain the global_scope, rather than copying a handful of scope fields.
The RobotStore is shared by all Browsers. While every browser has a single
flight to prevent duplicate requests to the same robots.txt, that's limited to
that specific browser. So, 2 browsers can ask for the same robots.txt and then
put try to store the result. The RobotStore _is_ thread safe, but it's a simple
last-one-wins which overrites the previous record, without freeing either the
key or value.
This replaces the last-write-wins with a first-write-wins, avoiding the leak.
ScriptManager, XMLHttpRequest.zig, Fetch, Workers, etc. all take ownership (aka
dupe) the HTTP response from HTTPClient. They all have a headerCallback that
does something like:
```zig
if (transfer.getContentLength()) |cl| {
try self.body.ensureTotalCapacity(self.arena, cl);
}
```
But in all non-streaming cases (which is most cases), the HttpClient buffers
the response and only calls the headerCallback _after_ the body has been
received. Rather than relying on "Content-Length" header, the body buffer can
be sized to the exact body length. Why does this matter? Because the
Content-Length is the length of the body on the wire, and if the body is
compressed (like almost all .js files are), it will under-report the final
body length AND, because most callers are using an arena, the buffer growth
will retain more memory than it should.
This adds a `transfer.bodyLen()` method. Callers which dupe the body now use
this rather than the Content-Length (Content-Length is still used, e.g. for
XHR progress report).
This is a small step towards WebDriver supports (non-bidi). It allows creating
and deleting a BiDi "Session" (e.g. a worker). It also allows attaching a BiDi
driver to an HTTP-created BiDi session (the typical selenium startup flow).
This change unblocks the most basic setup/teardown of Selenium, so it still
isn't enough to actually use a Selenium script as-is. But it's significant
because it models a worker (thread) that isn't tied to a WebSocket, something we
haven't had before.
A consequence of a pure HTTP Session is that we don't have a clear cleanup
signal. There is no "the socket is disconnected". There's a new HTTP reaper
which kills HTTP Sessions after --http-session-timeout. It's expected that
drivers properly DELETE /session/:id. I imagine we're going to run into
--cdp-max-connections limits and need to tweak this code. BUT, this entire flow
is only enabled with --protocol webdriver, so it won't impact exiting CDP users.
An alternation or a repeat keeps only whether its text may start and
end with a token character, and read that off one marker. An optional
non-token stretch there (`\/?x`) was taken as a definite non-token,
so `\/ads(\/?x|\/y)` was filed under "ads" while `/adsx` carries no
such token. What follows the stretch answers now.