Commit Graph
148 Commits
Author SHA1 Message Date
Muki Kiboigo a246a67d93 add CorsStore 2026-09-23 07:09:01 -07:00
Karl Seguin 0b66a5ed05 http: dont' re-use connections which are likely in a bad state.
Some status-codes should never have a body except for a single trailing blank
line. If we don't handle these, then we end up with a dirty connection in our
connection pool:

1 - read the header, but not the body
2 - put the connection back in the pool
3 - try to read the header, but actually get the body from #1

WPT /fetch/api/basic/response-null-body.any.html exercises this path and is
flaky (because it depends whether the request goes back out on a keep-alive
connection)..but for a given run,you'll almost always get 1-3 failures.

This commit processes the request, but tells libcurl not to re-use the
connection.
2026-09-22 10:08:41 +08:00
Karl Seguin 3213342055 http: improve caching
1 - Centralized cache-awareness into Cache and pulled header details out of
    SqliteCache and HttpClient

2 - Added support for expires header

3 - Support caching more status types (but not all, since HttpClient would need
    to be aware of what caching a 3xx/206 means)

4 - Revalidate cares about  "not specified" vs "no-store" vs "stale"
    (e.g. expires=0 means "stale", not fallthrough the last-modified logic)
2026-09-21 11:38:34 +08:00
Karl Seguin f70705b2ce webapi: include Sec-Fetch-Site and Sec-Fetch-Mode headers 2026-09-18 12:34:19 +08:00
Karl Seguin 437a9c27d6 Merge pull request #3543 from lightpanda-io/fix-import-crash
crash: fix a rare import crash
2026-09-18 05:53:27 +08:00
Karl Seguin a4e1fa95b9 crash: fix a rare import crash
Currently, our waitForImport blocks the caller, but continues to process any
already-queued requests. This can result in new JavaScript running while v8
is linking modules and that JavaScript can itself import a module that is
part of the still-being-linked graph.

waitForImport now works like a syncRequest. While HttpClient will continue to
make progress on all transfers, all other transfers will gate behind the waiting
one (using the same infrastructure that exists for syncRequest).

This crash was seen on an unknown srape URL.
2026-09-17 08:17:13 +08:00
Karl Seguin ddf0fa2ce9 Merge pull request #3538 from lightpanda-io/webdriver-navigate
WebDriver: add navigate
2026-09-17 07:37:36 +08:00
Karl Seguin bc5f0777fb Merge pull request #3540 from lightpanda-io/cors-redirect-credentials
webapi: limit redirect with credentials
2026-09-17 07:24:18 +08:00
Karl Seguin b15707973b webapi: limit redirect with credentials
A cors request with credentials can only follow redirects when staying on the
same origin. WPT cors-redirect-credentials
2026-09-16 17:41:32 +08:00
Karl Seguin afbc8378b1 Merge pull request #3522 from lightpanda-io/regex-shared
`Regex`: share the PCRE2 wrapper; `findElement` matches names by `/regex/`
2026-09-16 16:55:40 +08:00
Karl Seguin 7ecefa9e92 WebDriver: add navigate
This feature is significant because it adds support for processing an HTTP
request via the worker. It requires parking the connection and then having the
worker notify the loop when the response is ready. A lot of this was already
in-place (e.g. worker -> loop notification) but not quite do this extent.
2026-09-16 14:05:49 +08:00
Scott Taylor 8339361043 http: drop request-body headers when redirects rewrite the method
A multipart form POST followed by a 302 changed to GET and lost its body,
but retained Content-Type: multipart/form-data. Servers could then try to
parse an absent multipart body and return 400. This was reproduced on a
local redirect server and a storefront localization flow.

Delete Fetch's request-body header names when rewriting to GET. Preserve
method and body on 307/308, rewrite only POST on 301/302, and preserve GET
and HEAD on 303 rather than rewriting every request indiscriminately.

Test method/header transitions and header handling through the existing
CDP fulfilled-redirect path.
2026-09-15 07:48:18 -04:00
Adrià Arrufat d2f72a9fa9 Regex: share the PCRE2 wrapper beyond adblock
Compiled patterns are useful anywhere someone else writes the pattern:
the adblock lists today, agent tool arguments next. The wrapper moves
out of the adblock directory and gains an options struct (case, UTF-8
subjects) and a compile diagnostic the caller can log or show. The App
owns the one context every consumer compiles through, the blocker
included.
2026-09-15 09:33:39 +02:00
Halil Durak 93381010f1 Merge branch 'main' into nikneym/lax-exception-RFC6265bis 2026-09-14 14:27:55 +03:00
Muki Kiboigo 00c98313c2 use curl no body option for head requests 2026-09-14 07:58:56 +08:00
Karl Seguin 35c0a9aeda chore: dedupe HttpClient.Owner using new GlobalScope
https://github.com/lightpanda-io/browser/pull/3447 made better use of the
GlobalScope to simplify various callsites. This changes HttpClient.Owner to
contain the global_scope, rather than copying a handful of scope fields.
2026-09-12 12:09:56 +08:00
Karl Seguin f81f6e4eb7 mem: reduce memory usage of cloned HTTP responses
ScriptManager, XMLHttpRequest.zig, Fetch, Workers, etc. all take ownership (aka
dupe) the HTTP response from HTTPClient. They all have a headerCallback that
does something like:

```zig
if (transfer.getContentLength()) |cl| {
  try self.body.ensureTotalCapacity(self.arena, cl);
}
```

But in all non-streaming cases (which is most cases),  the HttpClient buffers
the response and only calls the headerCallback _after_ the body has been
received. Rather than relying on "Content-Length" header, the body buffer can
be sized to the exact body length. Why does this matter? Because the
Content-Length is the length of the body on the wire, and if the body is
compressed (like almost all .js files are), it will under-report the final
body length AND, because most callers are using an arena, the buffer growth
will retain more memory than it should.

This adds a `transfer.bodyLen()` method. Callers which dupe the body now use
this rather than the Content-Length (Content-Length is still used, e.g. for
XHR progress report).
2026-09-11 11:50:30 +08:00
Karl Seguin 8ad9eaf48d webdriver: HTTP WebDriver session management
This is a small step towards WebDriver supports (non-bidi). It allows creating
and deleting a BiDi "Session" (e.g. a worker). It also allows attaching a BiDi
driver to an HTTP-created BiDi session (the typical selenium startup flow).

This change unblocks the most basic setup/teardown of Selenium, so it still
isn't enough to actually use a Selenium script as-is. But it's significant
because it models a worker (thread) that isn't tied to a WebSocket, something we
haven't had before.

A consequence of a pure HTTP Session is that we don't have a clear cleanup
signal. There is no "the socket is disconnected". There's a new HTTP reaper
which kills HTTP Sessions after --http-session-timeout. It's expected that
drivers properly DELETE /session/:id. I imagine we're going to run into
--cdp-max-connections limits and need to tweak this code. BUT, this entire flow
is only enabled with --protocol webdriver, so it won't impact exiting CDP users.
2026-09-11 05:11:26 +08:00
Karl Seguin 72165ef2a4 Merge pull request #3476 from lightpanda-io/make-private-if-private
chore: make declarations private if they don't need to be public
2026-09-11 05:09:17 +08:00
Adrià Arrufat d693f49872 Merge pull request #3463 from lightpanda-io/adblock-regex-pcre2
`AdBlocker`: run /regex/ filters with PCRE2
2026-09-10 14:04:14 +02:00
Karl Seguin 2e6999f20b chore: make declarations private if they don't need to be public
This change is 99%  s/pub//   + a handful of dead code removal.
2026-09-10 14:42:09 +08:00
Karl Seguin 33ddbdf0fc Merge pull request #3466 from lightpanda-io/locale-timezone
cli: add --locale and --timezone, derive language signals from one value
2026-09-10 08:06:39 +08:00
Halil Durak 380bdcfa00 change how Lax allowance computed (more RFC 2625bis compliance)
Specifically to distinguish cross-site iframe navigation from top-level navigation, this PR reworks how `SameSite=Lax` moved. Since we're not checking if its a navigation alone now, the field for it is also renamed to `lax_allowed`.
2026-09-09 19:43:41 +03:00
Adrià Arrufat 37bf7b68c1 cli: one Accept-Language parser for the header and navigator.languages
navigator.languages now lists the Accept-Language tags in order, which is
Chrome's contract, instead of a second derivation from the locale tag that
disagreed with the header (--locale de-DE sent de-DE,de,en but reported
["de-DE","de"]). HttpHeaders.AcceptLanguage owns both shapes and is also
the CDP override type.

ICU canonicalizes a BCP 47 tag read from LC_ALL itself, script subtag
included, so the POSIX id conversion is gone; it dropped the script and
turned zh-Hans-TW into Traditional Chinese.

Also: the CDP handler keeps validateUserAgent's verdict instead of scanning
for Mozilla twice, the override is cleared unconditionally on context
teardown instead of through a flag, and the flags are sentinel strings so
Platform passes them to setenv without copying.
2026-09-09 18:07:03 +02:00
Adrià Arrufat c55afc1df9 cli: add --locale and --timezone, derive language signals from one value
navigator.language was hard-coded to en-US and Accept-Language was a
constant, while Intl, toLocaleString and Date followed the host process
environment. On a de_DE host a page saw navigator.language === "en-US"
next to German number formatting, a mismatch fingerprinting scripts look
for, and the same page rendered differently across machines.

Follow Chrome's --lang rule: one configured tag drives navigator.language(s),
the Accept-Language header and ICU's default locale. --locale defaults to
en-US, so Intl is now en-US on every host instead of whatever LANG says.
--timezone sets the IANA zone Date and Intl use; absent, the host zone stays.

Both are applied by writing LC_ALL and TZ before V8 initializes ICU, which
reads them lazily. Platform.init is the first call in App.init, before any
thread exists, so setenv is safe there.

CDP Emulation.setUserAgentOverride.acceptLanguage, which Playwright sends
for its locale option, now overrides the header and navigator.languages
for the browser context's lifetime, mirroring the user agent override, and
applies even when the Mozilla user agent is refused.

Emulation.setLocaleOverride and setTimezoneOverride stay no-ops: changing
ICU's defaults at runtime needs new zig-v8-fork bindings.
2026-09-09 17:49:02 +02:00
Adrià Arrufat 433ca9b747 adblock: fold the regex path into the existing mechanisms
The raw URL lives on `pattern.Url` next to the lowercased one, so
`pattern.matches` owns the `.regex` arm like every other kind and the
engine stops special-casing it. `Request.init` does the lowercasing
itself, as `fromHttp` already had to, instead of asking callers for
both spellings.

The regex shape now spells its uncertain marker as `*` and keeps
non-token literals as one marker, so it is read by the same
bounded-token loop as a plain pattern rather than a copy of it. The
quantifier parser keeps only what it uses: whether the atom may be
absent.

`Regex.matches` runs on a stack-first allocator: PCRE2 wants a match
data block and 20KB of backtracking frames per call, which no longer
touches the heap in the common case. A filter holds a pointer to its
regex, keeping `NetworkFilter` at its previous size.
2026-09-09 13:13:18 +02:00
Adrià Arrufat c9329d5955 AdBlocker: run /regex/ filters with PCRE2
Filter lists carry a few hundred rules written as JavaScript regex
literals (24 in EasyList, 165 in uBO's badware list); they parsed but
were dropped as unsupported. PCRE2 reads that syntax as-is, `\/` and
friends included, its compiled patterns are immutable so the one
blocker shared by every HTTP client thread can run them, and 10.48
ships a build.zig for 0.16, so it is wired like sqlite3.

`Regex.Context` routes every PCRE2 allocation through the blocker's
allocator, which puts the compiled patterns under the test runner's
leak detection, and caps match and depth so a broken pattern costs a
false negative rather than a stalled request. As in uBO, a regex
tests the raw URL with the case-insensitive flag unless `$match-case`.

Regex filters are still never tokenized: they ride the fallback bucket.
2026-09-09 12:43:39 +02:00
Karl Seguin 66af1dc58d webapi: Origin header + request header validation
Driven by a handful of /fetch/ WPT tests, three changes:

1 - Prevent libcurl from auto-inserting a 'application/x-www-form-urlencoded"
    content type for types we really have no content-type for.

2 - Include origin header in all requests that should have it. This is something
    CorsGate was doing in most cases, but cors can be disabled, so the logic
    is now moved to HttpClient.

3 - Expands on the header guard added in  https://github.com/lightpanda-io/browser/pull/3374/
    Adds more modes and more header check. Request.init also uses the header
    guard now
2026-09-09 14:44:48 +08:00
Karl Seguin b633dc4b16 webapi: improve various WPT fetch apis
Headers strip whitespace and guard against invalid characters

Headers iterator sorts and combines PER step, so that mutations are picked up.
Not the most efficient, but this is a short list, and how often are these being
iterated?

XMLHttpRequest: has its own extra header validation

Mime support for multiple Content-Type headers (or a header with multiple values)
last value wins.

Add BufferSource js bridge type that accepts various types -> []const u8 (at the
cost of losing the actual type). Useful in fetch, where various types can be a
body, but we only care about the underlying bytes (e.g. we didn't support A
rrayBufferView before this)

Refactored response body getters so that they all go through the same consume
and resolve logic
2026-09-08 13:12:20 +08:00
Halil Durak 7ad73d0a6d changes after rebase 2026-09-07 11:14:15 +03:00
Halil Durak 0a6d43c2bb HttpClient.isAdBlocked should just delegate Adblocker.isBlocked 2026-09-07 11:00:47 +03:00
Halil Durak 486c888080 drop is_subframe and notification 2026-09-07 11:00:47 +03:00
Halil Durak e696a29da7 metrics: make adblocker verdicts measurable 2026-09-07 11:00:05 +03:00
Halil Durak b60f790a03 Adblocker: changes on request construction & tokenization
* Engine.Request.fromHttp(req, source_url, buffers) now builds the adblock request straight from HttpClient.Request.
* The URL is tokenized once per request (hashed into the Request, shared by all engines); capped at 128 tokens (same as adblock-rust).
* Document hostname longer than 253 bytes now skips adblocking.
2026-09-07 10:57:36 +03:00
Halil Durak 6a85b1f386 handle all switch cases (emerged after rebase to main) 2026-09-07 10:57:36 +03:00
Halil Durak 8aa0dda4b3 integrate request engine to Adblocker, apply changes required for filter list in HttpClient 2026-09-07 10:57:35 +03:00
Muki Kiboigo 34b743fbd6 origin is tainted on cross origin redirects 2026-09-04 07:02:01 -07:00
Muki Kiboigo 54518383c0 enforce cors response on redirects as well 2026-09-04 07:00:03 -07:00
Muki Kiboigo 0610d5ecd1 collapse isCrossOriginModeAllowed check in pipeline 2026-09-04 07:00:02 -07:00
Muki Kiboigo 0aeba826b4 use credentials_mode instead of cookie bool 2026-09-04 07:00:02 -07:00
Muki Kiboigo 827d0f0598 only check cross origin mode on obey cors 2026-09-04 07:00:02 -07:00
Muki Kiboigo 8a8bb3814f fix test running 2026-09-04 07:00:01 -07:00
Muki Kiboigo 15d27902ba use experimental features flag instead of obey cors 2026-09-04 07:00:01 -07:00
Muki Kiboigo 1fe8456cbd add modes to the tests 2026-09-04 07:00:01 -07:00
Muki Kiboigo 12ed38dfdd better no cors opaque behavior 2026-09-04 07:00:00 -07:00
Muki Kiboigo d1f4605459 non-default credentials and request mode 2026-09-04 07:00:00 -07:00
Muki Kiboigo feebb889ad cors check before cache check 2026-09-04 06:59:40 -07:00
Muki Kiboigo d3c0291bd1 add request mode for Fetch 2026-09-04 06:59:39 -07:00
Muki Kiboigo ddfa034310 add credentials_mode for proper CORS credentials handling 2026-09-04 06:59:39 -07:00
Muki Kiboigo f31b32ac4e don't store network in CorsGate 2026-09-04 06:59:38 -07:00