Rather than having everything tied to the page's memory (arena, factory), most
things are now tied to an IDBTransaction and its arena. The IDBTransaction is
reference counted and finalized with v8. The lifecycle is relatively complicated
compared to anything else we have, which is a concern, but using the page arena
seems like a dealbreaker to me.
Any "child" created for the transaction (e.g. IDBRequest) has its v8 acquireRef
and releaseRef forwarded to the Transaction. This should make v8's usage safe.
However, in addition to this, Zig itself must take a RC whenever a drain is
scheduled AND whenever the transaction is parked in the engine. And these
things, especially on cleanup, can be a little messy.
The _awful_ cursor allocations have been improved by a local re-used ArrayLists
for the key/primary key/value. So rather than accumulating _every_ allocation
we now only capture the peak.
One final memory-related area this commit addresses is the lifetime difference
between Transactions (Page) and Engine (Session). This is problematic because
the Engine can reference the Transaction. On js.Context deinit, the engine is
notified and all related Transactions are canceled/removed. This is...unusual
in our design. We have other similar cases that use a list on the frame/GWS to
track resources to cleanup. But, the Engine already _has_ to have this list so
rather than book-keeping in Engine AND Frame AND WGS, only Engine has a list at
the cost of js.Context having to notify it on teardown.
Because of the single-connection nature of our Engine (in-memory SQLite), the
IDB Engine now keeps a list of "gated" transactions, processing one at time.
When one transaction completes, waiting transactions are signaled so that they
can processes their queue.
This also adds a double-queue so that any requests that comes in while
processing requests goes into a new queue and is only drained on the next tick.
Note: this introduces a UAF, where Engine lives on Session and referneces
IDBTransactions which are tied to the page. The next bit work on IDB is the
memory re-work, so I'm punting that there.
Previously, every operation would run synchronously, and resolve the value
on the next drain (schedule task run on the next tick).
With 1 connection per DB, this doesn't work with SQLite: you cannot have nested
transactions. Imagine:
add
begin
insert into
JS callback (success)
add
begin <-- nested
insert
This approach captures the requested operation and does all necessary validation
(since most validation errors are returned synchronously), but only executes
operations on drain. Hence, rather than IDBTransaction._requests being a
queue of result values to emit on drain, it is now a queue of operations to
execute and then emit.
This will also potentially make it easier to address serious memory
issues by better scoping the lifetime of things (particularly requests).
Initial WIP on IndexedDB WebAPI. Uses sqlite and a new `--indexdb_dir` config.
Defaults to :memory:.
The main things missing are the IDBCursor, IDBIndex and IDBKeyRange, along with
a bunch of smaller apis.
Also, every object is currently tied to the Page arena / factory. It's possible
that's how it will have to be, but, once more of the API lands, I will check
if we can scope these better.
The `call_arena` is currently our shortest-lived arena. Its promise is that it
will be valid for at least 1 v8 -> zig -> v8 function call, which makes it ideal
for getting values from v8 into zig, temporary work, and getting values from
zig to v8.
VBut `call_arena`'s lifetime is actually much longer. It's only reset when the
call_depth reaches 0. And the reason for that are Zig functions that invoke
v8 callbacks. If you just do it when a function ends, then if you allocate in
ZigA and then call CallbackA which calls ZigB, then when ZigB ends, your ZigA
allocation is cleared. This is something we could solve by reaching into
Zig's ArenaAllocator to capture a position to rollback from.
So, `call_arena` can end up living relatively long and accumulating quite a bit
of memory. But most calls _don't_ invoke callbacks. Hence, `local_arena` IS
reset at the end of every function.
Using `local_arena` _is_ dangerous, mostly for the case where some of our APIs
receive a *Frame (or *Execution) and thus might use a `local_arena` because they
know they aren't invoking a JS callback. BUT, those APIs might not know that
they're also invoked by some other Zig code that _could_ be invoking JS.
Dangerous? Sure. But, github has code that looks like:
```js
function onDelegatedClick(e) {
for (const el of document.querySelectorAll('[data-action]')) {
if (el.matches('.menu > .item:not(.disabled) a[href]')) {
handle(el);
}
el.closest('.panel');
}
}
```
Anything allocated in the `call_arena` will only be freed when the caller of
`onDelegatedClick` ends. The `querySelectorAll` returns hundreds of elements
and thus builds hundreds of parsed CSS and other scrap (x2 for `matches` and
closest`). `call_arena` peaks at 15MB. With the local_arena? 1MB. If 10x more
elements were returned, the peak would be 10x higher. With local_arena, it
stays 1MB.
When I added this, I was going through the list in /dom/interface-objects.html
thinking that was exhaustive. But no, no interfaces should be enumerable and
various other WPT tests (usually the idlharness ones) assert that for their
respective types.
Make it _always_ non enumerable means we no longer need Meta.enumerable to
be declared true/false (it's always false).
std.crypto.random's default backend mmaps a thread-local 528-byte state
page on first use and never unmaps it — there is no thread-exit hook.
With one detached thread per CDP connection (Server.handleConnection),
that leaks one resident page per connection (uuidv4 in
Page.getOrCreateOrigin touches it), ~4KB/conn of unbounded RSS growth.
Route every std.crypto.random call to the getrandom syscall instead.
.crypto_always_getrandom = true,
Idea here is to skip re-parsing that happen for each connection; we already use BoringSSL, so we can take more advantage of it by directly mutating cert store of `SSL_CTX`.
Adds the infrastructure for [de]serializing Zig objects via structuredClone.
Adds support to Blob, File, FileList and ImageData. These are the easiest to
implement. Blob is used extensively by WPT IndexedDB tests, but this PR can be
merged prior to IndexedDB landing.
On for the main document parsing should a leading BOM be stripped. When setting
innerHTML, it should be preserved (and becomes a text node).
Fixes react hydration issue with theverge.com
Per review feedback: public querySelector doesn't guarantee reuse, so the cache
must be bounded regardless of the SelectorPath bypass. Move it off the Frame
(where it was wiped every navigation and unbounded) onto the Browser, since a
parsed selector references no Frame/Context — entries are now shared across the
browser's pages and survive navigation.
Selector.Cache is a StringArrayHashMap with per-entry arenas (so eviction can
free an individual entry, which a shared arena can't) and FIFO eviction of the
oldest entry past a capacity. The SelectorPath *Uncached bypass stays.
Replace the arbitrary 1024-entry cap with an explicit split: the public
querySelector/querySelectorAll/matches/closest entry points cache (page scripts,
waitForSelector — selectors that recur), while SelectorPath's synthesized one-off
candidates use new *Uncached variants that parse into a transient arena. The
cache now only ever holds genuinely-reused selectors, so it needs no size bound.
Lightpanda installs no SIGSEGV handler, so a segfault (or the abort() in
the panic path) falls through to the kernel and writes a core dump. When
many instances run under a shared core_pattern crash reporter -- e.g. a
containerized crawl fleet -- those cores become pure storage/alert noise,
and a browser core can capture the contents of arbitrary pages.
Crashes are already reported via telemetry, so this adds an opt-in
LIGHTPANDA_DISABLE_CORE_DUMP env var (mirroring LIGHTPANDA_DISABLE_TELEMETRY)
that zeroes the soft RLIMIT_CORE at startup. Default behavior is unchanged.
Remove sid. Include iid in every message. Booleans true/false => 1/0. Constant
string values => single letter. Example:
["8800df58-a5d5-4ca5-9a06-d7691a2a3780","H","fetch",0,"macos","aarch64","1.0.0-dev.7609+88b1bc671"]
["8800df58-a5d5-4ca5-9a06-d7691a2a3780","R"]
["8800df58-a5d5-4ca5-9a06-d7691a2a3780","N",1,"P"]
H => Header
R => Run
N => Navigate
B => Buffer Overflow / dropped
L => LLM
Navigate context are
P => Page
O => Open (popup)
I => Iframe
First, this adds 1 small piece of data to the navigate event: whether the
navigate was a page, frame or popup.
It also adds a session id, but as far as I'm concerned, this isn't "new"
information, or any new tracking/insight into users. Between the iid and the
"run" event, a "session" was always trackable. By giving it an explicit value,
we can shrink the size of all other messes.
This change reduces the telemetry payload by ~70% (despite the extra nav field).
I'm hoping this might remove a reason some people would consider turning it off.
It hits a /v2/ endpoint. The changes:
1 - A header is the first message in a session and contains all of the static
data, as well as a session id
2 - Every event is encoded as an array, [$SID, "event-type, params...]
```
{"sid":"92e98210a141f497","iid":"$UUID","mode":"fetch","os":"macos","arch":"aarch64","version":"$VERSION","proxy":false}
["92e98210a141f497","run"]
["92e98210a141f497","nav",false,"page"]
```
(the driver=cdp field was removed from nav, because it was always cdp).
Some notes for the server:
1 - The server can tell a header from an event based on the first character.
2 - The SID is 8 bytes, enough to be unique, but not globally unique. The
iid + sid + time window is how a events for the same SID can be grouped.
3 - A valid event is always an array of 2+ items, index 0 = SID, index 1 = type
Although positional data isn't expressive, it's still extendable.
The 4 events:
run, no parameters (mostly just used to flush the header now)
["sid", "run"]
// nav, tls, page/popup/iframe
["sid","nav",true,"page"]
// bof, # of lost telemetry events
["sid","bof",42]
// llm, provider, model (nullable)
["sid","llm","anthropic","claude"]