Element.scrollContainer read the inline style= attribute only, so a
scroller declared in a stylesheet was invisible to the scroll tool and to
wheel scrolling, which then fell through to the viewport.
StyleManager now tracks overflow-x and overflow-y alongside display,
visibility, opacity and pointer-events, and exposes scrolls(el, axes) as
an own-element probe. The overflow shorthand is expanded into its
longhands in declaration order, in both the attribute scan and the
materialized style object, so a shorthand and its longhands keep the
precedence of the source text. overlay counts as auto, as in Chrome.
Element.scrollContainer asks the style manager, and the two unused Props
bits hold the new flags, so the per-element memo does not grow.
The scroll tool (MCP, agent, LP.scrollNode) wrote scrollTop on the exact
node it was given, so a leaf inside an overflow:auto panel stored an
offset on a non-scroller, the panel's own scroll listener never ran, and
the tool reported the requested coordinates as if it had worked. It also
fired a synchronous bubbling scroll on top of the async non-bubbling
scroll/scrollend the setters already schedule.
actions.scroll now resolves the nearest ancestor-or-self scroll
container, falls back to the node itself, and returns the node that
moved plus the read-back position. The tool and LP.scrollNode report
that instead of the request.
The container query moves from user_input.zig onto Element as
scrollContainer(axes), so the wheel path, the tool and WebDriver share
one resolver. WebDriver's wheel scrolled the hit-test element directly
and fired its own bubbling scroll; it now goes through
user_input.wheelScroll like CDP and BiDi wheel.
Window and Element share one ScrollToOpts. Its offsets() helper
normalizes the positional and dictionary forms once, and an omitted axis
in the dictionary form leaves that axis untouched for the window too,
matching browsers, so scrolling the window on one axis no longer resets
the other.
Runtime.consoleAPICalled serialized object arguments with JSON.stringify,
which can run getters and toJSON callbacks. A callback can log again or
navigate, emitting further CDP events while the outer event is being built.
Resetting send_arena after each send and notification_arena after each
handler then invalidates the outer message's buffers. This produces
use-after-free in debug builds or malformed JSON that disconnects clients.
Scope both arenas to the outermost send or notification handler, including
inspector messages and error paths. Also match Chrome's console argument
representation: objects remain remote handles, primitives carry value,
and non-JSON numbers and bigints use unserializableValue. Console logging
must not invoke object getters or toJSON as a serialization side effect.
Test complete nested WebSocket messages, failure recovery, nested legacy
Console notifications, and primitive/object protocol shapes.
There's been recent work on improve select / options:
- https://github.com/lightpanda-io/browser/pull/3375
- https://github.com/lightpanda-io/browser/pull/3402
- https://github.com/lightpanda-io/browser/pull/3499
One of the main issues is that a select's options "selected" state wasn't always
kept in sync. Select.zig had a boolean flag to mark whether or not a
selectedIndex was explicit set and every read and update would need to dance
around it. Removing the selectedIndex option would not, for example, keep things
in sync.
This removes the flag and keeps the Option._selected in sync, i.e. the sync
happens on write, not on read and is thus naturally recorded in the state of
the Select and its Options.
The write path is more complicated, but the read path is simpler (though the
real win is always being correct).
We've auto-injected `*Frame` into WebApi since forever (used to be called *Page,
but then we split *Page / *Frame, but same same). And it worked wonderfully:
there was always a single *Frame, so the Frame/Context that the JS was being
executed in HAD to be the *Frame that a node belonged to.
But with the addition of iframe and popups, that truth no longer holds. The
*Frame executing the JS (which is the frame that we auto-inject) isn't
necessarily the *Frame that owns a Node.
This is particularly problematic because the *Frame holds a bunch of node/
element data, e.g. `_element_datasets`. So now the DataSet that you get back
depends on the context in which its called..they don't have identity and can
fall out of sync. Some code calls node.ownerFrame() / node.ownerDocument(), but
not all and the hope is to address this throughout the codebase once and for
all.
This is the first in a series of commits meant to fix this long-standing
issue. All it does is change the node.ownerFrame() return value from *Frame to
?*Frame. It's up to each caller to decide how to handle a frameless node, e.g.
clicking a frameless link should not navigate.
Each wheel axis now looks for the nearest ancestor whose inline overflow
along that axis is auto or scroll, and scrolls it through Element.scrollBy,
or the viewport through Window.scrollBy, so the trusted scroll and
scrollend events are scheduled once instead of firing a second, bubbling
scroll inline. The lookup reads the inline style straight from the style
manager instead of building a computed-style object per ancestor.
Both scrollBy implementations saturate the addition, so an oversized delta
from CDP or from a script no longer overflows the i32 position.
isHiddenSelf duplicated isHidden minus the ancestor walk, on both the
StyleManager and the AX writer. An ancestors flag on the existing options
struct keeps one entry point and states the precondition at the call site.
The agent-script tests only drive actions.click, so Input.dispatchMouseEvent's
own gate had no coverage: deleting the `if (!suppressed)` in triggerMousePress
left the whole suite green.
Press on a preventDefault-ing element and assert focus is kept, then press a
focusable one and assert focus moves, so the test pins the gate rather than an
inert default action. Verified to fail when the gate is removed.
The WebDriver path (performPointerSource) is still uncovered; WebDriver.zig has
no unit tests and is exercised through WPT testdriver, so that one needs a
harness rather than another test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ScriptManager, XMLHttpRequest.zig, Fetch, Workers, etc. all take ownership (aka
dupe) the HTTP response from HTTPClient. They all have a headerCallback that
does something like:
```zig
if (transfer.getContentLength()) |cl| {
try self.body.ensureTotalCapacity(self.arena, cl);
}
```
But in all non-streaming cases (which is most cases), the HttpClient buffers
the response and only calls the headerCallback _after_ the body has been
received. Rather than relying on "Content-Length" header, the body buffer can
be sized to the exact body length. Why does this matter? Because the
Content-Length is the length of the body on the wire, and if the body is
compressed (like almost all .js files are), it will under-report the final
body length AND, because most callers are using an arena, the buffer growth
will retain more memory than it should.
This adds a `transfer.bodyLen()` method. Callers which dupe the body now use
this rather than the Content-Length (Content-Length is still used, e.g. for
XHR progress report).
Visibility probes no longer decide whether an element's style attribute
gets parsed into a CSSStyleProperties object. The three layout readers
(getElementAxis, positionStyle, horizontalPosition) create it on demand
through Element.inlineStyle, and StyleManager only ever folds the
attribute text. That removes the scan/materialize mode threaded through
every probe, and a JS layout read now only materializes the elements
whose inline style it actually reads.
Adds a `clutter` option to --strip-mode. This is based on readability.js. It
isn't a direct port (e.g. it doesn't strop bylines). It fallsback to `shell` if
it strips too much (and shell itself can fallback to not stripping anything).
But clutter rarely fallback to shell, only when a page is very small or when
it strips out _a lot_.
Also expanded shell to look at class names and ids.
Some additional API changes:
- Add strip-mode support to pdf/png generation.
- Add LP.dump which provides greater content gathering capability to CDP,
exposing most `fetch` dump-related parameters (e.g. format, strip, selector,
...)
The new "--strip-mode shell" is designed to try to remove non-content elements
such as the header and footer. The end goal is to use readibility.js test cases
as a baseline, but this isn't a port of readibility.js.
This is just the basic implementation of this, e.g removing a few key tags, e.g.
<header>, <footer> and considering some specific roles.
Even if --strip-mode shell is used, we might decide to stick with a whole dump:
it's better to strip not enough than to strip too much. This currently works by
measuring the ratio of non-link text of the stripped vs unstripped page.
tighten socket ownership (on error paths)
allow reaper to be disabled
Handle window where link is being destroyed, worker is still alive, and client
attempts to re-link.
This is a small step towards WebDriver supports (non-bidi). It allows creating
and deleting a BiDi "Session" (e.g. a worker). It also allows attaching a BiDi
driver to an HTTP-created BiDi session (the typical selenium startup flow).
This change unblocks the most basic setup/teardown of Selenium, so it still
isn't enough to actually use a Selenium script as-is. But it's significant
because it models a worker (thread) that isn't tied to a WebSocket, something we
haven't had before.
A consequence of a pure HTTP Session is that we don't have a clear cleanup
signal. There is no "the socket is disconnected". There's a new HTTP reaper
which kills HTTP Sessions after --http-session-timeout. It's expected that
drivers properly DELETE /session/:id. I imagine we're going to run into
--cdp-max-connections limits and need to tweak this code. BUT, this entire flow
is only enabled with --protocol webdriver, so it won't impact exiting CDP users.
Three parameters every Playwright and Stagehand session sends were parsed
and then only logged as not implemented:
- Page.addScriptToEvaluateOnNewDocument runImmediately now also evaluates
the script in the current document, in the requested world.
- Emulation.setDeviceMetricsOverride screenWidth/screenHeight now back
window.screen, kept on the viewport override next to width/height; 0
keeps the current value, as for the other dimensions.
- Browser.setDownloadBehavior browserContextId is checked against the
loaded context, as the Storage commands already do.
Drivers focus a node before typing (chromedp's SendKeys calls DOM.focus,
then Input.dispatchKeyEvent). The method was unknown, so the keystrokes
went to the previously active element and every chromedp form fill was a
no-op. Resolve the node like the other DOM commands and call Element.focus,
which already handles focusability and the blur/focus event sequence.
Input.dispatchMouseEvent mouseWheel (and the BiDi wheel action, which
shares triggerMouseWheel) only ever moved the scroll position of the element
under the cursor, and gave up when no element was there. window.scrollY
reads a separate field that Window.scrollTo maintains, so a wheel never moved
the page and never fired the document's scroll event; every CDP driver's
scroll-by-wheel API was a no-op.
Scroll the nearest ancestor whose overflow is auto or scroll, as before, and
otherwise scroll the window through Window.scrollBy, which already schedules
the trusted scroll and scrollend events. A wheel over empty space targets the
root element instead of returning early.
The AX writer probed the whole ancestor chain three times per node
(prune, ignored, childIds) and elementFromPoint once per node, although
neither ever descends into a hidden element, so below the walk root only
the element's own display can change the verdict. The AX walk now
prunes on the element's own attributes and cascade and hands writeNode
the verdict; the root and the query walk keep the chain probe.
elementFromPoint carries the verdict down its stack while still counting
hidden nodes, so positions keep agreeing with getBoundingClientRect.
Visibility probes no longer decide whether an element's style attribute
gets parsed into a CSSStyleProperties object. The three layout readers
(getElementAxis, positionStyle, horizontalPosition) create it on demand
through Element.inlineStyle, and StyleManager only ever folds the
attribute text. That removes the scan/materialize mode threaded through
every probe, and a JS layout read now only materializes the elements
whose inline style it actually reads.
With the StyleManager memo, a VisibilityCache/PointerEventsCache hit
costs the same hash probe as a memo hit, so the caches only added a
second probe per ancestor, a call_arena allocation per element, and
plumbing through SemanticTree, AXNode, CDP, links, ResizeObserver and
elementFromPoint. checkVisibilityCached becomes isVisible.
Specifically to distinguish cross-site iframe navigation from top-level navigation, this PR reworks how `SameSite=Lax` moved. Since we're not checking if its a navigation alone now, the field for it is also renamed to `lax_allowed`.