The http max default was 4K with a 16KB hard limit. The default limit is now 1MB
with an initial default of 4K. This is to accommodate larger WebDriver payloads.
DOM.getFrameOwner answered with the child frame's document instead of the
element hosting it. Clients that map an <iframe> to its frame compare the
backendNodeId it returns with the one they resolved in the parent
(Stagehand's frameLocator/deepLocator, Playwright's contentFrame), so the
lookup never matched and they could not descend into frames.
Return the owner element, as Chrome does, and error for the main frame.
The node writer now emits frameId on frame-owner elements and on document
nodes, which is the other half of that pairing.
Continuation of cleaning up frame ownership (#3549, #3536, #3520, ...).
Element.blur/focus are noop on a frameless element.
AXNode.Writer rejects frameless nodes. DOM.resolveNodes keeps its fallback
bacause it's stateless - just needs the context.
Every key path resolves its target through user_input.focusedElement,
which is document.activeElement with the <body> fallback. WebDriver and
the MCP press action used to fall back to the document node instead, a
target no default action or char step handles.
The Enter and Space activation rules share isButton, and the text-entry
rule (no text goes into a checkbox or radio) lives on the TextEntry
mixin as acceptsTextEntry rather than as a type test inside the shared
insertion helper. pressKey takes its extra ref only when a keypress will
be built from the keydown.
Chrome clicks a button on the keypress Enter produces, after the keypress
fires; only links follow Enter on the keydown. Clicking on the keydown put
the click before the keypress and, with the char step also submitting for
button-type inputs, submitted the form twice.
The MCP press action relied on the same char step for Enter, so its own
implicit submission is gone with the double submit it caused on buttons.
The #3549 test only covered LP.getSemanticTree's text format. Cover
the JSON format with interactiveOnly on the same fixture, and drive the
MCP tree and nodeDetails tools through call() with a child-frame
backendNodeId. Passing the main frame in either tool now fails the
suite instead of only tripping a debug assert.
Every caller resolved node.ownerFrame(frame) before building a
SemanticTree or calling getNodeDetails, and every one made the same
decision on a frameless node. Move that into SemanticTree.init, which
returns error.FramelessNode, and turn getNodeDetails into a method so
it goes through the same constructor.
With init as the only entry point the per-method assertOwns is
redundant, so drop it along with its copy of StyleManager's helper.
A nodeId root inside a child frame was walked with the main frame. The
style checks already resolve the owner frame per element, but the
listener map was built from the main frame's event manager, so
listener-only elements in the iframe were reported non-interactive, and
relative hrefs resolved against the parent's base URL.
Resolve the root's owner frame before the walk, as getSemanticTree and
getNodeDetails do.
Follow Chrome's model. A keydown only runs its own default action (Tab,
caret moves, Backspace, Enter activation). Text is typed by the char half
of the press, which fires keypress and then beforeinput/textInput, so
either can veto the edit.
A keyDown carrying `text` runs the char step inline (Puppeteer,
Playwright). A text-less keyDown followed by a `char` message runs it on
the char (chromedp), so each character is typed once. A cancelled
text-less keyDown drops the char that follows it, as Chrome does.
BiDi, WebDriver and the MCP press action go through the same pressKey,
deriving the text from the key. pressKey holds a ref on the keydown for
the keypress it builds, so nothing reads the event after dispatch.
Input.insertText fires beforeinput like a key press does.
Applies the frame-ownership pass to SemanticTree, copying what we did for
StyleManager (1). SemanticTree doesn't visit iframes, so the frame of the root
is the frame/frame._style_manager we need to target for all visited nodes.
Like #3536, it's up to the callers to (a) get the correct frame and (b) decide
what to do on a frameless-node.
(1) https://github.com/lightpanda-io/browser/pull/3536
Fold the two loose Page fields (input_pressed_buttons, input_mousedown_suppressed)
into a PointerButtons struct in user_input.zig that owns the chord mask and the
press/release transitions, so triggerMousePress/Release keep only dispatch.
Add runMouseDownFocus so the three click callers stop repeating the
!suppress_mouse and !suppress_focus gate; a focus-error policy param keeps each
caller's warn-vs-propagate behavior.
Trim the verbose dispatch/trigger and CDP-test comments to a single sentence
each, moving behavior contracts onto the function doc-comments; also drop the
stale buttonsMask comment orphaned by the earlier buttonsBitmask reuse.
Round-1 review feedback from arrufat on #3507 (share pointer/mouse click
dispatch across click paths):
- CDP mousePressed now threads clickCount into mousedown's detail
(mouse_detail), matching the MCP path (actions.click) instead of
always firing detail 0. Threaded through BiDi's press call too, which
was already tracking click_count for release but silently dropping it
on press.
- dispatchClickAsPointer now carries the still-held buttons mask instead
of hardcoding 0, so a primary click firing mid-chord (the primary
button releasing while another button is still held) reports the
correct PointerEvent.buttons.
- Fixed a regression caught while verifying the above against live
Chrome: the first fix's fallback forced clickCount 0 (CDP's default
when the field is omitted) to detail 1. Chrome and Firefox both
preserve 0 there, and Chrome doesn't fire `click` at all in that case
— the fallback now preserves 0 instead of forcing 1.
Left two pre-existing issues alone, flagged in code comments instead of
fixed, per arrufat's own scoping:
- triggerMouseRelease's parallel clickCount-0 fallback has the same
mismatch on the release side; not introduced by this PR.
- The chord's activation-event ordering (contextmenu on the right press
vs. this codebase's release-only contextmenu, and no auxclick) is a
real, pre-existing deviation from both Chrome and Firefox, confirmed
on Firefox too during this round, not Chrome-specific.
Test coverage: a parametrized CDP test over clickCount 0/1/2 asserting
mousedown's detail; a chord test asserting a mid-chord primary click's
buttons mask. Two more tests were added independently during review
audit and kept as-is: clickCount 2 on a press+release pair asserting
detail 2 on mousedown/mouseup/click plus dblclick, and a chorded press
asserting its own mousedown branch also carries the press message's
clickCount.
Confirmed against live Chrome (headless, raw CDP) and Firefox (headless,
WebDriver BiDi) for both the mousedown-detail and chord-buttons fixes.
zig fmt --check clean; cdp.input 20/20 and bidi.input 5/5 pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A second independent Codex audit on the rebased bdabfea42 found a real
regression the chord fix itself introduced: the chorded-press branch in
triggerMousePress dispatched the compatibility mousedown but discarded its
cancellation result and never called focusForMouseDown, unlike the
first-button path. Pressing a second button on a different element while
the first is still held (e.g. right-click a second field while holding
left on the first) left focus on the original element instead of moving
it to the new mousedown's target, diverging from real Chrome.
Independently verified before committing:
- Traced the diff: the chorded branch's `_ = try dispatchMouseEventOn(...)`
discarded the return value entirely, so suppress_focus was never
computed and focusForMouseDown was never reachable from that branch —
confirmed this matches the first-button path's own
`if (!press.suppress_mouse and !press.suppress_focus) try
focusForMouseDown(...)` structure, which the chord branch should mirror
but didn't.
- Reverted the one-line fix (kept the new test staged) and confirmed the
new "chorded mousedown focuses its target unless pointerdown or
mousedown was cancelled" test fails against the pre-fix code, then
restored it.
- Ran the full suite: zig build test and -Dwpt_extensions both 1521/1521.
- Rebuilt the binary and ran the extended tools/shared-click-audit.mjs
--assert-chord --assert-chord-focus against it: chord_focus_contract
PASS, with the raw event trace confirming focus actually landed on the
second element ("ticket").
- Ran the same harness with --chrome: chord_focus_contract PASS against
real Chrome too, independently confirming this is the correct target
behavior and not just the audit's claim.
Fix mirrors the existing first-button path exactly: capture the chorded
mousedown's own cancellation result and run focusForMouseDown when it
wasn't cancelled, still gated by the gesture-level suppression flag so a
cancelled initiating pointerdown still suppresses every mousedown in the
chord, as before.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XLXnBHBxQNskg2MAke3Lhv
An independent audit of the shared click-dispatch refactor (Codex, on
b2cd5a190) reproduced a real defect: triggerMousePress/triggerMouseRelease
treated every button press/release as its own complete gesture, with no
notion of another button already held. Pressing a second button while the
first was still down (a chord) fired a second pointerdown with the wrong
buttons mask (single-button bitmask, not the aggregate), and
Page.input_mousedown_suppressed — a bare bool — got overwritten by the
second press's own suppression outcome, discarding whether the gesture's
initiating pointerdown had actually been cancelled. The eventual release of
the first (cancelled) button then incorrectly fired mouseup. Verified
against real Chrome 152's behavior for the identical input (one pointerdown,
button changes as pointermove, one final pointerup, no compatibility
mousedown/mouseup once the initiating pointerdown is cancelled) per
https://www.w3.org/TR/pointerevents3/#chorded-button-interactions.
Fixed by adding Page.input_pressed_buttons (an aggregate mask) and
dispatching pointerdown/pointerup only at its 0/nonzero transitions;
a button change while another remains held fires pointermove instead, and
input_mousedown_suppressed is now set once at gesture start and held for
the whole chord rather than being overwritten per press. Scoped to
triggerMousePress/triggerMouseRelease only — dispatchPointerPress/Release
(used by actions.click and WebDriver.click, which can't chord) are
unchanged. BiDi's input path calls the same two functions, so it gets the
fix for free.
One wrinkle caught while fixing: an existing test dispatches two
mousePressed calls on the *same* button with no release between them (to
test focus in isolation, not a real chord). Keyed the fresh-gesture check
off "some *other* button already held" rather than "any button held" so
that pattern still starts a fresh gesture, matching its own expectation.
New test: "a mouse chord fires pointermove for the mid-gesture button
change, not a second pointerdown/pointerup" on a new #btnChord fixture
element (mcp_actions.html). Confirmed it fails against the pre-fix code
(TestUnexpectedResult, verified by stashing the fix, running the test, and
restoring) before trusting it.
Verification: zig build test and zig build test -Dwpt_extensions:
1517/1517 both ways. zig fmt --check and git diff --check clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XLXnBHBxQNskg2MAke3Lhv
An automated review pass (Grok) on the shared click-dispatch diff flagged
comments that narrated the refactor's history ("for the first time",
"deliberately-untouched", "shares its dispatch mechanics with...") instead
of stating the invariant a reader needs. Rewrote four: WebDriver.click's
never-fails contract, actions.click's focus-error policy, the mousedown
click-count comment (also fixed to attribute the gap to triggerMousePress's
signature, not a CDP-only quirk — bidi/input.zig hits the same gap), and
the two new CDP regression tests' descriptions.
One suggested fix (clear input_mousedown_suppressed defensively before
dispatchPointerPress) was reviewed and rejected: suppress_mouse=true is a
constant once computed, with no fallible call between that assignment and
return, so the flag can only be lost on a throw when suppress_mouse=false —
which is the value already held by design. The scenario Grok describes
additionally requires an unpaired mousePressed (protocol misuse), not a
leak in this code.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XLXnBHBxQNskg2MAke3Lhv
Follows up on the maintainer's review note on PR #3431: WebDriver.click
(testdriver), Frame.user_input.triggerMousePress/Release (CDP's
Input.dispatchMouseEvent, i.e. Puppeteer/Playwright), and actions.click
(MCP) each dispatched their own near-identical pointerdown/mousedown/
pointerup/mouseup/click sequence. Adds three shared functions to
frame/user_input.zig -- dispatchPointerPress, dispatchPointerRelease,
dispatchClickAsPointer -- and routes all three call sites through them.
This gives the CDP path pointerdown/pointerup for the first time (it
previously only fired bare mousedown/mouseup/click), and a PointerEvent
click (previously a plain MouseEvent there). It also gives WebDriver.click
suppress/focus handling it never had: that function used to dispatch its
fixed five-event sequence unconditionally, ignoring preventDefault() and
never moving focus.
Because CDP's mousePressed and mouseReleased arrive as two independent
Input.dispatchMouseEvent messages with no shared call stack, a new
Page.input_mousedown_suppressed field carries whether the press half's
cancelled pointerdown should suppress this gesture's mouseup on the
release half. It's read-and-reset unconditionally at the top of
triggerMouseRelease (and reset on a press that finds no element), so an
unmatched or missed message can't leak stale state into the next gesture.
dispatchPointerPress returns PressResult{suppress_mouse, suppress_focus}
rather than running the focus default action itself: focusForMouseDown
can fail, and the three callers don't agree on what that should mean for
the click (actions.click: warn and continue; WebDriver.click and CDP:
propagate), so each runs it against the result with its own handling.
WebDriver.click also now reads frame._page.input_modifiers so a held
modifier key still reaches its dispatched events, matching its own
pre-existing local helpers' behavior (only compiled under
-Dwpt_extensions; actions.click and CDP don't track modifier state, so
they pass an empty Modifiers{}).
WebDriver.actionSequence's performPointerSource (a fourth, more complex
copy -- click counts, drag chords, touch) is deliberately left alone, as
is its own pre-existing gap (no pointerdown-suppresses-mousedown there).
Two new CDP tests reuse the existing mcp_actions.html fixture (#btn
records the full event sequence; #btnPreventDefault's pointerdown
listener calls preventDefault()) to pin the new pointer events and the
cross-message suppression. Both were confirmed to fail against the
pre-refactor triggerMousePress/Release bodies. Extended #btn's recorder
with event.detail after an independent review caught mousedown/mouseup's
click count silently dropping to 0 at all three call sites in an earlier
version of this change; the MCP click test's assertion was updated to
match.
zig build test and zig build test -Dwpt_extensions: 1516/1516 both ways.
zig fmt --check clean on all seven changed files.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XLXnBHBxQNskg2MAke3Lhv
Follow up to a chain of iframe context correctness (3520, 3510, 3501). Every
caller of StyleManager now must make sure the they call use the correct
StyleManager for a given Element. The StyleManager enforces this correctness
with a debug-only assertion.
It's tempting to think that this could be handled internally by the
StyleManager. It would be a lot cleaner..it can do the element -> ownerFrame
lookup. The issue is with frameless elements/documents which the StyleManager
cannot handle: every caller needs to decide how to handle this case.
- window.devicePixelRatio: turn from a constant property into an accessor
mirroring innerWidth/innerHeight. Reads the page viewport scale (set via
Emulation.setDeviceMetricsOverride's deviceScaleFactor, default 1.0) and
remains [Replaceable] through a setter that delegates to replaceGlobalProperty.
- Input.dispatchKeyEvent:
- For `char` events, insert `params.text` into the focused element when
not default-prevented by a keypress listener. Non-text keys without a `text`
payload (e.g. Enter) leave the value untouched.
- For `keyDown` events, fall back to `params.text` when `params.key` is
omitted, allowing text-only keyDown dispatches from CDP clients to insert
the character into the focused element.
Blink does not dispatch a second event for the legacy name. Per target,
it runs the wheel listeners if there are any, else the mousewheel ones
with the event retyped for the call. A target registering both never
sees mousewheel. The event manager now does the same for trusted
events, so one wheel dispatch covers both names and the fallback also
decides what counts as a non-passive listener on the path.
This feature is significant because it adds support for processing an HTTP
request via the worker. It requires parking the connection and then having the
worker notify the loop when the response is ready. A lot of this was already
in-place (e.g. worker -> loop notification) but not quite do this extent.
Requires: https://github.com/lightpanda-io/zig-v8-fork/pull/207
Inspired by https://github.com/lightpanda-io/browser/pull/3504 this simplifies
v8::Value serialization (e.g. as used in console.log(...)).
1 - It doesn't executes JS and thus can't have a side effect, which is otherwise
possible if we invoke a getter or through a proxy
2 - It removes the debugValue debug-only path
(2) is potentially a loss in debug builds, but I think the usefulness of that
was always, at best. The upside is code elimination and consistency in how
values are reported in debug/release
ScrollResult's container arm carried the node the caller had passed in,
so both consumers re-derived what they already held. The payload is now
just the scroller, the CDP handler folds two arms into one, and the tool
formats the element from its own argument instead of re-registering it.
Element.scrollContainer returns on the first hit for callers positioning
one scroller; scrollContainers keeps the per-axis walk for wheel. That
drops ScrollTargets.nearest, which only one caller read and which made
the walk keep climbing for an axis that caller discarded. ScrollTarget
gains scrollBy, so the wheel path no longer needs a free function to
pick between an element and the window.
StyleManager folds declarations through the existing declaration
iterator rather than the linked list, so Property.fromNodeLink goes back
to private. The overflow shorthand re-enters Slots.apply with its
longhand names instead of a comptime slot lookup, and both fold loops
are inline so the property name matching folds away.
The listener-list walk moves to EventManagerBase next to getListeners,
where it also skips removed listeners like findListener does. The
propagation walk takes a comptime type, so the lookup key is built once
instead of per node.
WebDriver's wheel passes the action's own coordinates through rather
than synthesizing them from a bounding rect, and its touch dispatch uses
the owner frame it already resolved.
Element.scrollContainers walks the ancestor chain once for every
requested axis and returns a ScrollTarget per axis, viewport or
container, plus the nearest of the two. The wheel path no longer walks
twice on a diagonal wheel, and the tool's "no container means the node
itself" policy is a visible switch arm rather than an orelse on null.
actions.ScrollResult names what scrolled as a tagged union: the window,
the given node, or its container together with the node. The tool and
LP.scrollNode format the result without touching the request.
CDP/BiDi wheel and WebDriver wheel each had their own definition of what
a wheel does. WebDriver computed cancelability from listener passivity
and fired Blink's legacy mousewheel; the CDP path fired a single always
cancelable wheel. Only the target lookup was legitimately different.
user_input.wheel now owns the sequence: wheel, mousewheel, then the
scroll unless either was canceled, each non-cancelable when every
listener on its path is passive. The passivity query moves onto
EventManager, next to the listeners it inspects. Callers supply the
target and the deltas.
For CDP this means a page's document-level wheel listener can no longer
cancel scrolling unless it opts out of the default-passive behavior, and
mousewheel listeners now fire, both as in Chrome.
Element.scrollContainer read the inline style= attribute only, so a
scroller declared in a stylesheet was invisible to the scroll tool and to
wheel scrolling, which then fell through to the viewport.
StyleManager now tracks overflow-x and overflow-y alongside display,
visibility, opacity and pointer-events, and exposes scrolls(el, axes) as
an own-element probe. The overflow shorthand is expanded into its
longhands in declaration order, in both the attribute scan and the
materialized style object, so a shorthand and its longhands keep the
precedence of the source text. overlay counts as auto, as in Chrome.
Element.scrollContainer asks the style manager, and the two unused Props
bits hold the new flags, so the per-element memo does not grow.
Follow up to https://github.com/lightpanda-io/browser/pull/3510
Moves the element/node lookups, e.g. `element_class_lists` from Frame to Page.
Elements and nodes can outlive a Frame (it's the reason the identity map lives
on the Page, not the frame). These maps are merely properties on Node/Elements
optimized for a specific usage-pattern (i.e. most Node/Elements don't have these
or they are never materialized from JS). So if a Node/Element can outlive the
Frame, than so too can all of their properties. And, even when an frame is alive
the properties belong to the *Node* or *Element*, NOT the Frame...accessing
those properties across frames should yield the same value / identity.
More mechanically, frame._page => frame.page and all of these lookups lose their
_ prefix. Short summary of _ prefix is:
1 - It's used to deal with Zig not allowing shadowing. This is particularly true
in the WebApis were it happens a bit more often
2 - Early prototype was built as a stand-alone library, and the _ was used to
signal "private" (again, working around Zig). Frame.page shouldn't be
"private" and neither should these lookups (if we aren't going to provide
getter/setters for them).
The scroll tool (MCP, agent, LP.scrollNode) wrote scrollTop on the exact
node it was given, so a leaf inside an overflow:auto panel stored an
offset on a non-scroller, the panel's own scroll listener never ran, and
the tool reported the requested coordinates as if it had worked. It also
fired a synchronous bubbling scroll on top of the async non-bubbling
scroll/scrollend the setters already schedule.
actions.scroll now resolves the nearest ancestor-or-self scroll
container, falls back to the node itself, and returns the node that
moved plus the read-back position. The tool and LP.scrollNode report
that instead of the request.
The container query moves from user_input.zig onto Element as
scrollContainer(axes), so the wheel path, the tool and WebDriver share
one resolver. WebDriver's wheel scrolled the hit-test element directly
and fired its own bubbling scroll; it now goes through
user_input.wheelScroll like CDP and BiDi wheel.
Window and Element share one ScrollToOpts. Its offsets() helper
normalizes the positional and dictionary forms once, and an omitted axis
in the dictionary form leaves that axis untouched for the window too,
matching browsers, so scrolling the window on one axis no longer resets
the other.
Runtime.consoleAPICalled serialized object arguments with JSON.stringify,
which can run getters and toJSON callbacks. A callback can log again or
navigate, emitting further CDP events while the outer event is being built.
Resetting send_arena after each send and notification_arena after each
handler then invalidates the outer message's buffers. This produces
use-after-free in debug builds or malformed JSON that disconnects clients.
Scope both arenas to the outermost send or notification handler, including
inspector messages and error paths. Also match Chrome's console argument
representation: objects remain remote handles, primitives carry value,
and non-JSON numbers and bigints use unserializableValue. Console logging
must not invoke object getters or toJSON as a serialization side effect.
Test complete nested WebSocket messages, failure recovery, nested legacy
Console notifications, and primitive/object protocol shapes.
There's been recent work on improve select / options:
- https://github.com/lightpanda-io/browser/pull/3375
- https://github.com/lightpanda-io/browser/pull/3402
- https://github.com/lightpanda-io/browser/pull/3499
One of the main issues is that a select's options "selected" state wasn't always
kept in sync. Select.zig had a boolean flag to mark whether or not a
selectedIndex was explicit set and every read and update would need to dance
around it. Removing the selectedIndex option would not, for example, keep things
in sync.
This removes the flag and keeps the Option._selected in sync, i.e. the sync
happens on write, not on read and is thus naturally recorded in the state of
the Select and its Options.
The write path is more complicated, but the read path is simpler (though the
real win is always being correct).
We've auto-injected `*Frame` into WebApi since forever (used to be called *Page,
but then we split *Page / *Frame, but same same). And it worked wonderfully:
there was always a single *Frame, so the Frame/Context that the JS was being
executed in HAD to be the *Frame that a node belonged to.
But with the addition of iframe and popups, that truth no longer holds. The
*Frame executing the JS (which is the frame that we auto-inject) isn't
necessarily the *Frame that owns a Node.
This is particularly problematic because the *Frame holds a bunch of node/
element data, e.g. `_element_datasets`. So now the DataSet that you get back
depends on the context in which its called..they don't have identity and can
fall out of sync. Some code calls node.ownerFrame() / node.ownerDocument(), but
not all and the hope is to address this throughout the codebase once and for
all.
This is the first in a series of commits meant to fix this long-standing
issue. All it does is change the node.ownerFrame() return value from *Frame to
?*Frame. It's up to each caller to decide how to handle a frameless node, e.g.
clicking a frameless link should not navigate.