A BCP 47 tag in LC_ALL is not a POSIX locale, so setlocale(LC_ALL, "")
failed for the rest of the process and for any child. In the REPL,
isocline took the terminal for non-UTF-8 and dropped every non-ASCII
keystroke. Set ICU's default locale through the new
v8__V8__SetDefaultLocale binding and leave the C library alone.
Pins zig-v8-fork to lightpanda-io/zig-v8-fork#209; CI links once that is
tagged and action.yml's zig-v8 is bumped.
/save kept a model's backendNodeId call by swapping in the selector the
tool layer resolved, but a typed `/scroll backendNodeId=3` never asked
for one and was dropped from the script. Ask for it on this path too.
Recording swaps a call's backendNodeId for the selector the tool layer
resolved, but scroll had no selector parameter: the replayed
scroll({ selector, y }) dropped the field and scrolled the window. Take
a selector as click and hover do.
`call` reports a failed navigation in-band, as an is_error result, so
the catch in gotoStart never saw it: a start page that could not
resolve left the agent on a blank page with no message.
isocline copied the submitted line into a buffer one byte short whenever
the locale is not UTF-8. A line whose length the allocator serves
exactly, such as the 24-byte `/screenshot path=out.png`, corrupted the
heap on Enter.
Picks up the per-effort `max_tokens` floor, so a tight budget no longer
buys mostly thinking and a truncated answer, and the search clients
logging a non-2xx like the model clients already did.
The bump also carries zenai's `ErrorDetail.body` -> `.message` rename,
which the search failure path reads.
Review pass over the branch.
`selectorForArgs` ran for any call carrying a `backendNodeId`, before
`isRecorded` was consulted. `SelectorPath.build` is the expensive part of
a tool call -- uncached full-document matches per ancestor -- and the
read-only tools that take an id (tree, markdown, html, nodeDetails) are
exactly the ones the system prompt steers toward ids, so the common case
built a selector and threw it away. `nodeDetails` built it twice.
The recorded selectors were duped into `self.allocator` and never freed:
`clearRetainingCapacity` drops the only pointers. They belong to the
conversation arena, which already owns the args and the `Command` built
from them, so the ownership question disappears rather than being
patched with a free loop.
`searchKeyStatus` returned null for `.auto`, which is the default -- so
the keyless-endpoint notice never fired in the one configuration that
reaches a keyless endpoint. It now reports the rung the cascade lands on.
`gotoStart` was inserted between `printUsageSummary`'s doc comment and
the function, leaving the `$usage` wire format -- which wrapper scripts
grep -- documenting the wrong function. It also returned `?ToolError`
where the rest of the file uses error unions.
Plus: the three search collectors differed only in a field name, the
failure writer had five `catch return OutOfMemory` in ten lines beside a
neighbour doing it with one, `visitAll` was `pub` with no caller outside
its file, and the tests named themselves after a function that no longer
renders anything.
`/save` synthesis never looked at `finish_reason`, and `stripCodeFence`
accepts a fenced block with no closing fence. A script cut off at
`max_tokens` therefore looked exactly like a complete one: it was written
to disk, the save path was remembered, the buffer was reset and the
terminal said "Saved synthesized script to ...". The truncation only
turned up on replay.
It now aborts like any other failed synthesis, leaving the previous file
and the buffer alone.
`reset` clears the lookup maps and deliberately leaves `node_id` where it
is, so a stale id fails closed with a miss instead of resolving to
whatever registers next. Nothing asserted that, and it is the invariant
every id-addressed tool call leans on across a navigation.
`jsonStringify` and `textStringify` each carried their own copy of the
same six-line preamble -- xpath buffer, listener target map, label index,
walk, error log -- differing only in the visitor and the wording of the
log line. Either one could drift from the other silently, since nothing
compares them.
`visitAll` takes the visitor and owns the walk; both dumpers are now one
line. Their existing tests cover it.
A `--task` run that knows where it is going still spent a model turn
navigating there, and the REPL had no way to start anywhere but blank.
`--url` opens the page first, through `browser_tools.call` rather than
around it, so a bad URL fails like any other tool call and the opening
navigation is recorded for `--save` -- which is the first line any
replayable script needs anyway.
The search engine could only be set from the REPL's `/searchEngine`, so a
one-shot run had no way to pin one -- and a benchmark that wants its
results to mean something has to record which API answered.
`/searchEngine` also carried the only warning about key state, which is
the half that matters more. A keyless engine is not an error and starts
fine, so a run that silently falls back to a rate-limited public endpoint
looks exactly like a working one until every search begins failing, and
then it looks like a bad agent. Both messages now fire wherever the
engine is resolved, not just from the command.
`resolveSearchEngine`'s doc comment said there was no CLI flag. There is
one now.
A failed search reported only the error name: "keenable search failed:
ApiError". The status code and the provider's own message were logged
once inside `apiSearch` and then died with the client on its deferred
deinit, so nothing downstream could tell a rate limit from a rejected key
-- and a model reading the tool result had nothing to act on.
`Failure` carries the status and body out before the client goes, and the
message now names both. A 429 additionally says it is a rate limit and
that waiting or reading the page instead are the ways out, because that
is the one failure where the right move is not "try another query".
Found by a benchmark run where a keyless search engine quietly hit its
hourly cap; every search failed for twenty minutes and the run just
looked like a bad agent.
The four per-engine markdown writers differed only in which response
field holds the title, the url and the snippet; the numbering, the
newline flattening and the "No results." case were copied four times.
Each engine now contributes a `collect` that maps its response onto
`[]Hit`, and one `renderResults` turns those into the markdown the model
reads. Adding an engine is a field mapping rather than another copy of
the renderer, and the intermediate `SearchResults` is a shape a caller
can use directly instead of re-parsing markdown.
Output is byte-identical -- the existing tests assert the exact strings
and are unchanged apart from calling the new pair.
`--save` drops any tool call that names its element by `backendNodeId` and
carries no selector (`Command.ToolCall.isRecorded`), because a registry id
means nothing in a later session. That is a silent gap on the chat path:
whenever the model addresses a node by id, the action runs and is then
quietly omitted from the saved script.
The tool layer is the one place that already resolves the node before
acting, so it is where the selector can be taken while the node still
exists -- a navigation takes it away a moment later. `CallOpts.record`
asks for it and `ToolResult.selector` returns it, for every tool, with no
per-tool knowledge anywhere. Recording then swaps `backendNodeId` for the
selector and keeps every other argument, so `scroll`'s `y` and
`waitForState`'s arguments survive.
One `save_selectors` entry per call, errors included, so the index lines
up with `RunToolsResult.tool_calls_made`.
TryCatchRethrow, JsException and ExecutionTerminated all mean V8 already
has something pending: creating an error value and aborting the stream
would replace the exception the script is meant to see, or hand a killed
script a catchable Error. Drop the streaming handle instead, matching the
early exit in Caller.handleError.
Element.focus() would "focus" the element even when it shouldn't. We already
have the logic to determine if an element is focusable in `user_input.zig`, so
this was moved to Element and is now used in el.focus().
Some status-codes should never have a body except for a single trailing blank
line. If we don't handle these, then we end up with a dirty connection in our
connection pool:
1 - read the header, but not the body
2 - put the connection back in the pool
3 - try to read the header, but actually get the body from #1
WPT /fetch/api/basic/response-null-body.any.html exercises this path and is
flaky (because it depends whether the request goes back out on a keep-alive
connection)..but for a given run,you'll almost always get 1-3 failures.
This commit processes the request, but tells libcurl not to re-use the
connection.
The http max default was 4K with a 16KB hard limit. The default limit is now 1MB
with an initial default of 4K. This is to accommodate larger WebDriver payloads.
1 - Centralized cache-awareness into Cache and pulled header details out of
SqliteCache and HttpClient
2 - Added support for expires header
3 - Support caching more status types (but not all, since HttpClient would need
to be aware of what caching a 3xx/206 means)
4 - Revalidate cares about "not specified" vs "no-store" vs "stale"
(e.g. expires=0 means "stale", not fallthrough the last-modified logic)
Responses with specific status (e.g. 204) should always have a null body. Also
adds validation to Response constructor (e.g. can't provide an invalid status).
Improves a handful of WPT tests:
fetch/api/response/response-error.any.html
fetch/api/response/response-static-json.any.html