Saw a site return "Content-Type: charset=UTF-8;charset=UTF-8" on scrape. Not a
valid header and it correctly fails to parse. Except...CDP fails to RI on these
since it tries to parse and bubbles. Safe to swallow this error here and
continue with the response serialization.
Ultimate goal is to improve the logging to be more useful by:
1 - identifying the page being navigating
2 - downgrading most page-loading errors to debug
3 - provide page statistics (in logs, via cdp).
So, the logs would look something like:
```
info page-navigate page=3392 url=https://.....
debug urlblocked page=3392
debug timeout page=3392 url=script.js
info page-done page=3392 url=https://....
```
The log level for page-done will be the @max() of any log level for that page
so that if you run --log-level warn, and you get:
```
info page-navigate page=3392 url=https://... <-- filtered out
debug timeout page=3392 url=script.js <--filtered out
warn unknown module lookup page=3392 identifier=blah.js
warn page-done page=3392 url=https://....
```
So that any page event can be linked to the top leveL URL, while still
respecting the --log-level.
This first parts adds the log-level page context, and the thread local
variable.
Resolving the cascade for every element on each mutation made an
append + scrollHeight read up to 12x slower. Virtualizers set their
spacer height inline.
If the server gives a status reason, report it as-is. Only default to the
zig code->reason map when one isn't given.
Also, don't force a 407 status code when auth_challenge is present.
https://github.com/lightpanda-io/browser/pull/3619 added proper support for
Sanitizer. A consequence of that is that every setHTML/parseHTML without an
explicit creates the default Sanitizer (with its ~300 entries).
This commit creates 1 app-level default sanitizer and uses it, internally, when
none is explicitly given. The sanitizer is immutable so can safely be used
across threads.
Adjacent placeholder anchors, like a JS-driven nav, ran together as
HomeAbout. Keep the standalone line placement linked anchors get and
drop only the link syntax.
`lightpanda version.io` was rejected as a misspelt `version`, 3 edits
within the length-scaled limit. Like curl, it fetches now: the command check
shares --dump's isUrlLike.
Following some experimentation, this sets M_MMAP_THRESHOLD to 128K on linux
builds IF it isn't explicit set, e.g. via the `MALLOC_MMAP_THRESHOLD_` env. This
allows users explicitly set it if they want.
I've been looking at this for a while, but my understanding is still basic. The
simplistic explanation is that glibc's allocator internally uses arenas.
Allocations above THRESHOLD skip these arenas and are allocated directly from
and freed directly to the OS. Allocations under this threshold are released
to the arena where they can be re-used. Everything works OK, EXCEPT for when it
comes time to release the memory from these arenas back to the OS. If free arena
memory is pinned under live memory, it cannot be released.
Various online posts / comments reference glibc as suffering from significant
fragmentation and "holding" onto memory. That's the behavior that we've seen.
Further, calls to malloc_trim reclaim the lost memory, so what we're seeing
aren't leaks.