The multiline line reader returned EOF with buffered data, so record
parsing discarded the last row when the file had no trailing newline.
Clear EOF while data remains, as the normal line reader does.
Add regressions for final records, empty input and unterminated quotes.
The normal core suite passes 627 tests with the new CSV tests included.
core:encoding/base64 can only produce padded output, so callers that need
the canonical unpadded ("raw") form -- JOSE base64url, Go's
RawStdEncoding/RawURLEncoding -- have to encode and then trim, which
forces an extra allocation and cleanup dance around the returned string.
Add an Encode_Options bit set, mirroring Decode_Options, and honor
{.No_Padding} in `encode`, `encode_into_buf`, `encode_into`, and
`encoded_len`. Existing calls and padded output are unchanged.
Tests cover the RFC 4648 section 9 illustrations and section 10 test
vectors across the standard and URL alphabets, padded and raw, plus Go's
reference corpus. Newline characters are asserted to be rejected per
RFC 4648 section 3.3. Comments in `encode_impl` document the 24-bit group
packing and the padding decision for partial groups.
core:encoding/base64 currently decodes leniently: it infers padding from the
last one or two bytes, accepts missing/extra/interior padding, and does not
check the trailing padding bits required by RFC 4648 Section 3.5.
Go exposes canonical decoding as `Encoding.Strict()`; Odin has no equivalent,
so downstream ports (e.g. age) have to reimplement base64 from scratch.
This adds opt-in strict decoding through a `Decode_Options` bit set.
Existing calls and behavior are unchanged.
unquote_string replaces each byte that is not valid UTF-8 with U+FFFD,
which is three bytes for one, but sized its buffer as len(s) + 2*UTF_MAX:
slack for a single replacement, not for one per invalid byte. A string
holding several ran the write cursor past the end, an out-of-range slice
under bounds checking and a memory-safety bug without it.
Count the invalid bytes in the remainder up front and size for them. The
escape sequences never grow their input, so they need no allowance.
parse_object_body allocates an object key, then may fail in parse_colon or
parse_value before that key is ever inserted into the object. Its cleanup defer
only walks `obj`, so a key that never got there is unreachable to it. The caller
cannot free it either -- a failed parse returns a nil Value -- so it leaks.
The same applies to the parsed element on the duplicate-key path, and to both on
the out-of-memory path.
JSON5 makes this reachable from ordinary malformed input, because an unquoted
ident is a legal key and anything other than a colon after it fails. Plain JSON
leaks it too, via a quoted key.
before, measured with a tracking allocator over 8 inputs x 2 specs:
LEAK JSON5 colon fails after unquoted key 1 alloc / 7 bytes
LEAK JSON colon fails after quoted key 1 alloc / 2 bytes
LEAK JSON5 colon fails after quoted key 1 alloc / 2 bytes
LEAK JSON value fails after key 1 alloc / 2 bytes
LEAK JSON5 value fails after key 1 alloc / 2 bytes
LEAK JSON nested value fails 2 alloc / 4 bytes
LEAK JSON5 nested value fails 2 alloc / 4 bytes
LEAK JSON deep nesting fails 3 alloc / 6 bytes
LEAK JSON5 deep nesting fails 3 alloc / 6 bytes
LEAK JSON array element fails 1 alloc / 2 bytes
LEAK JSON5 array element fails 1 alloc / 2 bytes
total leaked allocations: 17
after, same probe:
total leaked allocations: 0
The leak scales with nesting depth -- one orphaned key per enclosing object -- so
a service parsing untrusted JSON leaks a little on every malformed request.
The fix marks the key and the element as owned by the loop iteration until they
are stored, and frees them otherwise. The duplicate-key path loses its explicit
delete, which the same mechanism now covers.
Found via odinfmt, which reported a 7-byte leak in a downstream test that parses
`{ broken not json` to check that invalid input is rejected.
Regression test added to tests/core/encoding/json: it reports
`17 leaks and 0 bad frees` without this change and passes with it. The existing
11 tests pass unchanged under -define:ODIN_TEST_FAIL_ON_BAD_MEMORY=true.
* json: fix user unmarshaller example
- Returning `.None` in the custom unmarshaler is wrong, should be `nil`
- `advance_token` has to be called
Besides the fixes I made it an actual example that will show up on the package docs
Also had to vendor `core:encoding/ini` into `core:os/os2` for the user directories on *nix,
as it used that package to read `~/.config/user-dirs.dirs`, causing an import cycle.