Files
browser/orderfile/tools/gen_order.py
T
Karl Seguin ee1fd9207c mem: On linux CI builds, use an orderfile
CI build, for e2e-test and nightly now build with:
-Dorderfile/lightpanda.ld

This file informs the build on how to organize the code in the binary, grouping
hot code together so that we have to load less of the binary into memory.

lightpanda.ld will drift: we'll refactor our code, add new features, update
dependencies, update Zig, ... So it has to be re-generated. But we can do that
automatically in the CI (say, before the nightly build). That's for a follow up
PR.

This does not currently cover V8. V8 is being build with
`-no-unique-section-names`, so we don't get names that we can correctly
organize. The real win comes from doing this in V8, since a lot of V8 is cold.
This PR can land as-is, a zig-v8-fork PR will remove that flag, and then we can
have a follow up PR with an lightpanda.ld that includes the v8 symbols.

This is opt-in (via the -Dorderfile flag) because it adds ~20 seconds of
linking time.
2026-08-25 21:02:31 +08:00

59 lines
2.9 KiB
Python
Executable File

#!/usr/bin/env python3
"""usage: gen_order.py <hot.text> <hot.rodata> <out.ld> <obj-or-archive>...
Builds a symbol -> (file, section) map from the objects and emits an INSERT
linker script that places the hot sections in .text.hot / .rodata.hot ahead of
.text / .rodata. Patterns are scoped to their object file (`*api.o(...)`):
LLD tests every input section against every unscoped pattern, which turns a
26k-pattern script into an 80s link; scoped, it is a few seconds.
"""
import sys, subprocess, collections, re, os
hot_text, hot_rodata, out = sys.argv[1:4]
objs = sys.argv[4:]
loc_of = collections.defaultdict(set) # symbol -> {(file, section)}
for o in objs:
p = subprocess.run(["objdump", "-t", o], capture_output=True, text=True).stdout
fname = os.path.basename(o)
for line in p.splitlines():
m = re.match(r"^(.+?):\s+file format elf", line)
if m:
fname = os.path.basename(m.group(1))
continue
# value flags section<TAB>size name -- the section can contain spaces (zig)
if "\t" not in line: continue
left, right = line.split("\t", 1)
parts = right.split(None, 1)
if len(parts) != 2: continue
name = parts[1].strip()
for vis in (".hidden ", ".protected ", ".internal "):
if name.startswith(vis): name = name[len(vis):]
fl = left[17:24]
if "d" in fl: continue # section symbol
sec = left[24:].strip()
if sec.startswith((".text", ".rodata")):
loc_of[name].add((fname, sec))
def esc(s):
return re.sub(r'([*?\[\]\\])', r'\\\1', s)
stats = collections.Counter()
def emit(hotfile, prefix):
by_file = collections.OrderedDict(); seen = set()
for name in open(hotfile).read().split("\n"):
if not name: continue
locs = loc_of.get(name)
if not locs: stats[prefix + " nomap"] += 1; continue
for fname, sec in sorted(locs):
if not sec.startswith(prefix): continue
if sec == prefix or sec.endswith("."): stats[prefix + " generic"] += 1; continue
if '"' in sec or '"' in fname: stats[prefix + " quote-skip"] += 1; continue
if (fname, sec) in seen: continue
seen.add((fname, sec)); by_file.setdefault(fname, []).append(sec)
stats[prefix + " sections"] = len(seen); stats[prefix + " files"] = len(by_file)
# LLD unquotes section names but not the file pattern, so that one stays bare.
return [f' *{esc(f)}(' + " ".join(f'"{esc(s)}"' for s in secs) + ")" for f, secs in by_file.items()]
t = emit(hot_text, ".text"); r = emit(hot_rodata, ".rodata")
with open(out, "w") as f:
f.write("/* Generated by orderfile/tools/gen_order.py, see orderfile/README.md. */\n")
f.write("SECTIONS {\n .text.hot : {\n" + "\n".join(t) + "\n }\n} INSERT BEFORE .text;\n")
f.write("SECTIONS {\n .rodata.hot : {\n" + "\n".join(r) + "\n }\n} INSERT BEFORE .rodata;\n")
for k, v in sorted(stats.items()): print(f"{k}: {v}")