rt_malloc was a bare mmap-per-call wrapper: every allocation, even a 32-72B AST/symbol node, consumed a page-rounded 4KB region and was never freed. w6a/w6c emit ~4 such nodes per .s line, so assembling a 145K-line file burned ~625K pages (~2.5GB); test 995's 5 concurrent self-rebuilds then OOM'd. The defect is linear and str-independent -- the str->24B codegen merely enlarged .s files past the cliff. Replace it with the no-free SUBSET of Hare's allocator (ref/hare/rt/malloc.ha): 2MiB chunk-bump (CHUNKSZ malloc.ha:24, ALIGN malloc.ha:14), oversized (>CHUNKSZ) requests direct-mmap'd. The bin/freelist/META machinery exists only to support free, which ww does not have, so it is omitted. Policy lives in rt/malloc.ww; the raw mmap primitive stays in rt/alloc.s as rt_segmalloc -- Hare's malloc/segmalloc split. rt_free becomes a documented no-op (os.free re-exports it for the public API, so the symbol must stay); rt/ensure.ww drops its now-impossible reclaim. Zero-init is preserved: the bump never reuses memory, so every byte is fresh MAP_ANONYMOUS-zeroed. w6a_ww on a 145K-line .s: 2165MB -> 28MB (~glibc parity, C w6a 24MB). .o output byte-identical; both stages emit identical asm. make test 134/134.
45 lines
1.7 KiB
Plaintext
45 lines
1.7 KiB
Plaintext
// rt/malloc.ww — bump allocator over 2MiB chunks. Archived into
|
|
// libwwrt.a, compiled standalone (no `module`) with bare symbols, like
|
|
// rt/ensure.ww.
|
|
//
|
|
// WHY bump, replacing the page-per-call mmap (rt_segmalloc, ex-rt_malloc):
|
|
// the old rt_malloc page-rounded every request, so a 32-72B node cost a
|
|
// whole 4KiB page. w6a/w6c emit ~4 such nodes per .s line; a 145K-line
|
|
// assembly burned ~625K pages (~2.5GB) vs C w6a's 24MB. Packing nodes
|
|
// into shared chunks collapses that to the live set. (task #8)
|
|
//
|
|
// Mirrors Hare rt::malloc's chunk-bump (ref/hare/rt/malloc.ha:52-66,
|
|
// CHUNKSZ :24, ALIGN :14), minus the bin/freelist machinery — that
|
|
// exists only to support free, which ww does not have.
|
|
//
|
|
// Zero-init is preserved: every byte handed out is fresh from
|
|
// MAP_ANONYMOUS (kernel-zeroed) and never reused, so alloc(T{})'s
|
|
// zero-fill still holds. Lazy first chunk falls out of cur=rem=0.
|
|
|
|
def CHUNKSZ: u64 = 2097152u64; // 1<<21, ref/hare/rt/malloc.ha:24
|
|
def ALIGN: u64 = 16u64; // ref/hare/rt/malloc.ha:14
|
|
|
|
@symbol("rt_segmalloc") fn segmalloc(n: u64) *void;
|
|
|
|
let cur: u64 = 0u64;
|
|
let rem: u64 = 0u64;
|
|
|
|
@symbol("rt_malloc") export fn rt_malloc(n: u64) *void = {
|
|
let need: u64 = (n + (ALIGN - 1u64)) & ~(ALIGN - 1u64);
|
|
// Oversized requests bypass the chunk so one huge alloc can't
|
|
// strand most of a chunk (ref/hare/rt/malloc.ha:38).
|
|
if (need > CHUNKSZ) { return segmalloc(need); };
|
|
if (rem < need) {
|
|
let c: *void = segmalloc(CHUNKSZ);
|
|
// mmap failure: keep the nomem null contract that the
|
|
// alloc-builtin `!`/`?` lowering checks.
|
|
if (c == nil) { return nil; };
|
|
cur = c: u64;
|
|
rem = CHUNKSZ;
|
|
};
|
|
let p: u64 = cur;
|
|
cur += need;
|
|
rem -= need;
|
|
return p: *void;
|
|
};
|