wcc,ww,os: atomic pkgcache store via temp+rename, both stages (#104)

The out/.pkgcache content-keyed store copied each artifact IN-PLACE
(cp -f / copyfile) to the fixed paths P.wwi/P.o/P.key. Key-last gave
crash-consistency but NOT concurrent-read safety: two same-stage builds
of a shared lib pkg (rt/time/os) target one out/.pkgcache/<pkg>/P.{wwi,o};
once an early finisher writes P.key, a later build's cache_lookup copies
P.wwi/P.o while a mid-finisher is still mid-write -> torn read -> corrupt
link / cs!=ww. The key is content-only, so it is purely the non-atomic
write.

Fix (Go-build-cache pattern, both stages in lock-step, rule 10): write
each artifact to a per-pid same-dir temp (P.wwi.tmp.<pid> etc.) then
rename() into place. Same dir => rename is atomic (cross-fs is not);
per-pid temp => concurrent writers don't clobber each other mid-copy;
content-keyed => last-writer-wins is byte-identical. Key renamed LAST so
a reader that sees the new key always finds complete artifacts. On any
mid-store error the per-pid temps are unlinked so a failed store leaves
no litter (cstage goto cleanup; wwstage cachermtmp helper).

  cstage cmd/ww/main.c cache_store: libc rename(2) + getpid().
  wwstage selfhost/cmd/ww/main.ww cachestore: new os.rename + cachetmp.
  lib/os/os.ww: add rename(2) (RENAME=82), ref/hare/os/os.ha:17 -- returns
    raw i32 errno like sibling remove/mkdir/rmdir (ww's os is the flat
    syscall floor, no fs:: layer); a second pathbuf2 slot holds newpath
    since kpath's single pathbuf can't carry both paths.

cache_lookup is unchanged: it reads cache->private scratch, and an atomic
source is never torn.

The torn-read race is closed BY CONSTRUCTION; a deterministic behavioral
regression-guard isn't feasible through the product build path (content-
keying => concurrent COLD builds all MISS+STORE, never HIT-read a mid-store
entry; a warm cache is never re-stored). The deferred white-box guard is
TASK #105. A WHY-comment at both fix sites records this.

Tests: 989_sepbuild_run KEEPS its private per-pid WW_PKGCACHE -- the
comment is corrected: the pin is NOT a torn-read mask (closed by
construction) but cold-compile isolation for the test's INTERMEDIATE
(.s/.unit.ww) byte-id compare, which a cache HIT legitimately skips
producing. The former 989_pkgcache_atomic_run is renamed to
989_pkgcache_concurrent_run and HONESTLY relabeled: it is a concurrent
shared-cache build-correctness smoke (N concurrent --sep builds sharing
one cache -> every binary byte-identical to an isolated reference + correct
run, both stages), NOT a torn-read/atomicity proof (a review revert-
experiment proved the original claim vacuous). Shrunk to 4 concurrent
builds x 1 batch x both stages. COLD/dev-only, off every byte-id/bootstrap
gate.

selfhost/cmd/ww/main.combined.ww remains stale (its writer was deleted at
the M4 E3-C1 flip; #90 deletes the file) -- not regenerated.

make test: all 445 passed; make sizelint clean; 990-997 byte-id hold.
This commit is contained in:
2026-06-18 20:20:37 +09:00
parent 33edc386f1
commit a9778ec000
6 changed files with 383 additions and 23 deletions

View File

@@ -809,28 +809,56 @@ cache_lookup(struct sepgraph *g, int pi, const char *manifest,
return 1;
}
/* On MISS, persist the freshly compiled artifacts then the manifest. The key
* is written LAST so a crash mid-store never leaves a key whose artifacts are
* absent/partial (the next run simply re-misses). */
/* On MISS, persist the freshly compiled artifacts then the manifest. Each is
* copied/written to a per-pid same-dir temp then rename()d into place: rename
* is atomic within one filesystem (cross-fs is not), so a concurrent
* cache_lookup never observes a half-written P.wwi/P.o/P.key (#104). The
* per-pid temp name keeps two concurrent writers from clobbering each other
* mid-copy; content-keyed ⇒ last-writer-wins is byte-identical. The key is
* renamed LAST so a reader that sees the new key always finds complete
* artifacts, and a crash mid-store never leaves a key without them.
*
* The torn-read race is thus closed BY CONSTRUCTION. A deterministic
* behavioral regression-guard isn't feasible through the product build path:
* content-keying means concurrent COLD builds all MISS at lookup and STORE —
* none HIT-reads a mid-store entry — and a warm cache is never re-stored, so
* "store concurrent with a HIT-read of the same entry" can't be forced. The
* deferred white-box guard is TASK #105; 989_pkgcache_concurrent_run smokes
* that concurrent shared-cache builds stay correct. On any mid-store error the
* per-pid temps are unlinked so a failed store leaves no litter. */
static void
cache_store(struct sepgraph *g, int pi, const char *manifest,
const char *wwi, const char *obj)
{
char dir[1024], keyp[1100], cwwi[1100], cobj[1100], cmd[4096];
char twwi[1200], tobj[1200], tkey[1200];
int pid = (int)getpid();
FILE *f;
pkgcache_dir(g, pi, dir, sizeof dir);
snprintf(cmd, sizeof cmd, "mkdir -p '%s'", dir);
if (run(cmd) != 0) return;
snprintf(cwwi, sizeof cwwi, "%s/P.wwi", dir);
snprintf(cobj, sizeof cobj, "%s/P.o", dir);
snprintf(keyp, sizeof keyp, "%s/P.key", dir);
snprintf(cmd, sizeof cmd, "cp -f '%s' '%s'", wwi, cwwi);
if (run(cmd) != 0) return;
snprintf(cmd, sizeof cmd, "cp -f '%s' '%s'", obj, cobj);
if (run(cmd) != 0) return;
FILE *f = fopen(keyp, "wb");
if (f == NULL) return;
snprintf(twwi, sizeof twwi, "%s/P.wwi.tmp.%d", dir, pid);
snprintf(tobj, sizeof tobj, "%s/P.o.tmp.%d", dir, pid);
snprintf(tkey, sizeof tkey, "%s/P.key.tmp.%d", dir, pid);
snprintf(cmd, sizeof cmd, "cp -f '%s' '%s'", wwi, twwi);
if (run(cmd) != 0) goto cleanup;
snprintf(cmd, sizeof cmd, "cp -f '%s' '%s'", obj, tobj);
if (run(cmd) != 0) goto cleanup;
f = fopen(tkey, "wb");
if (f == NULL) goto cleanup;
fputs(manifest, f);
fclose(f);
if (rename(twwi, cwwi) != 0) goto cleanup;
if (rename(tobj, cobj) != 0) goto cleanup;
if (rename(tkey, keyp) != 0) goto cleanup;
return;
cleanup:
unlink(twwi);
unlink(tobj);
unlink(tkey);
}
/* build_one_sep — the --sep orchestration: discover_deps, reverse_topo,