42 KiB
WW Project Continuation Plan
Status date: 2026-08-05
Canonical baseline: master at 3d7973596422f1c9e2275016a5f264a2ae504ceb
This is the durable handoff for the next Codex session. Read this file before starting implementation. It records the repository audit, the test-suite audit, the branch cleanup, the first T0 implementation slice, and the recommended order of work.
0. T0 first-slice update
The focused topic branch is t0-harness-integrity, based directly on the
canonical master commit above. This slice did not touch compiler features,
CSP, bootstrap design, the cast-diagnostic worktree, recovery artifacts, or the
existing unrelated generated/untracked paths.
Completed in this slice:
test/runnow materializes the complete 337-source corpus before fan-out, records atomic started and terminal records, captures launcher failures, and checks discovered, started, completed, classified, cached, excluded, and not-run counts by identity.- Worker results distinguish pass, fail, explicit exit-77 skip, timeout, and
harness error. Every C wrapper has a 900-second outer GNU
timeoutwith a kill-after grace period. A private child-status record plus a separate GNUtimeout --verbosediagnostic distinguishes actual expiration from tests that themselves return reserved codes 124-127 or 137. make test-unitprints all eight intentional phase-2 exclusions. Unexpected skips, timeouts, missing binaries, malformed/incomplete records, launcher failures, empty corpora, and corpus shrinkage make required targets fail.- Successful wrapper output is visible. The whole-wrapper prerequisite skips in
530_w6a_parsenum,810_dyn,989_dynentry_run, and989_structlocal_framenow return the explicit skip code instead of success. - The cache key includes tracked and relevant nonignored untracked inputs,
unstaged-deletion markers, built tools, the runtime archive, cached gate
executables, and harness configuration. Ambient
WW_W6*/library/cache overrides are removed from required-suite workers. Inputs are fingerprinted again after both cache hits and misses, publication is atomic, custom workers disable caching, and the host-dependent996_dyn_wwgate is never cached. test/run_test.sh, wired asmake test-harness, exercises pass, fail, skip, TERM-resistant timeout, genuine reserved exit codes, timeout-wrapper failure, missing binary/capture/result, empty/shrunken corpora, failed launchers, zero-duration rejection, cache publication/hits, unstaged tracked deletion, uncached host-dependent gates, and visibletest-unitexclusions entirely in temporary directories.- The regex gate now runs the external black-box unit and the in-package module
root as two separate
-Tbinaries. It requires all 33 black-box tests and all 15 named white-box tests exactly once. The obsolete zero-testlib/regex/wb/regex_test.wwscaffold and its duplicate byte-ID row are gone. 989_test_filter.cuses one named four-entry bound for both expectation arrays and loops; the redundant fifth initializer is gone, and its dedicated Make rule compiles with-Werror.989_lib_byteid.cnow pins its enrolled corpus at 42 entries, so deleting a byte-ID row cannot silently shrink that cross-stage gate.
Verification on 2026-08-05:
| Command | Result | Reported scope and time |
|---|---|---|
sh test/run_test.sh |
Pass | All 16 synthetic harness scenarios |
make test-harness out/bin/test_regex_run out/bin/test_test_filter |
Pass | Harness self-test plus focused binaries |
BIN="$PWD/out/bin" timeout 300 ./out/bin/test_regex_run |
Pass | Regex black-box 33/33; white-box 15/15 |
BIN="$PWD/out/bin" timeout 60 ./out/bin/test_test_filter |
Pass | Five filter rows; ww == ww_ww output |
BIN=out/bin out/bin/test_lib_byteid |
Pass | 36 byte-identical, four pinned divergences, two pinned WW rejects across 42 units |
make test-unit |
Pass | Discovered 337; started/completed/pass 329; fail/skip/timeout/harness-error 0; eight explicit exclusions; 76 seconds |
make test |
Pass | C harness 337/337 with zero fail/skip/timeout/harness-error; 159 language files containing 1,224 tests; 153 language byte-ID files; 106/106 runww cases; 308 seconds |
The pre/post generated-artifact path scan contained the same 533 paths; neither
required run introduced a new source-tree artifact path. Final logs are in
/tmp/ww-t0-test-unit-final.log and /tmp/ww-t0-test-full-final.log, which are
not durable.
Remaining T0 and adjacent risks:
runwwstill uses shared/tmp/runww.errand/tmp/runww.bin, has no per-case timeout, and has no pinned 106-case minimum. Close these paths in the first replacement-runner vertical slice rather than extending the legacy wrapper.- The WW-native
lib/testrunner still needs T2's child completion token, timeout/capture protocol, skip reason API, and failure-count clamping; an explicit stale-record/256-failure regression therefore remains outstanding. - Several legacy wrappers have partial stage-missing branches that print a skip and continue after other assertions. Required Make prerequisites made none of those branches fire in either verified run, and their output is now visible, but converting them to structured whole-test/subtest outcomes remains work.
- Warning-fatal compilation is local to
989_test_filter.c; broader harness-Werrorrequires resolving existing unrelated missing-initializer and path truncation warnings first. - The regex module-root route is deliberately temporary and C-driver-only.
T1 must implement package-aware white-box/external test planning symmetrically
in
wwandww_ww, after whichregex_whitebox.wwcan become a canonical*_test.wwfile. - Manifest and planning appends still rely on the fresh temporary filesystem
remaining writable, and cached gates treat ordinary host
PATHutilities as part of the machine baseline rather than fingerprinting each utility. These are reasons not to expand the legacy cache/harness design.
Strategic direction recorded on 2026-08-05: the 337-wrapper/manual-Make test
architecture is a temporary migration oracle, not a design to keep extending.
After landing this branch, design the declarative replacement and migrate
runww as its first vertical slice; the new runner must provide unique temp
roots, per-case timeouts, a pinned corpus, and structured outcomes from the
start. Retire old wrappers and manual rules only after coverage-equivalence
checks pass. Do not integrate the cast-diagnostic WIP into this branch.
1. Executive decision
The project is alive and its current test commands pass. The first T0 slice has
closed the host C harness's missing-worker and vacuous-regex false greens, but the
remaining runww, package-driver, and WW-native runner contracts still prevent
the test system from being the final trustworthy development gate. The largest
immediate risk remains test architecture, not a single compiler feature.
Proceed in this order:
- Preserve the current state and make the test result trustworthy.
- Split compiler unit tests, compiler conformance tests, package unit tests, integrations, differential tests, and bootstrap tests into explicit suites.
- Finish the small compiler-stability tail, including the active cast-diagnostic work.
- Produce a usable developer release: SDK installation, documentation, CI, and import/export and error-handling cleanup.
- Start the CSP/runtime vertical slice only after renewed explicit approval.
- Treat bootstrap independence as the final milestone, not the next task.
Do not start by rewriting the compiler, changing the backend, or doing a large test-directory rename. First eliminate false-green paths and define stable suite contracts.
2. Preservation and branch cleanup
2.1 Repository state
- There is no configured Git remote. Local history has no off-machine safety net.
- The main worktree is
/home/kimchi/src/wwont0-harness-integrity, based directly onmasterat the canonical commit above. docs/spec.mdexists but is untracked. It appears to be intended as a normative language specification and must not be discarded.ref/hareandref/plan9frontare gitlinks, but there is no.gitmodulesfile.ref/inferno-osandref/qbeare untracked trees. Make their provenance reproducible before calling the repository portable.- Ignored bootstrap binaries exist under
bootstrap/amd64. - Generated output currently exists under
build/, severaltest/wcc/data/*/out/directories,scratch/, and other untracked paths. Classify these paths before deleting them; they were not part of the branch cleanup. .ai/is ignored and contains older handoffs. Use this document as the current source of truth unless an older note is explicitly cited.
2.2 Recovery artifacts created before cleanup
The following files are in the repository root and are currently untracked:
| File | Purpose | SHA-256 |
|---|---|---|
branch-backup-20260805.bundle |
Complete pre-cleanup Git history and refs | a96f48588aa406ab74c52ba8e94fdd635bfbb4e80c2602811d33c5b28a7793aa |
wip-29-castdiag-20260805.patch |
Tracked dirty changes from the active cast-diagnostic worktree | c7a5de8606ee8d9cbba32dbc5fd420c694b0c2929b673f35155db4f7066ba2d7 |
wip-29-cast-fixture-20260805.tar.gz |
Untracked cast_unknown_type fixture |
27a16b43ee83dde28c207f22050738591c1414ab917509529fb2d9d5519b5192 |
git bundle verify branch-backup-20260805.bundle reports a complete history.
The bundle contains the original 107 local branch refs plus the stash, tag, and
worktree refs. Copy these three artifacts to a different disk or remote storage;
an untracked backup on the same disk is not a permanent backup.
To recover one deleted branch without disturbing current refs:
git bundle verify branch-backup-20260805.bundle
git fetch ./branch-backup-20260805.bundle \
refs/heads/OLD_NAME:refs/heads/restore/OLD_NAME
To reconstruct the dirty cast work on a fresh worktree:
git worktree add ../ww-cast-restore -b restore/castdiag master
git -C ../ww-cast-restore apply \
/home/kimchi/src/ww/wip-29-castdiag-20260805.patch
tar -xzf /home/kimchi/src/ww/wip-29-cast-fixture-20260805.tar.gz \
-C ../ww-cast-restore
2.3 Branch cleanup result
The repository started with 107 local branches. Cleanup deleted 105:
- 90 branch tips were direct ancestors of
master. - 14 branches looked unmerged by graph ancestry but were audited and found to be
patch-equivalent to, or superseded by, later work on
master. wip-bugawas already merged; its clean worktree was removed and the branch was deleted.
Before T0 implementation, only these branches remained. The focused T0 branch was then added:
| Branch | State | Decision |
|---|---|---|
master |
Clean tracked baseline at 3d79735 |
Keep |
wip-29-castdiag |
Checked out at /home/kimchi/ww-29; contains uncommitted source and fixture changes |
Keep until reviewed and integrated |
t0-harness-integrity |
Checked out at /home/kimchi/src/ww; first T0 gate-integrity slice |
Review and merge separately |
The active WIP adds a clean diagnostic for a cast to an unknown type in
selfhost/cmd/wcc/check.ww, plus a fixture under
test/wcc/data/cast_unknown_type/. Do not delete this worktree merely because
its branch tip is already reachable from master; the valuable part is the
dirty working tree. The patch and fixture archive above are secondary recovery
copies.
A detached worktree remains at /home/kimchi/src/ww-f0 on e8977a4. It is not
a branch and was therefore not removed. Audit it separately before deciding
whether to remove it.
The 14 non-ancestor branches were matched to later master work as follows:
e3-c1-flip,e3-oracle: separate-compilation flip later landed around33edc38.f1-stack-v1,f1-stack-v2,f1-stack-v3: superseded by9bd0d8b,b9dd297,c138605,ac63951, and later type-chase work.impl-26: superseded by9fcb3be.impl-ascii-bytes-cut,impl-ascii2-complete: superseded by07fed80and2b893b9.impl-strconv: finalized through6a5cdbd,a117852, anddb5c5b6.impl-strings-cutrcut: patch-equivalent to38a906c.t28,t28-parked: superseded byc40b2dfand330792b.wip/io5-handle-seeker: superseded by06c00e0and532b0a8.worktree-impl-e8: superseded byb842f93,10cb835, and7d39f6d.
The bundle remains the authoritative record if any of those classifications must be revisited.
3. Verified build and test baseline
The following commands were run on master at the baseline commit before this
document and branch-only cleanup:
| Command | Result | Observed scope and time |
|---|---|---|
make test-unit |
Pass, 329/329 C harness entries | About 95 seconds |
make test |
Pass | C harness 337 entries; test/lang 159 files; language byte-ID pass over 153 files; runww 106 cases; about 313 seconds total |
The full audit log was written to /tmp/ww-tests-health-full.log; /tmp is not
durable, so rerun the commands when a fresh log is needed.
These historical green results prove that the canonical baseline is buildable. They did not prove that every named test executed useful assertions; the first T0 slice above closes the missing-worker and empty regex-table paths without changing this preserved baseline record.
No source code changed during the branch cleanup and English-document rewrite, so the full suite was not rerun afterward.
4. Current test architecture
The project already contains substantial test material. The problem is that several different kinds of tests are mixed behind numeric C wrappers and broad Make targets.
4.1 C harness under test/wcc
- 337 C test sources are registered through roughly 337 manual Make rules.
- The sources contain about 89,519 lines.
- Static inspection found subprocess orchestration in approximately 333 tests
(
fork,exec,system, orpopen). - Only a handful, approximately the
000,100,200,300, and738families, resemble direct compiler-component unit tests. - Most files are integration wrappers around compiler commands, generated programs, output matching, byte-ID checks, or bootstrap behavior.
Therefore make test-unit is currently misnamed. It omits a few heavy phase-2
tests, but still runs most integrations and includes 989_lib_byteid, which
alone took roughly 42 seconds in the audit. It does not run the 1,224-test
test/lang corpus or the runww fixtures, but C wrappers do indirectly launch
the library @test files. The first T0 slice makes all 458 library tests
effective, including the 15 regex white-box tests.
4.2 Language tests under test/lang
- 159 source files contain exactly 1,224 source-level
@testfunctions. - Six files are
_runonly_test.wwvariants. - The current Make targets compile/run this corpus and also perform byte-ID comparison for most of it.
This is valuable language and compiler coverage, but file layout and target naming do not clearly separate package unit behavior from compiler conformance behavior.
4.3 Library tests
- 35 test-named library source files contain 443
@testfunctions. Eleven use_test.ww; 24 use other names ending intest.ww. lib/regex/regex_whitebox.ww, a production-named source file, contains 15 additional@testfunctions.- There are therefore 36 files that actually contain tests and 458 source-level library tests in total.
Together, test/lang and lib contain 1,682 real source @test functions,
excluding generated combined files.
4.4 Directive fixtures through runww
There are 106 directive-driven cases:
- 100
//ww:errorcases. - 5
//ww:run-exitcases. - 1
//ww:runcase. - 0
//ww:compilecases.
The rejection coverage is useful, but this is not yet a balanced compiler conformance suite. Rejections do not identify whether failure was expected in lexing, parsing, checking, code generation, assembly, or linking. Accept/run cases use only the C stage, while rejects exercise two stages.
4.5 Current @test implementation
The compiler's -T path synthesizes a __wwtests value table and test main.
@test functions are expected to have the shape fn() void. Normal builds
type-check tests and then strip them. The runtime in lib/test forks once per
test, so an abort, signal, or nonzero exit does not stop subsequent tests. The
runner also supports glob filtering and treats zero matches as success.
The per-test fork model and compiler-generated value table are good foundations and should be preserved. The package driver and child/parent result protocol need to be completed.
4.6 Target coupling and differential caveats
- Despite its name,
make test-runis not an isolatedrunwwtarget. Its order-only prerequisites build the C harness and runtest-langplustest-lang-byteidfirst. The redesigned target graph must remove this surprising coupling. test-lang-byteidreads onlytest/lang. The six positiverunwwfixtures are not checked for stage-by-stage assembly identity.989_lib_byteidhas 42 rows: 36 require identity, four are declared divergences, and two are expected WW-stage rejects. A green result does not mean complete standard-library stage parity.runwwitself is compiled by the C-stage compiler and then invokes that same stage. It is useful coverage but not an independent sole oracle. Retain an independent host-side runner or cross-validate the harness during migration.
5. Correctness risks in the current test result
Treat the following as gate-integrity bugs, not cosmetic cleanup.
5.1 Vacuous regex test: closed in the first T0 slice
The former lib/regex/wb/regex_test.ww root contained no tests while importing
regex as a dependency, so -T stripped all 15 intended white-box tests. The T0
gate now resolves regex itself as the root module, keeps it separate from the
external regex_test binary, and asserts every expected name plus the exact
15/15 summary. The empty scaffold was removed. T1 still needs to replace this
temporary module-root route with ordinary package-aware test assembly.
5.2 Harness false-green paths
Closed in the first T0 slice:
test/runno longer masks eitherxargslauncher. It collects from a complete manifest and fails on missing starts, missing/malformed terminal records, launcher errors, or any count/identity mismatch.- Pass, fail, explicit skip, timeout, and harness error are distinct outcomes; required targets fail on every unexpected non-pass outcome.
- Successful wrapper output and every expected
test-unitexclusion are visible. - The cache hashes nonignored untracked inputs plus effective built inputs, represents tracked deletions, revalidates hits and misses at run end, and publishes atomically only after complete green accounting. The host-dependent dynamic gate is deliberately uncached.
- The dead nested
test/langscan is gone. 989_test_filter.chas one named expectation count, no excess initializer, and a rule-local-Werrorgate.
Remaining:
- Existing generated artifacts outside the new harness result/temp directories have not been broadly relocated or classified.
- Partial prerequisite skips inside legacy wrappers are not yet structured subtest outcomes. None fired in the recorded runs, and whole-wrapper skips now use exit 77.
- Broad warning-fatal harness compilation remains blocked by pre-existing
warnings outside
989_test_filter.c.
5.3 runww limitations
- It uses shared fixed paths such as
/tmp/runww.errand a fixed binary name, which prevents safe parallel execution and risks stale-file interference. - It supplies a nil environment to subprocesses.
- Negative cases check a diagnostic substring but do not identify the failure phase.
- There is no per-case timeout.
- Missing arguments fail, but there is no expected/minimum corpus count. A suite accidentally reduced from 106 fixtures to one can still pass without a corpus-shrinkage alarm.
5.4 Package-driver limitations
ww test DIRdiscovers only immediate_test.wwfiles and does not recurse.- Files are not sorted deterministically.
- Each discovered file is built separately instead of assembling one package test unit. Same-package helper sharing therefore does not work as expected.
- Production enumeration excludes all
*test.ww, while test discovery includes only_test.ww; the naming rules disagree. - Only the first glob argument reaches the runner even though the runner can support more than one.
- A single file with no tests can silently return success, while a directory with no recognized tests can error. Define one consistent contract.
- Imported dependency tests are stripped. This is normally correct, but root package files containing tests must be assembled into the root test build.
-cwithout-ocan leave a temporary test executable behind.
5.5 Runner-protocol limitations
- A test can call
os.exit(0)before completing and the parent reports it as a pass. The parent needs an explicit completion token from the child. - The runner returns the raw number of failures. At 256 failures, the Unix exit status can wrap to zero.
- There is no timeout, structured capture, duration reporting, skip reason, expected-abort protocol, cleanup hook, or per-test temporary directory.
- Invalid
@testdeclarations are checked too narrowly and mainly under-T. Exported tests, prototypes, arguments, variadics, non-void returns, duplicate annotations, and other invalid forms need deterministic diagnostics in normal and test builds.
6. Target model: Go/Hare-style separation
The desired result is not a copy of either project, but it should adopt their separation of concerns.
6.1 Lessons from Go
The local Go 1.25.4 source tree demonstrates two distinct systems:
- Compiler/runtime conformance fixtures live outside ordinary package unit tests and use explicit execution/error directives.
- Ordinary
*_test.gofiles are assembled per package into a generated test main, with support for same-package tests, external-package tests, filtering, caching, timeouts, parallelism, and structured output. - Compiler internals also have direct component unit tests rather than wrapping every assertion in a compiler subprocess.
WW needs the same conceptual split: compiler correctness fixtures are not the same thing as user-facing package unit tests.
6.2 Lessons from Hare and harec
The vendored ref/hare test command demonstrates useful package-test behavior:
- Recursive module discovery when no explicit module is given.
- Deterministic sorting, multiple globs, and a test-only build tag.
- First-class skip, require, expected-abort, current-test, capture, timing, and backtrace behavior.
The upstream Hare compiler suite uses feature-grouped positive and negative
fixtures and distinguishes outcomes such as success, user error, lex failure,
parse failure, check failure, and abnormal failure. The vendored ref/harec
copy is incomplete and lacks the full compiler test corpus; pin an upstream
version or import a documented reference snapshot before using it as a durable
oracle.
Keep WW's fork isolation and generated source value table. Do not copy Hare's same-process recovery mechanism or unrelated linker-specific structure.
7. Proposed suite contracts
Use semantic directories and explicit metadata. Numeric prefixes may remain as stable case IDs during migration, but they must no longer determine what a test means or when it runs.
test/
unit/ Direct host-side compiler/library API tests; no subprocesses
compiler/ Source conformance fixtures grouped by language feature
package/ Fixtures for `ww test` discovery and package assembly
integration/ CLI, import graph, archive, ELF, FFI, and toolchain behavior
differential/ C-stage versus WW-stage and byte-ID comparisons
bootstrap/ Fixed-point, self-host, and no-C-compiler scenarios
harness/ Shared runners, manifests, and assertion helpers
Actual package unit tests should stay beside their code as canonical
*_test.ww files. test/package is for testing the package-test driver itself,
including multi-file layouts, helper sharing, external test packages, empty
packages, recursion, and failures.
Define these Make targets:
| Target | Contract |
|---|---|
test-unit |
Direct in-process compiler and library API tests only; zero child compiler processes |
test-compiler |
Positive, negative, run, compile, assembly, and multi-file language conformance fixtures |
test-package |
Real ww test package tests and package-driver fixtures |
test-integration |
CLI, separate compilation, archives, ELF, FFI, and graph behavior |
test-differential |
C-stage/WW-stage output and byte-identity checks |
test-bootstrap |
Fixed point, self-host, and nocc; explicitly slow |
test-lint |
Harness integrity, fixture metadata, duplicate IDs, empty corpora, and warning checks |
test-fast |
Unit tests plus a documented compiler/package smoke set |
test |
Complete required pre-merge set, with no hidden skips |
test-all |
test plus bootstrap and any explicitly labeled slow suites |
Every target must print discovered, passed, failed, skipped, timed out, and not run counts. A target may intentionally exclude another named suite, but it may not silently skip entries within its declared corpus.
Compiler fixtures need explicit metadata for at least:
- Mode: check, compile, run, run-with-exit, reject, assembly match, or multi-file.
- Expected stage: lex, parse, check, codegen, assemble, link, run, signal, or success.
- Expected diagnostic, stdout, stderr, exit status, or signal.
- Timeout and optional platform/tool prerequisites.
- Stable feature group and case ID.
Use one declarative manifest or one clearly documented directive grammar. Do not continue adding one manual Make rule and one custom C subprocess wrapper for every source fixture.
8. Required ww test semantics
The language specification requires both same-package white-box tests and an
external <package>_test black-box package. A package-under-test invocation must
produce one top-level result covering both, but the implementation cannot simply
link two current -T units: every -T unit owns global main and __wwtests
symbols, dependencies are compiled without -T, and .wwi does not carry test
descriptors.
Use this implementation sequence:
- Treat
*_test.wwonly as the candidate-file discovery rule. Classify a file as white-box or black-box from its declared package (package fooversuspackage foo_test), not from its suffix alone. - Compile production files plus same-package test files into one white-box root unit.
- Compile the external package against only the public API. Initially, build a second test binary and aggregate both binary results in the driver. A later design may instead emit linkable test descriptors and generate one aggregator main, but it must first solve the current symbol and interface limitations.
- Exclude test bodies from imported dependencies unless that dependency is the package under test.
- Give each test a canonical package-qualified ID such as
<import-path>.<function>. Bare leaf names are ambiguous across the white-box and black-box units and are not sufficient for output or filtering. The generated descriptor must also refer to the callable through its owning package rather than assuming a bare function reference. - Produce one top-level package result even if it is internally implemented by two test binaries.
- Support multiple glob filters and a recursive package pattern equivalent in
purpose to
./.... - Use one canonical file naming rule:
*_test.ww.
Normal production enumeration must exclude the entire *_test.ww file,
including test-only helpers and imports. An @test declaration placed in a
production file must still be fully checked in a normal build before its function
is stripped. These are separate driver and checker responsibilities.
Make ordering an explicit compatibility decision. The recommended first contract is to sort package and source discovery paths, then preserve declaration order within each source. Do not silently change the existing documented source/collection behavior to a global name sort.
Define all empty cases separately:
- An explicitly requested package with no tests reports
[no tests]and succeeds. - An explicitly requested test file containing no
@testdeclaration is an error because it is likely a fixture mistake. - A filter matching zero tests reports
[no matches]and succeeds, preserving the current useful filtering behavior. - Recursive discovery reports packages with no tests visibly, but a pattern that resolves to zero packages is an error.
- A required suite with an unexpectedly empty or shrunken corpus is a harness error regardless of the user-facing command rules.
Validate every @test declaration during normal checking, even when test bodies
are later stripped. At minimum reject parameters, variadics, non-void returns,
prototypes without bodies, invalid export combinations, @test(...), unknown
attributes, placement on non-functions, and duplicate attributes with stable
diagnostic text. Put both the C-stage and self-hosted parser/checker through the
same acceptance fixtures.
Define compile-only behavior rather than deleting its output: ww test -c
should preserve a deterministic <package>.test binary by default, and -o
should select its path. Specify whether -o without -c is invalid; the
recommended contract is to require -c so output preservation is unambiguous.
Keep one process per test, but add a control pipe from child to parent. A normal
return must write a completion record. os.exit(0) without that record is a
failure. The same protocol can carry skip, expected-abort, failure reason, and
captured-output metadata. Clamp the process exit code to 0 for no failures and
1 for any failure instead of returning the raw failure count.
9. Test-system implementation plan
T0: Make the existing green result trustworthy
Do this before moving directories. Status after the first focused slice:
- Make
test/runcompare discovered, started, and completed test counts. - Remove
|| truemasking around worker execution; collect worker failure as a harness failure. - Represent pass, fail, skip, timeout, and harness error separately.
- Print skip reasons/output and fail required targets on unexpected skips.
- [~] Harness records and new regression artifacts are private temporary files, and relevant source/configuration inputs are in the cache key. Broad legacy source-tree artifact relocation remains.
- Give every host C-harness subprocess an outer timeout and kill its process group on expiry. WW-native enforcement remains T2 work.
- Replace shared
runwwpaths with a unique per-case temporary directory. - Fail the C-harness target on zero or shrunken corpora and incomplete worker
results. Add the equivalent pinned minimum to
runwwnext. - Fix the
mustnot[4]initializer and related loops through one named count. - Prove all 15 regex white-box tests execute in a separate root-module
binary without combining conflicting
-Tunits. - [~] Sixteen harness scenarios cover fail, skip, TERM-resistant timeout, reserved child exits, timeout-wrapper failure, lost/incomplete workers, failed launchers, missing binary/result/captures, empty/shrunken corpora, invalid bounds, cache safety, and expected exclusion. A fresh per-run result directory prevents stale reuse; the WW-native 256-failure case remains.
The host C-harness portion of the completion criterion is met: a killed or
skipped worker cannot leave a required target green, and the regex test reports
15 discovered and 15 passed white-box tests. Finish the unchecked runww items
before declaring all of T0 complete.
T1: Implement package-aware ww test
- Introduce a package test-plan data structure in the driver.
- Unify production-file exclusion and candidate discovery on
*_test.ww, then classify candidates by their package declarations. - Assemble production files and same-package tests into one white-box root.
- Build the required external
<package>_testunit as a second test binary first, and aggregate its result with the white-box binary at the driver level. Do not attempt to link two current-Tmains together. - Keep imported dependency tests stripped.
- Add package-qualified test IDs, deterministic recursive discovery, and multiple filter arguments.
- Implement the documented
-coutput contract and all four no-test/no-match cases above. - Add C-stage and self-hosted compiler diagnostics for all invalid
@testand attribute shapes. - Convert the regex white-box case into an ordinary, non-vacuous package test.
Completion criterion: two test files in one package can share a private helper,
dependency tests do not leak in, filters work deterministically, and empty or
invalid packages produce the documented result. The C-stage ww and self-hosted
ww_ww must discover the same fixtures and produce matching output classes and
exit statuses.
T2: Finish runner semantics
- Add the required OS/runtime wrappers for pipes, process-group creation, and process-group signals before depending on them in the WW-native runner.
- Add the child completion/control protocol.
- Add skip and require helpers, expected abort, current test identity, cleanup, per-test temporary directories, timeouts, captured output, and duration.
- Drain stdout, stderr, and control pipes while the child is running; waiting before draining can deadlock when a pipe buffer fills. Add polling/nonblocking support as needed.
- Ensure signals and premature
os.exit(0)are classified correctly. - Offer a stable human-readable output and a machine-readable event stream.
Completion criterion: the runner's own conformance suite covers every outcome and never wraps failures into exit status zero.
T3: Build the compiler conformance engine
- Create the explicit fixture schema and stage-aware outcome model.
- Migrate
runwwfirst because it is already directive-driven. - Keep an independent host-side fixture runner as the gate, or cross-check the WW-native harness so the compiler under test is not its own sole oracle.
- Add balanced positive/check/compile/run cases beside rejection cases.
- Add multi-file, diagnostic location, exit, signal, stdout/stderr, and assembly matching support.
- Pin the external Hare/harec reference version used for design comparison.
This is a replacement program, not a permanent adapter around test/run.
Freeze the current C wrapper layer as a comparison oracle, migrate by declared
suite and outcome, and remove each old path when the new engine proves the same
coverage. Do not add new numeric wrappers or hand-written Make registrations for
cases that fit the new fixture schema.
Completion criterion: adding a language fixture normally requires source plus metadata, not a new C wrapper or Make rule.
T4: Migrate and shrink test/wcc
Classify every existing C wrapper as unit, compiler fixture, package, integration, differential, bootstrap, or obsolete duplicate. Move direct API tests first. Mechanically migrate simple subprocess wrappers into fixtures. Keep C only where the assertion genuinely needs host APIs, ABI inspection, or custom orchestration.
Completion criterion: test-unit launches no compiler subprocess, suite
membership is discoverable without numeric-prefix knowledge, and the giant
manual Make registry is replaced by generated discovery plus a small manifest.
T5: Add CI cadence
- Pre-commit/local:
test-fast. - Required pre-merge:
testplus harness lint. - Scheduled or release:
test-all, bootstrap, sanitizers, and expensive differential/byte-ID jobs. - Archive structured results and fail on unexpected skips or corpus shrinkage.
Completion criterion: a fresh checkout can reproduce the declared test matrix without ignored local binaries or undocumented reference trees.
10. Broader project roadmap
P0: Preserve and normalize the repository
- Copy the recovery artifacts off-machine.
- Add a remote or another durable backup.
- Decide which of
docs/spec.md, reference trees, probes, and scratch artifacts are source, vendored dependencies, or disposable output. - Track reference provenance through submodules, pinned archives, or documented fetch scripts.
- Commit this plan and the chosen specification files. Do not commit generated
output or the bundle merely to make
git statusquiet.
P1: Establish the test foundation
Execute T0 through at least T3 before broad compiler feature work. T4 can proceed incrementally after the contracts stabilize.
P2: Finish the compiler-stability tail
- Review and integrate
wip-29-castdiag, including a regression fixture. - Finish or explicitly close the large aggregate return (
sret, especially values over 24 bytes) edge cases. - Re-audit stale issue-number tests and convert unresolved cases into explicit fixtures.
- Finish type-information (
tinfo) consistency work and remove duplicated or stage-specific behavior where possible. - Preserve separate-compilation and byte-ID guarantees while changing these paths.
Each item needs a focused branch, a regression test in the correct new suite, and a clean full required test run before merge.
P3: Produce a developer-usable v0.1 release
- Finish import/export membership rules and diagnostics.
- Normalize
errnoand error propagation across the standard library, especially the existing error/I/O tail. - Turn
make installinto an actual SDK installation. It currently copies only thewwexecutable andlibwcc.a; define and install the compiler, standard library sources/artifacts, headers or metadata, and version information. - Add a README with bootstrap prerequisites, build, test, install, and first program instructions.
- Track the language specification and label implemented, reserved, and planned behavior.
- Add CI from a fresh checkout and publish one reproducible release artifact.
- Document the current allocation lifetime honestly: both the compiler's
free(x)lowering andos.free(p, n)ultimately reach the current no-oprt_free, so neither reclaims memory today. Decide and test a real reclamation policy separately before promisingfreesemantics.
Release criterion: a developer can clone, build, run the required tests, install to a staging prefix, and compile a small program using the installed SDK without knowledge of this worktree.
P4: CSP/runtime vertical slice, only after explicit approval
The existing design direction is a cooperative M:1 runtime with swapctx, a
run queue, and an unbuffered channel implemented library-first, followed by
syntax. epoll, alt/select, broader scheduling, and optimization come later.
However, the recorded project status says CSP work was explicitly put on hold on 2026-06-18. Do not infer that a general request to resume the project cancels that hold. Ask for renewed explicit approval before implementing CSP/runtime features. The specification may reserve the design in the meantime.
When approved, implement one narrow end-to-end slice: spawn two cooperative tasks, rendezvous over one unbuffered channel, shut down deterministically, and cover it in package, integration, and differential tests.
P5: Bootstrap independence last
make bootstrap already checks a fixed point, make nocc exists, and ignored
stage-0 binaries are present. The process is not yet a clean independent bootstrap
because the C driver and local artifacts still orchestrate important steps.
Freeze bootstrap only after the language, package graph, tests, SDK layout, and runtime surface are stable. Then pin stage 0, document trust inputs, build stage 1 and stage 2 in a clean environment, compare the intended artifacts, and prove the process without a C compiler where promised.
11. Next-session checklist
The next Codex session should do the following, in order:
- Read
CLAUDE.md, this file, anddocs/spec.md; do not rely on.ai/notes without checking their date and status. - Inspect
t0-harness-integrityand its final commit/diff againstmasterat3d79735; do not repeat the repository-wide audit unless that state changed. - Verify the three recovery artifact hashes and arrange an off-machine copy.
- Inspect
/home/kimchi/ww-29and compare its working tree withwip-29-castdiag-20260805.patchand the fixture archive. Do not delete it. - Review and merge the first T0 branch separately, preserving all unrelated untracked/generated paths. Re-run required targets if the diff changes.
- Specify the replacement test manifest, structured result protocol, suite boundaries, and coverage-equivalence report; keep fork-per-test isolation.
- Migrate
runwwas the first vertical slice, implementing unique temp roots, per-case timeout, and the pinned 106-case corpus in the new engine rather than further entrenching the old wrapper architecture. - Begin package-aware
ww testassembly after that first replacement slice; keep white-box and external-package binaries separate while generatedmain/__wwtestssymbols still conflict. - Integrate the cast-diagnostic WIP as its own reviewed branch once the new test
placement is available, then remove
wip-29-castdiagand its worktree. - Do not start CSP or bootstrap redesign unless the prerequisites and approval above have been satisfied.
12. Working rules for future branches
- One branch should represent one reviewable outcome, not an indefinite work queue.
- Record the intended suite and completion criterion before implementing.
- Rebase or merge promptly, then delete the branch after verifying reachability and a durable backup.
- Never classify a branch only with
git branch --merged; inspect dirty worktrees, patch equivalence, fixture files, and later superseding commits. - Keep generated output out of topic branches.
- A green target is acceptable only when it reports a complete corpus and explicit skips.
The immediate success condition is not more test count. It is the ability to state exactly which compiler unit tests, compiler fixtures, package tests, and integrations ran—and to know that none of them passed vacuously.