live via QMP, ring sizing decided
include/starkernel/xhci.h: Capability/Operational/Runtime register
layouts, Port Register Set, Interrupter Register Set, Doorbell Array,
16-byte TRB struct -- all from the xHCI 1.2 spec, no existing
reference in this tree to build from (unlike virtio-blk). volatile
fields, no packed attribute, matching virtio_blk.c's documented
riscv64/QEMU-MMIO precedent. Compile-checked clean, sizeof(xhci_trb_t)
verified == 16.
QEMU qemu-xhci's PCI vendor:device ID (0x1B36:0x000D) confirmed live
via QMP query-pci against a real running instance -- not assumed from
memory, matches the Milestone 1 QMP infrastructure just built.
Ring sizing decided: fixed 256-TRB (one page) Command Ring and Event
Ring, single interrupter -- documented rationale in the header.
Bonus finding: src/starkernel/pci/pci.c already has more reusable
infrastructure than Milestone 2b assumed (pci_find_first is ID-based
lookup already existing; pci_bar/pci_map_bar/pci_enable are already
generic) -- 2b is smaller than originally scoped, noted in the punch
list.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Leftover from the milestone renumbering -- cross-references elsewhere
already said "Milestone 2e" etc., but the bare sub-item labels inside
Milestone 2's own section still said "3a."-"3h.". Now consistent.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
launch targets, three-arch verified
Makefile.starkernel: added -qmp unix:$QMP_SOCK,server=on,wait=off to
the amd64/aarch64/riscv64 qemu targets, matching the existing serial
chardev socket pattern exactly (same discoverability, same cleanup on
exit). Verified on all three architectures: QMP greeting arrives on
connect, qmp_capabilities handshake succeeds, device_add/device_del
round-trip correctly.
Real finding surfaced during device_add testing (recorded in
FABRIC-2.md's punch list for Milestone 2): the q35 machine's pcie.0
root bus doesn't support runtime PCI hotplug without a bridge --
Milestone 2's qemu-xhci USB controller needs to be present in the
static launch command, with USB devices hot-attached to its bus at
runtime, not the controller itself hot-added.
Also noted: g_doe_log_enabled's default-on per-tick heartbeat CSV
export was briefly mistaken for a hang during aarch64 verification --
it isn't one, just a large volume of routine diagnostic output before
reaching ok>. Not changing the source default; adopting HB-OFF
immediately after boot as the working pattern for the rest of this
punch list's dev work.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
execution order, physically reordered in the file
Previously kept stable IDs with a separate "read this note for the
real order" translation layer. Renumbered so the milestone numbers
themselves read top-to-bottom in execution order and physically
reordered the ### Milestone blocks in Section X to match -- no
translation needed anymore. New order: 1=QEMU monitor/QMP socket,
2=USB hardware stack, 3=block subsystem extensions, 4=drive/credential
security, 5=console/VM key-match binding, 6=PKI signing chain,
7=contributor capsules/trust tiers, 8=bare-metal USB boot (still
second-to-last, not first -- QEMU-first per Captain Bob's explicit
reinforcement), 9=networking (unchanged, still last). All cross-
references between milestones (including Milestone 2's internal
sub-item labels 2a-2h) updated to match throughout Sections U, W, and
X.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
continuous chain, not two parallel paths; resolves Milestone 7's
bootstrapping question precisely, not just for dev/test
Captain Bob's precise correction: "my own real root CA -> snakeoil
embedded cert -> blob & capsule + MANIFEST.md" is one chain. The
snakeoil intermediate is CA-signed, not self-signed/untrusted --
"snakeoil" names its informal/private-project status, not that it
lacks a real trust root. This means Milestone 7's CA-bootstrapping
question (how does the CA public key get into the kernel without
being just another unverifiable capsule) is resolved outright, not
just worked around for dev/test builds as the previous draft of
Section U's fourth addendum implied: trust is established once, at
build time, by whoever holds the real root CA and produces the build.
No kernel-boot-time verification against a hardcoded CA public key is
needed at all. Corrected both Section U item 20 and Milestone 7's
punch list in Section X to match.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
cert resolves Milestone 7's bootstrapping question for dev/test builds
Three distinct user-role tiers clarified (builder+dev, SDK-only dev,
no-SDK user) -- refines the trust-tier question beyond a simple core/
contrib boundary. More importantly: a snakeoil cert embedded directly
into each build (not loaded as a capsule, not verified against an
external CA at boot) is the actual answer to the CA-bootstrapping
problem Section X's Milestone 7 punch-list surfaced, at least for dev/
test builds -- trust is established at build time, sidestepping the
runtime chicken-and-egg entirely for that case. Two paths into a build
confirmed: snakeoil-signed, or code review + inclusion in the source
repo -- the latter likely makes at least one of Milestone 8's four
spitballed trust-tier directions (signature-authority tiers)
redundant, worth revisiting before picking one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
in QEMU first, defer real-hardware USB boot to last; add UEFI-only
boot-path constraint
Captain Bob's explicit correction: real-hardware boot is "difficult
and lots of blind guesswork" (no serial log, unknown firmware quirks)
and isn't worth attempting until there's a working Artemis subsystem
to actually demonstrate, not just an empty kernel proving UEFI boot
works. QEMU's own USB hotplug emulation is sufficient to build and
validate the entire home-blocks subsystem without touching real
hardware at all.
Milestone IDs in Section X kept stable (not renumbered, to avoid
breaking cross-references between milestones) with an explicit
execution-order note instead: 2 -> 3 -> 4 -> 5 -> 6 -> 7 -> 8 -> 1,
networking (9) still deferred past all of them.
New scope constraint captured for whenever Milestone 1 resumes: UEFI-
only boot path, no legacy BIOS/MBR, no GRUB2 -- starkernel_loader.efi
is meant to be the entire boot path, not one stage in a longer chain.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
breakdown across all 9 milestones, no code written
Nine milestones, in Section W's sequencing order, each broken down to
single-function granularity per Captain Bob's explicit request:
1. Bare-metal USB boot (demo mode) -- the actual near-term milestone,
~9 concrete steps from ISO build through real-hardware validation.
2. QEMU monitor/QMP socket -- small dev-workflow unblock.
3. USB hardware stack -- the hard prerequisite, broken into 8 sub-areas
(spec groundwork, PCI discovery, controller bring-up, interrupt/
event handling, hotplug detection, device enumeration, Bulk-Only
Transport read/write, block-subsystem integration) since nothing
in this tree has ever touched USB before.
4. Block subsystem extensions (identity-derived ranges, drive map,
migration state machine, sk_repl_idle() body, unclean-removal
handling).
5. Drive/credential security (foreign-drive signature, zuse one-way
burn extending the existing acl_pinned mechanism).
6. Console/VM key-match binding.
7. Kernel/capsule PKI signing chain, including a real open
bootstrapping question (how the CA public key itself gets into the
kernel without being just another capsule) not previously
surfaced in Sections U/V/W.
8. Contributor capsules/trust tiers.
9. Networking -- deliberately left unexpanded, deferred per Captain
Bob's own sequencing, not premature-detailed.
Every item cross-references back to the specific Section U requirement
and Section V verified-status finding it comes from. Still direction
only -- no code, no capsule work started.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
one place, pulling together Sections U and V
Single-read consolidation of the whole Artemis thumbdrive/home-blocks/
PKI vision: one-paragraph statement of intent, a component table
cross-referencing each piece to its detail (Section U) and verified
status (Section V), and explicit sequencing with the actual near-term
milestone first -- bare-metal boot from a physical USB stick in demo/
"try it" mode, confirmed NOT gated on the USB hardware driver gap
since booting from USB is a UEFI firmware responsibility, not a kernel
one. The existing starkernel.iso/raw-disk-image build artifacts (built
on every QEMU launch already) are, mechanically, what gets dd'd onto a
physical drive for this. Two 64GB SanDisk drives confirmed on hand and
available now. Everything else in the concept board explicitly waits
on the USB hardware driver as the singular hard prerequisite.
Still direction only -- nothing implemented, no hardware testing done.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
blocks/thumbdrive/PKI subsystem against the actual codebase
Systematic pass over every requirement gathered in Section U, each
verified directly against the tree rather than recalled. Organized by
area (physical block layout, USB hardware, drive/credential security,
console/VM binding, capsule signing/PKI, contributor trust tiers,
networking, dev workflow).
Found more than expected already exists and is reusable as-is: the
block address space and device-chain abstraction, sk_repl_idle()'s
empty trigger hook, capsule_birth_baby()'s on-demand VM spin-up,
acl_pinned's one-way-ratchet mechanism (a direct precedent for the
zuse one-way-burn requirement), arbitrary binary payload capsule
embedding (proven by the font capsule), the manifest's already-
documented Ed25519 anchor point, and the live-boot ISO pipeline.
Confirmed real, clearly-scoped gaps with nothing partially started:
identity-to-block-range derivation, the block migration state machine,
drive-map format, console/VM key binding, all signature/cert
verification code, magic-number content-type/foreign-drive detection,
the contributor trust-tier flag, QEMU monitor socket exposure, and the
install path.
Identifies the USB stack itself as the one hard, load-bearing
prerequisite gating almost every other gap from being testable at all,
even in QEMU -- the honest first-cut recommendation if a concrete next
milestone gets picked from this analysis.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
sequencing, and spitballed trust-tier ideas -- still brainstorming,
nothing implemented
Captures the QEMU-dev-environment assumption, confirms capsules/contrib/
would fit the existing subdirectory convention (artemis/, common/,
fonts/, hermes/ already exist), and confirms via mkcapsule.c that no
provenance/trust-tier distinction exists today -- every non-Mama-init
capsule gets identical FLAG_PRODUCTION|FLAG_EXPERIMENT unconditionally.
Explicit sequencing: ACL/PKI work closes first, then contrib-directory/
trust-tier work, then networking (downloadable capsules named but not
scoped). Closes with four explicitly-unvetted spitballed directions on
the trust-tier question, requested as free brainstorm: a new
FLAG_CONTRIB bit, signature-authority tiers hanging off the cert chain,
block-namespace sandboxing for contrib capsules, and QEMU-vs-real-
hardware conditional signature enforcement.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
extending items 7-8 -- still brainstorming, nothing implemented
Five more points captured: (10) intermediate cert embedded as a capsule
blob (CA stays external/unrevocable), reusing the capsule system's
already-proven arbitrary-binary-payload capability (the font capsule is
existing precedent); (11) two-stage validation chain, both stages net
new code; (12) confirmed via tools/mkcapsule.c's own header comment
that Ed25519 signing hanging off the xxHash64 manifest column was
already the documented Phase 8 plan, independent of this conversation
-- strong validation of the whole direction; (13) signing granularity
is per-capsule, matching the existing hash column's 1:1 file
granularity exactly; (14) content-type detection via magic numbers
rather than a new MIME-type field, and confirmed to be the same
mechanism as item 7's foreign-drive detection -- one shared
byte-sniffing primitive plausibly serves both.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
burn, console/VM key-match attachment -- still brainstorming, nothing
implemented
Three more requirements captured while fresh, same design session:
(7) home-blocks write path must check for a home-blocks signature
before ever writing to an inserted drive, warn and refuse on foreign/
unrecognized/blank media instead of silently claiming it; (8) zuse
credential minting is one-way, asymmetric with an operator drive's
presumed re-provisioning path; (9) console/VM split -- console is
generic and shared, drive insertion spins up a per-identity VM (a
Tripod-birth-mechanism consumer), and console-to-VM attachment is a
key/lock match, structurally similar to ACL-PIN's existing key model
but not yet confirmed to reuse it directly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
direction notes -- brainstorming only, nothing implemented
Captures the direction from a design conversation immediately following
the ACL-TTL campaign close (Section T): Phase 8 PKI/thumbdrive context,
confirmation that zero USB code exists anywhere in the kernel today,
verification that the block-address-space layout the conversation
converged on independently already matches block_subsystem.h's own
documented (unimplemented) chained-device design almost exactly, and
six requirements gathered in order (no quota for now, re-insertion
consistency, identity-derived not attach-order-derived block ranges,
drive-carries-its-own-map, bidirectional transparent block migration
as a state machine, and sk_repl_idle() as the likely trigger hook --
already an empty coarse-cadence placeholder found during Section R).
Explicitly not a spec or plan of record -- written up so the next
session starts from an accurate baseline instead of re-deriving the
shape from scratch. No code, no design doc, no capsule work started.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Section T -- +0.0603%, final accepted figure
Extended Section S's 3-seed/9-pair campaign to 6 seeds/18 pairs (36
cells) per Captain Bob's request for a fuller campaign before moving
on. All 36 cells: 480/480 rows, 0 errors, 17,280/17,280 rows total.
Every one of 18 disabled cells reads exactly 261063 ticks -- CV=0.000%
across all 3 architectures and 6 seeds, zero exceptions. Every enabled
cell's tick count is fully determined by seed alone, identical across
all 3 architectures, zero exceptions. Pooled overhead across all 18
pairs: +0.0603% (mean +0.0603%, stdev 0.0008%, range +0.0598%-
+0.0617%) -- statistically indistinguishable from Section S's 9-pair
figure, now confirmed over double the data with 3 entirely new seeds.
This closes the ACL-TTL overhead measurement line of investigation
(Sections P, Q, R, S, T). +0.0603% is the final accepted figure.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
amd64/99999/enabled, aarch64/24680/disabled added (cell 26 needed a
retry after an unexplained external SIGTERM killed the qemu process
mid-boot -- matches a previously-noted, still-unexplained SIGTERM
recurrence from a process named "claude", first seen 2026-08-18;
1-line stub log from the killed attempt kept as audit trail). All
successful cells: 480/480 rows, 0 errors.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
amd64/24680/disabled, amd64/24680/enabled, aarch64/11111/enabled
added. All 480/480 rows, 0 errors. Cross-arch consistency continues
holding: seed 24680 gives 261219 on both riscv64 and amd64; seed 11111
gives 261222 on both amd64 and aarch64.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Extended from 3 to 6 seeds (added 24680/11111/99999) per Captain Bob's
request for a fuller campaign before moving on. amd64/11111/enabled,
riscv64/24680/enabled, riscv64/24680/disabled added. All 480/480 rows,
0 errors. Disabled-arm determinism (261063) holding across new seeds.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
-- +0.0604% mean, architecture-independent, fully deterministic
All 18 cells (9 arch/seed pairs x disabled/enabled, zuse-authenticated
throughout) complete: 8,640/8,640 rows, 0 errors. Every disabled cell
reads exactly 261063 ticks -- CV=0.000% across all 3 architectures and
3 seeds. Every enabled cell's tick count depends only on seed, identical
across all 3 architectures for a given seed. Pooled overhead: +0.0604%
(mean +0.0604%, stdev 0.0010%, range +0.0598%-+0.0617%).
This is now the accepted ACL-TTL overhead figure for this workload,
superseding Section P's invalidated wall-clock numbers (ACL never
actually armed) and refining Section R's single-pair pilot (+0.0448%,
n=1) to a tight, fully-reproducible, architecture-independent result
across 9 independent pairs.
One tooling bug fixed mid-campaign (cells 1-3): tick-extraction regex
missed the "[Hera] " console-tagger line prefix; underlying VM runs
were unaffected, affected cells' values recovered by hand from their
serial logs before the fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
aarch64/12345/disabled, aarch64/67890/disabled added. All 480/480
rows, 0 errors. Pattern holding: disabled delta=261063 identical
across every arch/seed so far; enabled clusters at 261219/261224
depending on seed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
find and fix the real ACL-TTL measurement bug (zuse session never
authenticated, ACL enforcement never active)
Two mistakes corrected in sequence, both documented in full in
FABRIC-2.md Section R:
1. HEARTBEAT-TICKS@ was swapped to read heartbeat_ticks() -- a newer,
kernel-only ISR hardware-timer counter (src/starkernel/heartbeat.c,
the M5 TIME-TRUST engine) -- based on a misreading of which counter
"the one clock" law refers to. Reverted to vm->heartbeat.tick_count,
Loop #7 "Adaptive Heartrate", the actual year-plus-old counter the
whole physics runtime is built on. Removed the now-irrelevant
HEARTBEAT-PERIOD-NS@ accessor added to diagnose the wrong counter's
adaptive re-arm period. Three-arch QEMU re-acceptance: POST 1012/0/0
on amd64/aarch64/riscv64, HEARTBEAT-TICKS@ confirmed returning 77
(matching the original pre-heartbeat_ticks() acceptance) on all three.
2. The real bug, found after the revert: every "ACL enabled" measurement
in this investigation (Section P's 18-cell campaign, Section Q's
pilot) loaded ACL.4th and ran EXEC-DOE from the bare `ok>` prompt
without ever authenticating a zuse session. repl.c:303 keeps
emergency_console=1 until zuse_session=1; vm_core.c:755 skips the
entire ACL check block (TTL decrement and acl_recheck()) whenever
emergency_console is set. ACL was configured but never armed.
capsules/zuse.4th's pre-existing self-pin bug means the documented
automatic zuse activation doesn't work either (still flagged, not
fixed) -- worked around by invoking the directly-registered
ZUSE-AUTHENTICATE word explicitly.
Validated pilot (amd64, seed 12345, 30 reps, same build, disabled vs.
genuinely zuse-authenticated-enabled): +117 ticks, +0.0448% overhead.
Disabled-arm determinism double-confirmed (261064 ticks, exact repeat
on a fresh boot) -- the 117-tick difference is real signal, not noise.
Reconciles with the original ACL-RWT campaign's own heartbeat-tick
result (+0.0054%-0.0088%, same order of magnitude). Section P's
wall-clock numbers and Section Q's "instrument blind" conclusion are
both marked invalidated/corrected in place, not deleted.
n=1 per arm, one architecture -- not yet a full campaign. Scoped as
next step, not undertaken in this pass.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
vm->heartbeat.tick_count (FORTH-dispatch counter)
Captain Bob's law is unambiguous: the adaptive heartbeat is the one and
only clock, full stop. The first cut of this word read the wrong
counter under that name -- vm->heartbeat.tick_count is a colon-word-
dispatch counter gated at a fixed cadence (frozen during idle, blind to
per-dispatch CPU cost, see FABRIC-2.md Section Q). The real adaptive
heartbeat is heartbeat_ticks() in src/starkernel/heartbeat.c, driven
directly by the ISR-latched 100Hz hardware timer -- genuinely
time-based, confirmed advancing during idle wall-clock time on all
three architectures (amd64 4039->5510, aarch64 6126->7607, riscv64
2965->4466, each over ~15s idle). Kernel build only (__STARKERNEL__);
hosted build has no ISR timer and keeps the old fallback.
Three-arch QEMU acceptance: POST 1012/0/0 on each, word live-tested.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
instrument blind to the effect, not a corrected number
Per Captain Bob's correction (heartbeat tick counter is the sole
canonical clock, not wall-clock), added HEARTBEAT-TICKS@ and re-ran the
ACL-TTL overhead measurement using tick deltas. Diagnostics confirmed
the counter is a FORTH-level colon-word-dispatch counter (frozen at
idle, jumps with real work) -- not a wall-clock proxy. A pilot pair
(amd64, seed 12345, 30 reps, ACL disabled vs enabled) produced
byte-identical deltas (261064 ticks both runs): acl_recheck() runs at
the C dispatch level and doesn't change which/how many colon words
execute, so it's invisible to a counter gated on colon-word-entry
count. Root-caused, not proceeding to the full 18-cell campaign --
every cell would read +0.00% by construction. Section P's wall-clock
numbers stand as the best estimate on record pending a instrument that
can see per-dispatch cost rather than control-flow shape.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
amd64/aarch64/riscv64 QEMU serial logs from the acceptance run for
commit 0b11f92 -- POST 1012/0/0 and HEARTBEAT-TICKS@ live-tested on
each architecture.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The adaptive heartbeat tick counter (vm->heartbeat.tick_count) is the
project's sole canonical clock for timing measurements -- host wall-clock
is not a valid substitute. Exposes it read-only so DoE/overhead campaigns
can measure elapsed ticks instead of wall-clock deltas.
Verified: three-arch QEMU acceptance (amd64/aarch64/riscv64), POST
1012/0/0 on each, HEARTBEAT-TICKS@ live-tested returning a real non-zero
count on all three.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Full 18-cell paired campaign (9 arch/seed pairs x ACL disabled/enabled)
complete: 8,640/8,640 rows, all 16 cfg values x 30 each in every cell,
zero errors. Confirmed capsules/zuse.4th's ACL-ZUSE-BOOT self-pin bug
(self-pin placed inside its own colon definition, causing a genuine
forward-reference failure) is isolated from the core ACL enforcement
mechanism -- verified via live VM state query on multiple cells that
ACL-INIT-PRIMITIVES and EXEC/BYE pinning both complete correctly
regardless.
Result: +5.30% pooled overhead, +4.42% unweighted mean across the 9
pairs (sd 6.77%), paired t=2.043 (df=8) -- not significant at p<0.05.
Documented honestly as a real positive trend that doesn't establish a
precise percentage with confidence, given wall-clock timing's noise
floor is comparable to the effect size -- unlike the original ACL-RWT
campaign's VM-internal tick-counter methodology. A tighter measurement
(more replications, or reading a VM-internal counter directly) is
scoped as a next step, not attempted here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Cells 4-6 complete and verified (480/480 rows, 16/16 cfg coverage, zero
errors each): aarch64/12345 enabled+disabled (a real pair: 310.5s vs
283.0s, +9.7% in the expected direction), riscv64/13579 disabled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
In-progress paired ACL-enabled/disabled campaign (18 cells: 3 seeds x 3
ISAs x 2 ACL states, randomized order, one continuous sitting per
Section O's naming/scoping ruling -- calling this "ACL-TTL overhead",
not "ACL-RWT", since the RWT mechanism no longer exists in the codebase).
Real finding along the way, not blocking: capsules/zuse.4th's
ACL-ZUSE-BOOT places its own self-pin inside its own colon-definition
body instead of after the closing ";", causing a genuine forward-
reference failure at capsule-load time. Confirmed via live VM state
query (EXEC's ACL-MODE@/ACL-PINNED? and DOE-WORK's ACL-MODE@) that this
does NOT affect the core ACL enforcement mechanism itself --
ACL-INIT-PRIMITIVES correctly stamps the whole dictionary, ACL-BOOT
correctly pins EXEC/BYE to STRICT -- so it doesn't invalidate this
measurement. Not fixed, flagged only.
3 cells complete and verified (480/480 rows, 16/16 cfg coverage, zero
errors each): riscv64/12345 disabled+enabled, aarch64/13579 disabled.
Cell 3's timing is mtime-based/approximate rather than precise
wall-clock -- a multi-hour session gap landed inside its measurement
window, contaminating the direct stopwatch reading; the log file's own
last-write mtime is used as a corrected proxy instead, noted as such in
timing.csv.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
ACL-RWT naming/dead-code finding found while scoping the next step
Section O: points to the 127-page deep-dive report and corrected
dataset already committed (dbe4b67, cf5b08f), records the headline
determinism findings, and documents a real finding surfaced while
looking at ACL's current state before implementing a paired
ACL-enabled/disabled measurement -- the "ACL-RWT" name traces to a
Rolling Window of Truth TTL mechanism that ACL.4th's own comments
record as dead code from the day it was written (removed 2026-07-08,
never reachable by the C hot path). The original June 2026 campaign
predates that removal by three weeks, raising a real historical-
accuracy question about what its own overhead numbers measured -- not
settled here, just flagged. Ruled: future paired measurements use an
accurate name ("ACL-TTL overhead") since RWT no longer exists in the
codebase at all.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
every factor interaction, and a raw-data appendix
Expanded the campaign-mechanism validation report from a condensed
6-page summary into the full depth Captain Bob asked for: analyze all 9
cells as a conglomerate Latin square, then dive into each cell's own
data, then cover every within-ISA and cross-factor interaction
explicitly rather than averaging it away.
Report structure (127 pages, compiled clean, no undefined references):
- Front matter: context, methodology, the SWAP-MTX bug narrative
(console-interleaving fix + the Fisher-Yates correctness bug and its
fix, both already committed separately)
- Layer 1: aggregate 3x3 Latin square (heatmap, invariant-metrics table)
- Per-Cell Deep Dive (9 sections): cfg-level distribution, summary
table, and a rep-order execution-trajectory chart per cell -- the
trajectory charts are what actually visualize the order-dependence
finding rather than just stating it
- Per-ISA Deep Dive (3 sections): within-architecture seed comparison
(violin plots, Kruskal-Wallis, per-factor main effects)
- Factor Interactions (6 sections, every pairwise combination of the 4
L8 binary factors): both infer_dec_q and early_exit interaction plots
faceted by architecture, plus the three-way
factor x factor x architecture significance test
- Per-Factor Response (4 sections): linear response by architecture,
with an explicit note that a true quadratic term isn't identifiable
from this 2-level factorial design
- Appendix: full run_id-ordered raw data, all 4,320 rows across all 9
cells, as the primary-source backing for every statistic above
Generated programmatically (analyse_stadium_relaunch_fixed.R for the
aggregate layer, generate_stadium_deepdive.R for the per-cell/per-ISA/
interaction/appendix layers) rather than hand-authored, since content at
this scale needs to be data-driven to stay honest.
Also includes analyse_stadium_relaunch.R, the earlier script built
against the pre-fix (buggy-shuffle) dataset -- superseded but kept for
the record, matching how the underlying data commits were handled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Full 9-cell campaign re-run from scratch (fresh clean build per cell,
fully randomized order, one continuous sitting) using the corrected
Fisher-Yates shuffle. All 9 cells now produce a genuinely valid uniform
permutation: 480/480 rows, all 16 cfg values represented exactly 30 times
each, zero errors -- across all three architectures and all three seeds.
The previous relaunch campaign (experiments/bare_metal/runs/
acl-rwt-20260820/, committed 79d160c) ran against the buggy shuffle and
is superseded by this one for any analysis; kept as-is per policy
(audit artifacts, not deleted), not treated as the canonical dataset.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Found while building the analysis report for the ACL-RWT relaunch
campaign: cfg=0 was missing from run coverage for 2 of 3 seeds, reproduced
identically across all three architectures. Root-caused rather than
worked around, per Captain Bob's "this is worrisome."
SWAP-MTX (capsules/doe.4th Block 2104) never actually swapped two
RUN-MATRIX cells -- it performed a lossy one-way copy (second MATRIX!
call mis-targeted mat[i] again instead of mat[j]). Confirmed by direct
empirical test on the hosted build: INIT-MATRIX gives mat[0]=0, mat[5]=5;
after 0 5 SWAP-MTX, mat[0]=0 (unchanged, should be 5) and mat[5]=0
(correct), with the original value 5 permanently destroyed. Every
Fisher-Yates shuffle this mechanism has ever run silently duplicated some
values and dropped others -- not a true permutation. Not new, not
introduced by item 4.6/Stadium work; predates this session.
Fixed with explicit temp variables (SW-I/SW-J/SW-VI/SW-VJ), trivially
verifiable by inspection over clever stack juggling. Verified on the
hosted build for all three seeds used by the relaunch campaign: each now
produces all 16 cfg values exactly 30 times, run_id 0-479 fully distinct.
Three-arch QEMU acceptance clean: 1012/0/0 POST on all three, identical
dict_hash (expected -- doe.4th isn't C-registered or auto-loaded at
boot). BLOCK_MAP.md correctly shows only doe.4th's own hash changed.
Also includes the R analysis/chart pipeline (analyse_stadium_relaunch.R)
built for the relaunch campaign report, and the three acceptance boot
logs.
Retroactive caveat: the relaunch campaign's own run-matrix coverage
(experiments/bare_metal/runs/acl-rwt-20260820/) is not a valid uniform
permutation, having run against the buggy shuffle. Whether to re-run it
against the fix is a separate call, not made here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Found and fixed a real bug before any campaign work could start: EXEC-DOE's
own CSV output was almost entirely lost to console interleaving with the
routine per-tick heartbeat export -- same bug class as Section L's PLOT
case. Fix: HB-OFF immediately before EXEC-DOE, HB-ON after DOE: complete.
Confirmed HB-ON-first (the reverse order) does NOT fix it -- tested
directly, row loss recurred identically.
Also found: L8-DOE/WL-HI/WL-LO (the mechanism bare_metal/README.md
describes as auto-run) don't exist anywhere in capsules/, and Makefile.
starkernel's DOE_SEED variable is declared but never referenced -- both
vestigial, matching Section K's earlier staleness finding.
Built QEMU-serial-socket injection tooling (socat) to drive EXEC-DOE
interactively after boot, since it requires live REPL input, not just
observation. Two real defects found and fixed in that tooling itself: a
log-discovery race (self-excluding the very log it needed to find,
causing two separate stuck-injector incidents, one overnight) and an
unredirected background launch that deadlocked socat on a full stdout
pipe. Both fixed by having the orchestrator pass exact log/socket paths
directly and always launching through the harness's tracked-background
mechanism.
First full campaign attempt ran all 9 cells as three ISA-blocked loops,
reusing one build per architecture -- caught mid-run: this confounds ISA
with time/session-order, invalidating the Latin square design. Discarded
(logs kept as audit artifacts, not treated as valid data) and re-run
clean: all 9 (arch, seed) cells in fully randomized order, fresh clean
rebuild before every single cell, one continuous sitting. Result:
4,320/4,320 rows captured, zero VM errors anywhere.
This validates the campaign mechanism runs cleanly and reproducibly under
the post-4.6 Stadium substrate -- satisfies item 5.1's own concern that a
green POST suite isn't evidence determinism holds post-migration. It does
NOT produce an ACL-RWT overhead number: ACL.4th is not self-activated in
this repo's default init.4th, so these 9 cells ran with ACL inactive.
Reproducing the original +0.0054%-+0.0088% measurement needs a paired
ACL-enabled/disabled run using this now-validated mechanism -- scoped,
not attempted here.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Closes the last open item in Section H. amd64 and aarch64 were already
confirmed post-quota-grant-fix; riscv64 was pending. Temporarily re-enabled
ART-STRESS-CAMPAIGN (block 4170, disabled since Section L) for this one
headless run, confirmed 30/30 reps / 1500/1500 trials passed with a clean
CAMPAIGN-DONE, then reverted the capsule back to its committed disabled
state (byte-identical to HEAD, mkcapsule --lint clean).
Two SUMMARY lines (reps 4, 15) printed visually garbled from concurrent
[HADES][DOE] console writes -- confirmed cosmetic only by grepping the full
log for refused (result=0) trials: zero matches across all 1500.
Also includes: the two DoE CSV exports and serial logs from this session's
riscv64 runs (audit artifacts per repo convention), and the resulting
Artemis disk image state from real block writes during the stress test.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Captain Bob went with the recommended path rather than measuring now: item
5.1 + ACL-RWT overhead re-measurement stays deferred until Artemis lands,
since Artemis's own storage/timing work would immediately perturb whatever
baseline gets captured today. The other F.3 item (4.4s -> 1.11 -> 4.3 ->
S17.4 chain) is unchanged -- still genuinely blocked on ACL Phase 8, no
ruling needed there.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Disabled capsules/artemis/init.4th block 4170's ART-STRESS-CAMPAIGN -- its
own comment already said to revert to disabled once the K-invariant/
heartbeat verification run (item 4.6, closed earlier this session) was
done. This was the actual ~25-30 minute wall blocking interactive REPL
access, unrelated to any DoE mechanism.
Verified capsules/turtle.4th and capsules/sdk.4th live in a gtk-display
QEMU session: a red hexagon (6 100 POLYGON) and a green self-intersecting
star (100 STAR) both render with correct geometry and color. Screenshot in
evidence/amd64/.
Two real obstacles found and worked around along the way: CS's full-
framebuffer PLOT loop is far slower under TCG than previously documented
(closer to 20+ minutes than "slow"), and the kernel's heartbeat CSV logging
draws to the same console surface PLOT writes pixels to, overwriting
drawings within a fraction of a second unless silenced first with the
existing HB-OFF word. Both HOWTOs updated to record this.
Re-verified full three-arch acceptance boot (POST, DoE, parity) with the
ART-STRESS-CAMPAIGN change: 1012/0/0 and matching dict_hash on all three,
identical to the pre-change baseline.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Loads turtle.4th and doe.4th, defines SDK-VERSION/SDK-HELP into an SDK
vocabulary, then calls FENCE once everything is loaded -- protecting the
base wordset and both cookbook capsules from FORGET. Kernel-only (EXEC
doesn't exist hosted), REPL-invoked via S" sdk.4th" EXEC, not part of
init.4th's boot sequence.
Verified before writing the capsule, not assumed: VOCABULARY/DEFINITIONS
does not actually scope word visibility in this interpreter -- vm_find_word
is a flat dictionary scan that never consults CONTEXT/CURRENT. Documented
plainly in the HOWTO so this isn't mistaken for namespace isolation later.
Block range 5109-5115 -- discovered along the way that user-block space is
capped at [2048, 5120) by mkcapsule, tighter than expected.
Verified: mkcapsule --lint clean, hosted-build trace runs SDK-HELP with
zero attributable VM errors, zero build warnings and identical 1012/0/0
POST results with matching dict_hash on all three kernel architectures.
HOWTO: docs/working/architecture/SDK-HOWTO-20260819.md
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>