Files
LithosAnanake/disk
Robert Allan JamesandClaude Sonnet 5 00e657019e stadium: make VM population bound RAM-derived, not a static array of 4
Replaces STADIUM_MAX_VM_COUNT (Kconfig, hardcoded default 4) with a
boot-time computation, mirroring the pattern stadium_boot_init() already
used for the cell pool. New Kconfig STADIUM_VM_MEMORY_PERCENT (default
50): max_vm_count = (kmalloc_get_stats().free_bytes after the cell array
* STADIUM_VM_MEMORY_PERCENT / 100) / VM_MEMORY_SIZE, floored to 1, no
ceiling (population is not knowable in advance - could be 4, could be
4000). stadium_quotas and word_slots (plus stat_promotions/stat_evictions)
are now kmalloc'd to the computed count instead of declared with a macro.
New accessor stadium_max_vm_count() replaces every STADIUM_MAX_VM_COUNT
reference, including capsule_birth.c's birth-refusal gate.

Two things found and fixed along the way:

- The existing cell-pool budget was sourced from pmm_get_stats(), which
  reflects physical pages PMM hasn't handed to any subsystem yet - but
  the actual allocation is kmalloc(), which draws from the separate,
  fixed-size heap kmalloc_init() (M6) already carved out of PMM before
  stadium_boot_init() ever runs. Budgeting against PMM's leftover and
  allocating from the kmalloc heap are two different pools. Both the
  cell budget and the new VM-count budget now source from
  kmalloc_get_stats() instead.

- stadium_owner[] (which VM's quota owns each cell) was uint8_t, capped
  at 255 slots by a compile-time assert tied to the old macro. Widened
  to uint16_t (65535 slots of headroom) with a runtime clamp + log if
  the computed count ever exceeds that, since there's no ceiling anymore.

Three-arch QEMU acceptance: all clean to ok>, computed VM count genuinely
differs by actual available RAM (amd64/riscv64: 50 slots at -m 1024,
aarch64: 101 slots), Stadium conservation invariant identical across all
three (resident_sum=43691 reservoir=21845 sum=65536).
logs/20260815-080526/amd64, logs/20260815-080826/aarch64,
logs/20260815-080952/riscv64.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-15 08:11:21 -04:00
..

disk/

QEMU disk images used for Artemis (block-storage VM) persistence testing. Mounted via Makefile.starkernel's ARTDISK variable (default disk/artemis.img) as a virtio-blk-pci device on all three architectures' kernel QEMU boots — not scripts/rundisk.sh, which targets a separate, currently-unused disks/ (plural) directory for the hosted VM's --disk-img= flag instead.

  • artemis.img — standard Artemis persistence test disk. Reformatted 2026-08-02: this image had been stuck in a corrupted state (valid LithosAnanke magic header, but data not matching what ART-READ-TEST expects) since before this repo's own git history begins (git log shows it already broken at the initial commit, carried over from the pre-split monorepo). Chasing down the resulting persistent FAIL: persist-read traced to the data, not the code — the write→reboot→read round trip works correctly on a fresh image (see artemis-debug-roundtrip.img below). Reformatted by blanking the file and letting a normal boot format+write-test it; verified PASS: persist-read on amd64, aarch64, and riscv64 against the same image afterward (cross-arch resume, matching the arch-neutral on-disk format .claude/ARTEMIS.md specifies).
  • artemis-debug-roundtrip.img — round-trip regression fixture created during that investigation. Known-good: format → self-test → write-test → reboot → resume → PASS: persist-read, confirmed 3 times in a row. Keep this in a passing state; if a future change breaks it, that's a real regression, not a stale-fixture artifact like artemis.img was.
  • artemis-persist-test.img — persistence round-trip test image (pre-existing; history/state not re-verified during the above investigation).
  • artemis-poison.img — a separate test image (exact scenario not documented elsewhere in the repo as of this writing; name suggests an adversarial/corruption test, not confirmed).
  • artemis-unrecognized-test.img — exercises ART-HALT-UNRECOG (.claude/ARTEMIS.md acceptance criterion #6). Regenerated 2026-08-02: the previous copy of this file had itself been silently reformatted by a since-fixed bug in the generic block subsystem (src/block_subsystem.c) — it carried a valid low-level 'STFR'/v2 header despite being meant to represent foreign disk content, direct forensic evidence of the bug described in .claude/ARTEMIS.md's Build Status item 6. Regenerated as 30MB of a repeating POISON-UNRECOGNIZED-DISK-TEST-FIXTURE--NOT-BLANK-NOT-STFR-NOT-ARTEMIS-- ASCII pattern — deliberately neither blank, nor the block subsystem's own 'STFR' magic, nor Artemis's "ARTEMIS\0" marker. Verified on amd64 and riscv64 post-fix: boot correctly halts (ARTEMIS HALT: unrecognized disk content) and the file's sha256 is now byte-for-byte identical before and after boot. Keep this fixture in this poisoned state — if a future change makes its sha256 change across a boot, that is exactly the regression this fixture exists to catch.

Note on incidental header churn: artemis.img picks up a few changed header bytes on every ordinary boot even though no user data changes — blk_subsys_attach_device() always records a fresh mounted_time on a successfully recognized disk, which gets flushed at shutdown. This is expected bookkeeping, not a bug; revert it before committing rather than carrying timestamp noise in git history.

These are regenerable QEMU raw disk images, not source — see .claude/ARTEMIS.md for the storage model they exercise.