Files
LithosAnanake/disk
Robert Allan JamesandClaude Sonnet 5 2b16daba16 Artemis Milestone 2d: xHCI Event Ring servicing, polled not interrupt-driven
Implements Event Ring TRB parsing and ERDP dequeue-pointer update
(xhci_poll_events(), src/starkernel/usb/xhci.c), called from
sk_repl_idle()'s existing ~1s idle cadence rather than a per-arch
interrupt handler.

A first attempt wired real interrupt delivery (PCI->IOAPIC GSI routing,
a dedicated isr_stub34/vector 0x22, GIC/PLIC routing mirroring
virtio_input.c). Checked live via QMP query-pci before trusting it: the
amd64 PIRQ swizzle formula predicted GSI 16 for the xHCI controller at
PCI slot 4; the real QEMU-assigned IRQ was 10, and embedded ICH9
functions contradicted the same formula too. Reverted all of it back to
the exact committed baseline rather than chasing chipset PIRQ routing
further, and reframed around Section U item 6's own design intent
("interrupt-driven, coarse cadence, cheap early-exit... quick check
blocks... done") via sk_repl_idle() instead -- USB insertion is a
human-timescale event, not a hot path.

Added -device qemu-xhci to all three QEMU launch targets (required for
any of this to be testable). Verified end to end via genuine post-boot
hotplug (QMP device_add/device_del usb-storage): all three architectures
detect a live attach within seconds. A false-alarm heartbeat "freeze"
found mid-verification traced to querying the wrong counter
(vm->heartbeat.tick_count, which only advances during word execution,
not the kernel's real ISR-driven heartbeat_ticks()) -- confirmed via a
temporary diagnostic word, captured and reverted.

Full writeup, including the discarded interrupt-routing attempt and the
false-alarm investigation, in FABRIC-2.md's Milestone 2c/2d entries.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HZ8kNoTuP63pbQtro4qvrm
2026-08-22 09:25:25 -04:00
..

disk/

QEMU disk images used for Artemis (block-storage VM) persistence testing. Mounted via Makefile.starkernel's ARTDISK variable (default disk/artemis.img) as a virtio-blk-pci device on all three architectures' kernel QEMU boots — not scripts/rundisk.sh, which targets a separate, currently-unused disks/ (plural) directory for the hosted VM's --disk-img= flag instead.

  • artemis.img — standard Artemis persistence test disk. Reformatted 2026-08-02: this image had been stuck in a corrupted state (valid LithosAnanke magic header, but data not matching what ART-READ-TEST expects) since before this repo's own git history begins (git log shows it already broken at the initial commit, carried over from the pre-split monorepo). Chasing down the resulting persistent FAIL: persist-read traced to the data, not the code — the write→reboot→read round trip works correctly on a fresh image (see artemis-debug-roundtrip.img below). Reformatted by blanking the file and letting a normal boot format+write-test it; verified PASS: persist-read on amd64, aarch64, and riscv64 against the same image afterward (cross-arch resume, matching the arch-neutral on-disk format .claude/ARTEMIS.md specifies).
  • artemis-debug-roundtrip.img — round-trip regression fixture created during that investigation. Known-good: format → self-test → write-test → reboot → resume → PASS: persist-read, confirmed 3 times in a row. Keep this in a passing state; if a future change breaks it, that's a real regression, not a stale-fixture artifact like artemis.img was.
  • artemis-persist-test.img — persistence round-trip test image (pre-existing; history/state not re-verified during the above investigation).
  • artemis-poison.img — a separate test image (exact scenario not documented elsewhere in the repo as of this writing; name suggests an adversarial/corruption test, not confirmed).
  • artemis-unrecognized-test.img — exercises ART-HALT-UNRECOG (.claude/ARTEMIS.md acceptance criterion #6). Regenerated 2026-08-02: the previous copy of this file had itself been silently reformatted by a since-fixed bug in the generic block subsystem (src/block_subsystem.c) — it carried a valid low-level 'STFR'/v2 header despite being meant to represent foreign disk content, direct forensic evidence of the bug described in .claude/ARTEMIS.md's Build Status item 6. Regenerated as 30MB of a repeating POISON-UNRECOGNIZED-DISK-TEST-FIXTURE--NOT-BLANK-NOT-STFR-NOT-ARTEMIS-- ASCII pattern — deliberately neither blank, nor the block subsystem's own 'STFR' magic, nor Artemis's "ARTEMIS\0" marker. Verified on amd64 and riscv64 post-fix: boot correctly halts (ARTEMIS HALT: unrecognized disk content) and the file's sha256 is now byte-for-byte identical before and after boot. Keep this fixture in this poisoned state — if a future change makes its sha256 change across a boot, that is exactly the regression this fixture exists to catch.

Note on incidental header churn: artemis.img picks up a few changed header bytes on every ordinary boot even though no user data changes — blk_subsys_attach_device() always records a fresh mounted_time on a successfully recognized disk, which gets flushed at shutdown. This is expected bookkeeping, not a bug; revert it before committing rather than carrying timestamp noise in git history.

These are regenerable QEMU raw disk images, not source — see .claude/ARTEMIS.md for the storage model they exercise.