Artemis Milestone 2h: hot-detach -- 2h complete

blk_subsys_detach_device() (block_subsystem.c) walks the device chain,
refuses removal of anything but the current tail (a mid-chain removal
would corrupt every later slot's start_lbn -- this architecture's own doc
already argues USB stays last specifically to avoid that), unlinks,
shrinks total_user_lbn, closes and frees the slot. Discards rather than
flushes dirty state -- the device is physically gone by the time this
runs (PORTSC disconnect only). Trigger wiring mirrors the attach path:
bot_msc_attached (set only once attach actually succeeds) gates a new
bot_msc_detach_pending flag set at PORTSC disconnect (not Disable Slot
completion, which is conditionally skipped and would miss concurrent
connect/disconnect pairs), consumed in sk_repl_idle().

Advisor flagged the real hazard ahead of time: block_words.c's VM block
window (blk_vm_lbn[]/blk_vm_cbuf[]) can go stale across a detach then a
same-LBN re-attach, and suggested a pointer-identity re-check in
blk_vm_load() as a minimal fix. That fix was implemented, then directly
falsified by its own designed-for-this test: attach a blank device, read
a block (populating the cache), detach, re-attach a device with distinct
content at the identical LBN, read again -- served stale content from
the first device. Root cause, confirmed live: glibc's allocator hands
free(slot) straight back to the very next same-size calloc(), so the
"fresh" and stale pointers were bitwise identical despite being two
different devices. Fixed properly with a monotonic blk_subsys_epoch()
counter (bumped on every attach/detach) checked by a new
blk_vm_check_epoch() helper at the one choke point (blk_vm_find(), plus
blk_vm_flush_all() which reads the same arrays directly) that covers
every path touching the window cache -- unfooled by address reuse.

Verified live with a new disk/usb-thumbdrive-test2.img fixture (distinct
content from the existing blank test image): attach A, read (cache hit
populated), detach, re-attach B at the same LBN, read again -- correctly
ran a fresh device read and returned B's real content, not A's stale
cached zeros. The failing pointer-comparison attempt's own capture log
kept as evidence, not deleted. All three architectures re-verified clean.
FABRIC-2.md Section X 2h marked complete -- enumeration through
hot-detach all live and verified; only WRITE(10) (2g's own still-open
item) remains unimplemented in the driver, not blocking anything here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
This commit is contained in:
Robert Allan James
2026-08-25 14:10:05 -04:00
co-authored by Claude Sonnet 5
parent 3b085dd875
commit af267a52a6
22 changed files with 46979 additions and 13 deletions
+13 -5
View File
@@ -58,11 +58,19 @@ carrying timestamp noise in git history.
- `usb-thumbdrive-test.img` — 64MB raw image backing a QEMU `usb-storage`
device attached to the xHCI controller's bus (`xhci0.0`) for Milestone 2e/
2h hotplug testing, added 2026-08-22. Not yet formatted with any
LithosAnanke/Artemis header — at this point in the driver's development
it only needs to exist as a backing store for a live Port Status Change
event; formatting comes once `blkio_usb.c` and `blk_subsys_attach_device()`
wiring exist (Milestone 2h).
2h hotplug testing, added 2026-08-22. Blank (all zero) — `blkio_usb.c` +
`blk_subsys_attach_device()` wiring (Milestone 2h, done 2026-08-25) attach
it as `BLK_FMT_PROVISIONAL` every time, which is the intended, exercised
state; not yet `BLK_FMT_FORMATTED` via `BLK-CONFIRM-FORMAT`.
- `usb-thumbdrive-test2.img` — 64MB raw image, added 2026-08-25 for
Milestone 2h hot-detach/re-attach verification. Filled with a repeating
`HOTDETACH-REATTACH-FIXTURE-2026-08-25--` ASCII pattern, deliberately
distinguishable from `usb-thumbdrive-test.img`'s all-zero content — the
point is proving a block read *after* detaching `usb-thumbdrive-test.img`
and re-attaching this one actually returns this pattern rather than
silently replaying the old device's cached (all-zero) content, which is
exactly the class of bug a same-LBN-range device swap can cause if the
VM block window cache (`vm->blk_vm_cbuf[]`) isn't re-validated on a hit.
**Convention, standing as of 2026-08-22: every virtual disk/thumb-drive image
used for testing — Artemis persistence disks above, and USB Mass Storage