diff --git a/FABRIC-3.md b/FABRIC-3.md index 2d4eb21..4933d67 100644 --- a/FABRIC-3.md +++ b/FABRIC-3.md @@ -3333,11 +3333,18 @@ failed"` gate, now with an else-branch for the recoverable-STALL case. is deferred to hardware (v2.5.0/Artemis bare-metal)**, where a real bad transfer can be staged. This is the last QEMU-verifiable storage-integrity gap and the recovery logic is in place; the one thing QEMU cannot prove is the live stall injection itself. -- **Follow-up (pre-existing, NOT G.1): ROOT-CAUSED and FIXED 2026-08-29 — see §G.4 below.** - During G.1 verification an intermittent boot-time attach race was observed (the - `sk_repl_idle()` `bot_msc_attach_pending` handoff occasionally does not progress on a cold - QEMU boot, independent of source, with baseline `HEAD` exhibiting it too). Unrelated to G.1; - root cause and fix are documented in §G.4, verified across six consecutive fresh boots. +- **Follow-up (pre-existing, NOT G.1): ROOT-CAUSED and FIXED 2026-08-29.** During G.1 + verification an intermittent boot-time attach race was observed (the `sk_repl_idle()` + `bot_msc_attach_pending` handoff occasionally does not progress on a cold QEMU boot, + independent of source, with baseline `HEAD` exhibiting it too). Unrelated to G.1. Root + cause: `xhci_poll_events()`'s event-ring drain loop had no hard ceiling — ERDP is written + back only when the loop exits, so the controller cannot reclaim event TRBs mid-drain, and + on pathological controller behavior the head can chase the software dequeue pointer + indefinitely, causing the drain (and hence the attach handoff) to livelock. Fixed by bounding + the drain to one full ring (`XHCI_EVT_RING_MAX_DRAIN`, `xhci.c` `xhci_poll_events()`); the + loop always terminates and always writes ERDP each call, and unprocessed events keep their + cycle bit and are re-read next poll (nothing is dropped). Verified across six consecutive + fresh boots. #### G.2 [v2.0.0] Real-hardware RNG driver plumbing, QEMU-verifiable slice (rest of it lands at v2.5.0)