G.4 (2h): bounded xHCI event-ring drain fixes boot-attach livelock

Root cause of the G.1 follow-up boot-time attach race: on pathological
controller behavior the xhci_poll_events() drain loop had no hard ceiling.
ERDP is written back only when the loop exits, so the controller cannot
reclaim event TRBs mid-drain; if it keeps producing events the head can
chase the software dequeue pointer forever. xhci_poll_events() never returns,
sk_repl_idle() never reaches its bot_msc_attach_pending check, and a fresh
USB BOT device that finished SET_CONFIGURATION is left flagged-but-never-
attached while the guest appears hung.

Fix: bound the drain to a full ring (XHCI_EVT_RING_MAX_DRAIN = 256), so
xhci_poll_events() always terminates and always writes ERDP each call.
Unprocessed events keep their cycle bit and are re-read next poll; nothing
is dropped. On the healthy path one drain processes only the one-or-few
events the controller posts per chained command, so the bound never triggers
except in the pathological case it breaks.

Beyond the G.1 additions: a new macro in include/starkernel/xhci.h and a
bounded loop in src/starkernel/usb/xhci.c. Builds clean on amd64. Verified
across six consecutive fresh QEMU boots (previously intermittently hung).
This commit is contained in:
Robert Allan James
2026-08-29 09:48:48 -04:00
parent 49a3faa331
commit dc2f38a1e1
5 changed files with 31 additions and 7 deletions
+5 -4
View File
@@ -3333,10 +3333,11 @@ failed"` gate, now with an else-branch for the recoverable-STALL case.
is deferred to hardware (v2.5.0/Artemis bare-metal)**, where a real bad transfer can be
staged. This is the last QEMU-verifiable storage-integrity gap and the recovery logic is in
place; the one thing QEMU cannot prove is the live stall injection itself.
- **Note (pre-existing, NOT G.1):** during verification an intermittent boot-time attach race
was observed (the `sk_repl_idle()` `bot_msc_attach_pending` handoff occasionally does not
progress on a cold QEMU boot, independent of source, with baseline `HEAD` exhibiting it too).
Unrelated to G.1; tracked for a separate follow-up.
- **Follow-up (pre-existing, NOT G.1): ROOT-CAUSED and FIXED 2026-08-29 — see §G.4 below.**
During G.1 verification an intermittent boot-time attach race was observed (the
`sk_repl_idle()` `bot_msc_attach_pending` handoff occasionally does not progress on a cold
QEMU boot, independent of source, with baseline `HEAD` exhibiting it too). Unrelated to G.1;
root cause and fix are documented in §G.4, verified across six consecutive fresh boots.
#### G.2 [v2.0.0] Real-hardware RNG driver plumbing, QEMU-verifiable slice (rest of it lands at v2.5.0)