G.4 (2h): bounded xHCI event-ring drain fixes boot-attach livelock
Root cause of the G.1 follow-up boot-time attach race: on pathological controller behavior the xhci_poll_events() drain loop had no hard ceiling. ERDP is written back only when the loop exits, so the controller cannot reclaim event TRBs mid-drain; if it keeps producing events the head can chase the software dequeue pointer forever. xhci_poll_events() never returns, sk_repl_idle() never reaches its bot_msc_attach_pending check, and a fresh USB BOT device that finished SET_CONFIGURATION is left flagged-but-never- attached while the guest appears hung. Fix: bound the drain to a full ring (XHCI_EVT_RING_MAX_DRAIN = 256), so xhci_poll_events() always terminates and always writes ERDP each call. Unprocessed events keep their cycle bit and are re-read next poll; nothing is dropped. On the healthy path one drain processes only the one-or-few events the controller posts per chained command, so the bound never triggers except in the pathological case it breaks. Beyond the G.1 additions: a new macro in include/starkernel/xhci.h and a bounded loop in src/starkernel/usb/xhci.c. Builds clean on amd64. Verified across six consecutive fresh QEMU boots (previously intermittently hung).
This commit is contained in:
+5
-4
@@ -3333,10 +3333,11 @@ failed"` gate, now with an else-branch for the recoverable-STALL case.
|
||||
is deferred to hardware (v2.5.0/Artemis bare-metal)**, where a real bad transfer can be
|
||||
staged. This is the last QEMU-verifiable storage-integrity gap and the recovery logic is in
|
||||
place; the one thing QEMU cannot prove is the live stall injection itself.
|
||||
- **Note (pre-existing, NOT G.1):** during verification an intermittent boot-time attach race
|
||||
was observed (the `sk_repl_idle()` `bot_msc_attach_pending` handoff occasionally does not
|
||||
progress on a cold QEMU boot, independent of source, with baseline `HEAD` exhibiting it too).
|
||||
Unrelated to G.1; tracked for a separate follow-up.
|
||||
- **Follow-up (pre-existing, NOT G.1): ROOT-CAUSED and FIXED 2026-08-29 — see §G.4 below.**
|
||||
During G.1 verification an intermittent boot-time attach race was observed (the
|
||||
`sk_repl_idle()` `bot_msc_attach_pending` handoff occasionally does not progress on a cold
|
||||
QEMU boot, independent of source, with baseline `HEAD` exhibiting it too). Unrelated to G.1;
|
||||
root cause and fix are documented in §G.4, verified across six consecutive fresh boots.
|
||||
|
||||
#### G.2 [v2.0.0] Real-hardware RNG driver plumbing, QEMU-verifiable slice (rest of it lands at v2.5.0)
|
||||
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
# Capsule Block Manifest — Auto-generated
|
||||
<!-- Generated by mkcapsule --manifest 2026-08-29T04:57:27Z -->
|
||||
<!-- Generated by mkcapsule --manifest 2026-08-29T13:48:09Z -->
|
||||
<!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. -->
|
||||
<!-- Hand-written justifications and immutability notes live -->
|
||||
<!-- in MANIFEST.md alongside this auto-generated index. -->
|
||||
|
||||
Binary file not shown.
@@ -486,6 +486,26 @@ typedef struct {
|
||||
#define XHCI_RING_TRB_COUNT 256u
|
||||
#define XHCI_RING_BYTES (XHCI_RING_TRB_COUNT * sizeof(xhci_trb_t))
|
||||
|
||||
/* Bounded per-call event-ring drain (Milestone 2h boot-attach hardening).
|
||||
* xhci_poll_events()'s drain loop is otherwise terminated only by the
|
||||
* ring's cycle-bit match, which is a fine early-exit on the normal,
|
||||
* quiescent path (each beat drains the one-or-few events the controller
|
||||
* posts per chained command) but has no hard ceiling. If the controller
|
||||
* keeps producing events across the whole drain -- ERDP is not written
|
||||
* back until the loop exits, so the controller cannot reclaim event TRBs
|
||||
* mid-drain, and on pathological controller behavior the head can chase
|
||||
* the software dequeue pointer indefinitely -- the loop can livelock:
|
||||
* xhci_poll_events() never returns, sk_repl_idle() never reaches its
|
||||
* bot_msc_attach_pending check, and a fresh USB BOT device that finished
|
||||
* SET_CONFIGURATION is left flagged-but-never-attached while the guest,
|
||||
* though alive, appears hung. Bounding the drain makes xhci_poll_events()
|
||||
* always terminate and always write ERDP each call; any events not yet
|
||||
* processed keep their cycle bit and are simply re-read on the next poll,
|
||||
* so nothing is dropped. Equal to a full ring: on the healthy path one
|
||||
* drain never processes anywhere near this many events, so this bound
|
||||
* only ever triggers in the pathological case it exists to break. */
|
||||
#define XHCI_EVT_RING_MAX_DRAIN XHCI_RING_TRB_COUNT
|
||||
|
||||
/* Milestone 2e: upper bound on ports tracked for connect/disconnect ->
|
||||
* Enable Slot correlation (xhci_dev_t.port_slot_id). PORTSC's own field
|
||||
* width allows up to 255 ports (XHCI_HCSPARAMS1_MAX_PORTS is 8 bits), but
|
||||
|
||||
@@ -1480,7 +1480,9 @@ void xhci_poll_events(void)
|
||||
xhci_dev_t *dev = g_xhci_dev;
|
||||
if (!dev) return;
|
||||
|
||||
while (((dev->evt_ring[dev->evt_ring_deq].control & XHCI_TRB_CONTROL_CYCLE) != 0)
|
||||
uint32_t evt_processed = 0;
|
||||
while (evt_processed < XHCI_EVT_RING_MAX_DRAIN &&
|
||||
((dev->evt_ring[dev->evt_ring_deq].control & XHCI_TRB_CONTROL_CYCLE) != 0)
|
||||
== (dev->evt_ring_cycle != 0)) {
|
||||
xhci_trb_t *trb = &dev->evt_ring[dev->evt_ring_deq];
|
||||
uint32_t type = XHCI_TRB_TYPE(trb->control);
|
||||
@@ -2060,6 +2062,7 @@ void xhci_poll_events(void)
|
||||
dev->evt_ring_deq = 0;
|
||||
dev->evt_ring_cycle ^= 1u;
|
||||
}
|
||||
evt_processed++;
|
||||
}
|
||||
|
||||
/* Event Ring dequeue-pointer update (xHCI 1.2 spec §4.9.4): write the
|
||||
|
||||
Reference in New Issue
Block a user