sk_repl_idle() auto-flush: implement Section V's "anything dirty? no? done" check

Makes blk_vm_flush_all() (block_words.c) non-static and declares it in
block_words.h -- it's already the entire implementation behind
SAVE-BUFFERS (block_word_save_buffers() is a one-line wrapper), so
sk_repl_idle() can call the exact same flush path outside word dispatch
without duplicating any logic. Cheap every idle tick regardless of dirty
state: every check inside is a small fixed-size scan, so no separate
pre-check was needed on top of it.

Caught a real bug via a live persistence test before trusting the
feature: the first version gated the flush on sk_repl_get_active_vm()
returning non-NULL, but NULL is that accessor's documented default
(Tripod's own USE-redirect override, "restore default dispatch") --
without an active USE redirect, the flush silently no-op'd for the
entire session. Confirmed live: wrote a byte via BUFFER (no
UPDATE/SAVE-BUFFERS), waited past the idle cadence, killed QEMU abruptly,
rebooted with the same disk image, read back 0 instead of the written
65. Fixed by threading the VM sk_repl_run()'s own loop already resolves
each iteration (g_repl_active_vm ? g_repl_active_vm : vm) down as a
parameter through sk_readline() into sk_repl_idle(), rather than trying
to re-derive it from an accessor with the wrong default. Re-ran the same
test after the fix: read back 65, matching the written byte -- the write
survived an abrupt kill with no explicit flush call anywhere in the
test, proving the idle-tick auto-flush genuinely ran.

All three architectures re-verified clean. FABRIC-2.md Section V item 6
and the corresponding Milestone 3 punch-list item marked done.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
This commit is contained in:
Robert Allan James
2026-08-25 14:51:02 -04:00
co-authored by Claude Sonnet 5
parent 1b80609cb0
commit 8eaefeb9ee
20 changed files with 63371 additions and 11 deletions
+41 -3
View File
@@ -2603,7 +2603,42 @@ chain mechanism that already exists, not a new design.
separate, deliberately coarser cadence for higher-level subsystem dispatch"* — and it is
currently a no-op placeholder (confirmed empty during the Section R investigation into the
ACL-TTL heartbeat bug). This is the natural home for a cheap "anything dirty? no? done"
block-sync check.
block-sync check. **Done 2026-08-25** — see writeup below.
**Implementation.** `blk_vm_flush_all()` (`block_words.c`) — already the entire
implementation behind `SAVE-BUFFERS`, `block_word_save_buffers()` is a one-line wrapper
around it — made non-`static` and declared in `block_words.h`, so `sk_repl_idle()` can call
the exact same flush path outside word dispatch rather than duplicating any of its logic.
Cheap to call every idle tick regardless of whether anything is actually dirty: every check
inside is a small fixed-size scan (`BLK_VM_SLOTS` here, `DISK_CACHE_SLOTS` per device inside
`blk_flush()`), so no separate "is anything dirty" pre-check was needed — the existing
function already *is* that cheap early-exit, no new logic required.
**A real bug, caught by the same live test that proved the feature.** The first version
called `sk_repl_get_active_vm()` from inside `sk_repl_idle()` and skipped the flush unless it
returned non-NULL. That accessor is Tripod's own `USE`-redirect override — NULL is its
*documented default* ("restore default dispatch (NULL = use REPL's own vm)"), not "no VM is
active." Since nothing in this test ever ran `USE`, the flush silently no-op'd for the
session's entire duration — confirmed live via a real persistence test (write a byte via
`BUFFER` with no `UPDATE`/`SAVE-BUFFERS`, wait past the idle cadence, kill QEMU abruptly,
reboot with the same disk image, read the byte back: got `0`, not the written `65`). Root
cause: `sk_repl_run()`'s own loop already resolves the correct VM every iteration
(`active = g_repl_active_vm ? g_repl_active_vm : vm`, right before calling `sk_readline()`)
`sk_repl_idle()` just had no way to see that resolution, since it's nested two calls deep
(`sk_repl_run()``sk_readline()``sk_repl_idle()`) with no VM parameter threaded through
either intermediate function. Fixed by threading the already-resolved VM down as a parameter:
`sk_readline()` gained a `VM *active_vm` argument, `sk_repl_idle()` gained one too (replacing
its own `sk_repl_get_active_vm()` call entirely), and both of `sk_readline()`'s callers now
pass the right value — `sk_repl_run()`'s own `active`, and `sk_repl_step()`'s own `vm`
parameter (a second, simpler REPL entry point with no `USE`-redirection concept at all).
Re-running the exact same persistence test after the fix: read back `65`, matching the
written `0x41` — the write survived an abrupt kill with no explicit flush call anywhere in
the test, proving the idle-tick auto-flush genuinely ran during the wait.
`logs/20260825-142510/amd64/` + `logs/20260825-142714/amd64/` are the failing before/after
reboot pair (kept as evidence of the bug, not deleted); `logs/20260825-143546/amd64/` +
`logs/20260825-143745/amd64/` are the same pair after the fix. All three architectures
re-verified clean with no dirty state pending: `logs/20260825-144008/amd64/`,
`logs/20260825-144126/aarch64/`, `logs/20260825-144517/riscv64/`.
**Not started:** no code, no design doc, no capsule work. This section exists so the next
session can pick up from an accurate baseline rather than re-deriving the shape from scratch.
@@ -3979,9 +4014,12 @@ sized around, not the size of every test image.
get its own independent implementation
- [ ] Implement the block-migration function itself (move one block's content + BAM entry
between two attached devices)
- [ ] Implement the `sk_repl_idle()` body — the cheap "anything dirty? no? done" check
- [x] Implement the `sk_repl_idle()` body — the cheap "anything dirty? no? done" check
(Section V confirmed this hook is empty and ready right now, doesn't even need
Milestone 2 to be written, only to be *tested end to end*)
Milestone 2 to be written, only to be *tested end to end*). **Done 2026-08-25**, see
Section V item 6's writeup — live-verified via a real abrupt-kill/reboot persistence
test, which also caught and fixed a real bug (`sk_repl_get_active_vm()`'s NULL default
silently no-op'ing the flush)
- [ ] Decide and implement unclean-removal handling (Section U's explicitly flagged open
question — never answered) — at minimum, detect a mid-flush disconnect via
Milestone 2e's disconnect signal and decide what state that leaves affected blocks in