sk_repl_idle() auto-flush: implement Section V's "anything dirty? no? done" check
Makes blk_vm_flush_all() (block_words.c) non-static and declares it in block_words.h -- it's already the entire implementation behind SAVE-BUFFERS (block_word_save_buffers() is a one-line wrapper), so sk_repl_idle() can call the exact same flush path outside word dispatch without duplicating any logic. Cheap every idle tick regardless of dirty state: every check inside is a small fixed-size scan, so no separate pre-check was needed on top of it. Caught a real bug via a live persistence test before trusting the feature: the first version gated the flush on sk_repl_get_active_vm() returning non-NULL, but NULL is that accessor's documented default (Tripod's own USE-redirect override, "restore default dispatch") -- without an active USE redirect, the flush silently no-op'd for the entire session. Confirmed live: wrote a byte via BUFFER (no UPDATE/SAVE-BUFFERS), waited past the idle cadence, killed QEMU abruptly, rebooted with the same disk image, read back 0 instead of the written 65. Fixed by threading the VM sk_repl_run()'s own loop already resolves each iteration (g_repl_active_vm ? g_repl_active_vm : vm) down as a parameter through sk_readline() into sk_repl_idle(), rather than trying to re-derive it from an accessor with the wrong default. Re-ran the same test after the fix: read back 65, matching the written byte -- the write survived an abrupt kill with no explicit flush call anywhere in the test, proving the idle-tick auto-flush genuinely ran. All three architectures re-verified clean. FABRIC-2.md Section V item 6 and the corresponding Milestone 3 punch-list item marked done. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
1b80609cb0
commit
8eaefeb9ee
+41
-3
@@ -2603,7 +2603,42 @@ chain mechanism that already exists, not a new design.
|
||||
separate, deliberately coarser cadence for higher-level subsystem dispatch"* — and it is
|
||||
currently a no-op placeholder (confirmed empty during the Section R investigation into the
|
||||
ACL-TTL heartbeat bug). This is the natural home for a cheap "anything dirty? no? done"
|
||||
block-sync check.
|
||||
block-sync check. **Done 2026-08-25** — see writeup below.
|
||||
|
||||
**Implementation.** `blk_vm_flush_all()` (`block_words.c`) — already the entire
|
||||
implementation behind `SAVE-BUFFERS`, `block_word_save_buffers()` is a one-line wrapper
|
||||
around it — made non-`static` and declared in `block_words.h`, so `sk_repl_idle()` can call
|
||||
the exact same flush path outside word dispatch rather than duplicating any of its logic.
|
||||
Cheap to call every idle tick regardless of whether anything is actually dirty: every check
|
||||
inside is a small fixed-size scan (`BLK_VM_SLOTS` here, `DISK_CACHE_SLOTS` per device inside
|
||||
`blk_flush()`), so no separate "is anything dirty" pre-check was needed — the existing
|
||||
function already *is* that cheap early-exit, no new logic required.
|
||||
|
||||
**A real bug, caught by the same live test that proved the feature.** The first version
|
||||
called `sk_repl_get_active_vm()` from inside `sk_repl_idle()` and skipped the flush unless it
|
||||
returned non-NULL. That accessor is Tripod's own `USE`-redirect override — NULL is its
|
||||
*documented default* ("restore default dispatch (NULL = use REPL's own vm)"), not "no VM is
|
||||
active." Since nothing in this test ever ran `USE`, the flush silently no-op'd for the
|
||||
session's entire duration — confirmed live via a real persistence test (write a byte via
|
||||
`BUFFER` with no `UPDATE`/`SAVE-BUFFERS`, wait past the idle cadence, kill QEMU abruptly,
|
||||
reboot with the same disk image, read the byte back: got `0`, not the written `65`). Root
|
||||
cause: `sk_repl_run()`'s own loop already resolves the correct VM every iteration
|
||||
(`active = g_repl_active_vm ? g_repl_active_vm : vm`, right before calling `sk_readline()`)
|
||||
— `sk_repl_idle()` just had no way to see that resolution, since it's nested two calls deep
|
||||
(`sk_repl_run()` → `sk_readline()` → `sk_repl_idle()`) with no VM parameter threaded through
|
||||
either intermediate function. Fixed by threading the already-resolved VM down as a parameter:
|
||||
`sk_readline()` gained a `VM *active_vm` argument, `sk_repl_idle()` gained one too (replacing
|
||||
its own `sk_repl_get_active_vm()` call entirely), and both of `sk_readline()`'s callers now
|
||||
pass the right value — `sk_repl_run()`'s own `active`, and `sk_repl_step()`'s own `vm`
|
||||
parameter (a second, simpler REPL entry point with no `USE`-redirection concept at all).
|
||||
Re-running the exact same persistence test after the fix: read back `65`, matching the
|
||||
written `0x41` — the write survived an abrupt kill with no explicit flush call anywhere in
|
||||
the test, proving the idle-tick auto-flush genuinely ran during the wait.
|
||||
`logs/20260825-142510/amd64/` + `logs/20260825-142714/amd64/` are the failing before/after
|
||||
reboot pair (kept as evidence of the bug, not deleted); `logs/20260825-143546/amd64/` +
|
||||
`logs/20260825-143745/amd64/` are the same pair after the fix. All three architectures
|
||||
re-verified clean with no dirty state pending: `logs/20260825-144008/amd64/`,
|
||||
`logs/20260825-144126/aarch64/`, `logs/20260825-144517/riscv64/`.
|
||||
|
||||
**Not started:** no code, no design doc, no capsule work. This section exists so the next
|
||||
session can pick up from an accurate baseline rather than re-deriving the shape from scratch.
|
||||
@@ -3979,9 +4014,12 @@ sized around, not the size of every test image.
|
||||
get its own independent implementation
|
||||
- [ ] Implement the block-migration function itself (move one block's content + BAM entry
|
||||
between two attached devices)
|
||||
- [ ] Implement the `sk_repl_idle()` body — the cheap "anything dirty? no? done" check
|
||||
- [x] Implement the `sk_repl_idle()` body — the cheap "anything dirty? no? done" check
|
||||
(Section V confirmed this hook is empty and ready right now, doesn't even need
|
||||
Milestone 2 to be written, only to be *tested end to end*)
|
||||
Milestone 2 to be written, only to be *tested end to end*). **Done 2026-08-25**, see
|
||||
Section V item 6's writeup — live-verified via a real abrupt-kill/reboot persistence
|
||||
test, which also caught and fixed a real bug (`sk_repl_get_active_vm()`'s NULL default
|
||||
silently no-op'ing the flush)
|
||||
- [ ] Decide and implement unclean-removal handling (Section U's explicitly flagged open
|
||||
question — never answered) — at minimum, detect a mid-flush disconnect via
|
||||
Milestone 2e's disconnect signal and decide what state that leaves affected blocks in
|
||||
|
||||
Reference in New Issue
Block a user