sk_repl_idle() auto-flush: implement Section V's "anything dirty? no? done" check

Makes blk_vm_flush_all() (block_words.c) non-static and declares it in
block_words.h -- it's already the entire implementation behind
SAVE-BUFFERS (block_word_save_buffers() is a one-line wrapper), so
sk_repl_idle() can call the exact same flush path outside word dispatch
without duplicating any logic. Cheap every idle tick regardless of dirty
state: every check inside is a small fixed-size scan, so no separate
pre-check was needed on top of it.

Caught a real bug via a live persistence test before trusting the
feature: the first version gated the flush on sk_repl_get_active_vm()
returning non-NULL, but NULL is that accessor's documented default
(Tripod's own USE-redirect override, "restore default dispatch") --
without an active USE redirect, the flush silently no-op'd for the
entire session. Confirmed live: wrote a byte via BUFFER (no
UPDATE/SAVE-BUFFERS), waited past the idle cadence, killed QEMU abruptly,
rebooted with the same disk image, read back 0 instead of the written
65. Fixed by threading the VM sk_repl_run()'s own loop already resolves
each iteration (g_repl_active_vm ? g_repl_active_vm : vm) down as a
parameter through sk_readline() into sk_repl_idle(), rather than trying
to re-derive it from an accessor with the wrong default. Re-ran the same
test after the fix: read back 65, matching the written byte -- the write
survived an abrupt kill with no explicit flush call anywhere in the
test, proving the idle-tick auto-flush genuinely ran.

All three architectures re-verified clean. FABRIC-2.md Section V item 6
and the corresponding Milestone 3 punch-list item marked done.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
This commit is contained in:
Robert Allan James
2026-08-25 14:51:02 -04:00
co-authored by Claude Sonnet 5
parent 1b80609cb0
commit 8eaefeb9ee
20 changed files with 63371 additions and 11 deletions
+41 -3
View File
@@ -2603,7 +2603,42 @@ chain mechanism that already exists, not a new design.
separate, deliberately coarser cadence for higher-level subsystem dispatch"* — and it is
currently a no-op placeholder (confirmed empty during the Section R investigation into the
ACL-TTL heartbeat bug). This is the natural home for a cheap "anything dirty? no? done"
block-sync check.
block-sync check. **Done 2026-08-25** — see writeup below.
**Implementation.** `blk_vm_flush_all()` (`block_words.c`) — already the entire
implementation behind `SAVE-BUFFERS`, `block_word_save_buffers()` is a one-line wrapper
around it — made non-`static` and declared in `block_words.h`, so `sk_repl_idle()` can call
the exact same flush path outside word dispatch rather than duplicating any of its logic.
Cheap to call every idle tick regardless of whether anything is actually dirty: every check
inside is a small fixed-size scan (`BLK_VM_SLOTS` here, `DISK_CACHE_SLOTS` per device inside
`blk_flush()`), so no separate "is anything dirty" pre-check was needed — the existing
function already *is* that cheap early-exit, no new logic required.
**A real bug, caught by the same live test that proved the feature.** The first version
called `sk_repl_get_active_vm()` from inside `sk_repl_idle()` and skipped the flush unless it
returned non-NULL. That accessor is Tripod's own `USE`-redirect override — NULL is its
*documented default* ("restore default dispatch (NULL = use REPL's own vm)"), not "no VM is
active." Since nothing in this test ever ran `USE`, the flush silently no-op'd for the
session's entire duration — confirmed live via a real persistence test (write a byte via
`BUFFER` with no `UPDATE`/`SAVE-BUFFERS`, wait past the idle cadence, kill QEMU abruptly,
reboot with the same disk image, read the byte back: got `0`, not the written `65`). Root
cause: `sk_repl_run()`'s own loop already resolves the correct VM every iteration
(`active = g_repl_active_vm ? g_repl_active_vm : vm`, right before calling `sk_readline()`)
`sk_repl_idle()` just had no way to see that resolution, since it's nested two calls deep
(`sk_repl_run()``sk_readline()``sk_repl_idle()`) with no VM parameter threaded through
either intermediate function. Fixed by threading the already-resolved VM down as a parameter:
`sk_readline()` gained a `VM *active_vm` argument, `sk_repl_idle()` gained one too (replacing
its own `sk_repl_get_active_vm()` call entirely), and both of `sk_readline()`'s callers now
pass the right value — `sk_repl_run()`'s own `active`, and `sk_repl_step()`'s own `vm`
parameter (a second, simpler REPL entry point with no `USE`-redirection concept at all).
Re-running the exact same persistence test after the fix: read back `65`, matching the
written `0x41` — the write survived an abrupt kill with no explicit flush call anywhere in
the test, proving the idle-tick auto-flush genuinely ran during the wait.
`logs/20260825-142510/amd64/` + `logs/20260825-142714/amd64/` are the failing before/after
reboot pair (kept as evidence of the bug, not deleted); `logs/20260825-143546/amd64/` +
`logs/20260825-143745/amd64/` are the same pair after the fix. All three architectures
re-verified clean with no dirty state pending: `logs/20260825-144008/amd64/`,
`logs/20260825-144126/aarch64/`, `logs/20260825-144517/riscv64/`.
**Not started:** no code, no design doc, no capsule work. This section exists so the next
session can pick up from an accurate baseline rather than re-deriving the shape from scratch.
@@ -3979,9 +4014,12 @@ sized around, not the size of every test image.
get its own independent implementation
- [ ] Implement the block-migration function itself (move one block's content + BAM entry
between two attached devices)
- [ ] Implement the `sk_repl_idle()` body — the cheap "anything dirty? no? done" check
- [x] Implement the `sk_repl_idle()` body — the cheap "anything dirty? no? done" check
(Section V confirmed this hook is empty and ready right now, doesn't even need
Milestone 2 to be written, only to be *tested end to end*)
Milestone 2 to be written, only to be *tested end to end*). **Done 2026-08-25**, see
Section V item 6's writeup — live-verified via a real abrupt-kill/reboot persistence
test, which also caught and fixed a real bug (`sk_repl_get_active_vm()`'s NULL default
silently no-op'ing the flush)
- [ ] Decide and implement unclean-removal handling (Section U's explicitly flagged open
question — never answered) — at minimum, detect a mid-flush disconnect via
Milestone 2e's disconnect signal and decide what state that leaves affected blocks in
+1 -1
View File
@@ -1,5 +1,5 @@
# Capsule Block Manifest — Auto-generated
<!-- Generated by mkcapsule --manifest 2026-08-25T18:05:55Z -->
<!-- Generated by mkcapsule --manifest 2026-08-25T18:44:46Z -->
<!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. -->
<!-- Hand-written justifications and immutability notes live -->
<!-- in MANIFEST.md alongside this auto-generated index. -->
Binary file not shown.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+28 -5
View File
@@ -38,6 +38,7 @@
#include "starkernel/blkio_usb.h"
#include "block_subsystem.h"
#include "word_source/include/keyboard_words.h"
#include "word_source/include/block_words.h"
#include <stdint.h>
#include <string.h>
@@ -77,7 +78,7 @@ VM *sk_repl_get_active_vm(void) { return g_repl_active_vm; }
static uint64_t g_last_beat_tick; /* zero-initialized (BSS) */
static void sk_repl_idle(void)
static void sk_repl_idle(VM *active_vm)
{
/* Artemis Milestone 2d: xHCI Event Ring servicing. This is exactly the
* "interrupt-driven, coarse cadence, cheap early-exit" trigger Section
@@ -124,6 +125,28 @@ static void sk_repl_idle(void)
xdev->bot_msc_detach_pending = 0;
blk_subsys_detach_device(&usb_blk_dev);
}
/* FABRIC.md/FABRIC-2.md Section V item 6: "a cheap 'anything dirty?
* no? done' block-sync check", the same "interrupt-driven, coarse
* cadence, cheap early-exit" trigger shape as the xHCI servicing
* above -- this was the one piece of that design already fully
* specified and waiting for this hook to actually be non-empty.
* blk_vm_flush_all() (block_words.c, the same code SAVE-BUFFERS
* itself runs) is cheap to call when nothing is dirty -- every
* check inside is a small fixed-size scan, no disk I/O happens
* unless something genuinely needs writing -- so no separate
* "is anything dirty" pre-check is needed here.
*
* active_vm is passed in by the caller (sk_readline(), itself passed
* through from sk_repl_run()/sk_repl_step()'s own already-resolved
* VM) rather than read via sk_repl_get_active_vm() here -- that
* accessor returns NULL whenever Tripod's USE word hasn't redirected
* it, which is the common case, not "no VM is active." An earlier
* version of this code called sk_repl_get_active_vm() directly and
* silently no-op'd for exactly that reason, confirmed live: a BUFFER
* write with no UPDATE, followed by an idle wait and an abrupt kill,
* did not survive a reboot until this fix. */
blk_vm_flush_all(active_vm);
}
/*===========================================================================
@@ -231,7 +254,7 @@ static int sk_kbd_getc(void)
* Returns the number of characters placed in buf (not counting '\0').
*===========================================================================*/
static int sk_readline(char *buf, int size)
static int sk_readline(char *buf, int size, VM *active_vm)
{
int n = 0;
@@ -255,7 +278,7 @@ static int sk_readline(char *buf, int size)
uint64_t now = heartbeat_ticks();
if (now - g_last_beat_tick >= SK_IDLE_BEAT_INTERVAL) {
g_last_beat_tick = now;
sk_repl_idle();
sk_repl_idle(active_vm);
}
/*
* Do NOT use hlt here: QEMU single-threaded TCG can't process
@@ -350,7 +373,7 @@ int sk_repl_step(VM *vm)
console_puts(SK_PROMPT_TEXT);
}
sk_readline(input, sizeof(input));
sk_readline(input, sizeof(input), vm);
if (input[0] == '\0') {
console_puts(" ok\n");
@@ -397,7 +420,7 @@ void sk_repl_run(VM *vm)
console_puts(SK_PROMPT_TEXT);
}
sk_readline(input, sizeof(input));
sk_readline(input, sizeof(input), active);
if (input[0] == '\0') {
console_puts(" ok\n");
+11 -2
View File
@@ -270,8 +270,17 @@ static vaddr_t blk_vm_assign(VM *vm, uint32_t lbn) {
* the epoch first (Milestone 2h, see blk_vm_check_epoch()) since this
* function walks vm->blk_vm_lbn[]/vm->blk_vm_cbuf[] directly rather than
* through blk_vm_find() -- a stale-epoch dirty slot must be discarded, not
* flushed onto whatever device now owns that LBN. */
static void blk_vm_flush_all(VM *vm) {
* flushed onto whatever device now owns that LBN.
*
* Non-static: this is the entire implementation behind SAVE-BUFFERS
* (block_word_save_buffers() below is a one-line wrapper) -- exposed
* (declared in block_words.h) so kernel-side code can reuse the exact
* same flush path outside the word-dispatch mechanism, e.g. sk_repl_idle()
* (see FABRIC.md/FABRIC-2.md Section V item 6). Cheap to call when
* nothing is dirty: every check below is a small fixed-size scan
* (BLK_VM_SLOTS here, DISK_CACHE_SLOTS per device inside blk_flush()),
* no I/O happens unless something actually needs writing. */
void blk_vm_flush_all(VM *vm) {
blk_vm_check_epoch(vm);
for (int i = 0; i < BLK_VM_SLOTS; i++) {
if (vm->blk_vm_cbuf[i] != NULL && vm->blk_vm_dirty[i]) {
+15
View File
@@ -137,6 +137,21 @@ void mark_buffer_dirty(VM * vm);
*/
void save_all_buffers(VM * vm);
/**
* @brief Sync every dirty VM block-window slot to the C-layer block cache,
* then flush the whole block subsystem to physical media --
* the complete SAVE-BUFFERS implementation (unlike
* save_all_buffers() above, which only calls blk_flush() and so
* misses any BUFFER-modified-but-not-yet-UPDATE'd content still
* sitting only in vm->memory's window). Cheap to call when nothing
* is dirty -- every check inside is a small fixed-size scan, no
* I/O happens unless something actually needs writing. Safe to
* call from outside word dispatch (e.g. a periodic idle-tick
* auto-flush); vm must be non-NULL.
* @param vm Pointer to the Forth virtual machine instance
*/
void blk_vm_flush_all(VM * vm);
/*
* @brief Marks all buffers as empty
* @param vm Pointer to the Forth virtual machine instance