Files
LithosAnanake/include/starkernel/capsule_wirebind.h
T
Robert Allan JamesandClaude Sonnet 5 2c1b3cd695
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run
Four bugs found live verifying the 8 identity thumbdrives (FABRIC-3.md §IX)
All found by actually running the identity workflow §VII/§VIII made
possible, not by code review:

1. Zuse/WIREBIND cross-contamination on detach: capsule_zuse_boot_logout()
   and capsule_wirebind_unclean_detach() both had no device parameter, so
   an unrelated device detaching (while the real owner's own stayed
   attached) incorrectly tore down the wrong session. Both now compare
   the departing device against their own tracked one, mirroring
   capsule_wirebind.c's pre-existing g_wirebind_attached_dev precedent.

2. Dictionary-entry memory leak: vm_create_word()'s sf_malloc()'d
   DictEntry (plus a second per-entry allocation for transition_metrics)
   was never freed by vm_cleanup(), in both the hosted and kernel
   implementations. Caused a real kernel PANIC after 8-9 repeated VM
   birth/kill cycles in one boot. Fixed by walking vm->latest in both.

3. sf_malloc/sf_free (alloc_kernel.c) was a 4MB bump arena with a
   deliberate no-op free, sized on "VM born once, never killed" -- fix #2
   alone didn't stop the panic because free() itself discarded the
   pointer regardless. Given a real free list (first-fit reuse).

4. Headless-console gate didn't re-engage after a mid-boot logout: the
   original fix (sk_console_mark_login(), one-way sticky) only gated the
   first login of the boot. Replaced with a live check
   (sk_console_identity_present()) re-evaluated continuously, including
   inside sk_console_readline()'s own blocking idle loop -- the console
   is normally sitting blocked there when a hot-unplug logout happens, so
   checking only at the top of the REPL loop wasn't enough.

Also: MINT now verifies its own write (verify_mint(), capsule_mint.c) by
reading back through the same check a real attach performs, rather than
trusting blkio_write()'s BLK_OK alone -- logged via log_message(), not
console_println(), per direct instruction.

Verified live, amd64: the full 8-identity repeated attach/detach cycle
that previously panicked at the same point every time now completes
clean, and a full serial-log sweep found zero bare unauthenticated
prompts anywhere in the run. Three-arch clean-qemu acceptance passed.

Still open, not fixed here: a 3+-simultaneous-device USB enumeration
failure found in a separate live test, not yet root-caused.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-06 01:49:13 -04:00

159 lines
7.2 KiB
C
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
/*
StarForth — Steady-State Virtual Machine Runtime
Copyright (c) 20232025 Robert A. James
All rights reserved.
Licensed under the StarForth License, Version 1.0
*/
/**
* capsule_wirebind.h - WIREBIND: the real thumbdrive-attach call site
* (FABRIC-2.md §F.5/§F.23). Assembles pieces already built and
* individually verified this session -- CERTVERIFY (vm_identity.h's
* vm_identity_from_cert()), RUNCAP (capsule_runcap.h), the console-VM +
* user-VM pair (capsule_console.h, sk_repl_dispatch_line() in repl.c) --
* into one automatic sequence, replacing the RUNCAP-TEST/PAIR-TEST
* diagnostic words that exercised each piece by hand.
*/
#ifndef STARKERNEL_CAPSULE_WIREBIND_H
#define STARKERNEL_CAPSULE_WIREBIND_H
#ifdef __STARKERNEL__
#include "starkernel/homeblocks_sig.h"
#include "starkernel/vm_identity.h"
#include "vm.h"
struct blkio_dev;
/**
* capsule_wirebind_verify_cert - Read the cert region off dev and verify
* it against mama_vm's own Zuse identity. Shared by both
* capsule_wirebind_try_attach() (the original attach) and BINDSTEP
* (mama_word_use(), mama_forth_words.c -- re-verifies live on every USE
* of an identity-locked VM, per FABRIC-2.md §F.9 decision 1) so both
* call sites check the exact same thing the exact same way.
*
* No-op-and-fail (-1) if sig->cert_offset is 0 (no cert region -- a
* genesis-mode Zuse drive, or simply not a regular identity drive) or
* mama_vm has no installed Zuse cert yet.
*
* @param dev Already-open block device to read the cert from.
* @param sig Its already-checked homeblocks_sig_t.
* @param mama_vm Hera's own VM -- the trust root (zuse_cert_pubkey).
* @param out Filled with the verified identity on success.
* @return 0 on success, -1 on any failure (read, verify, or precondition).
*/
int capsule_wirebind_verify_cert(struct blkio_dev *dev,
const homeblocks_sig_t *sig,
VM *mama_vm, VMIdentity *out);
/**
* capsule_wirebind_try_attach - Try to verify and bind a just-attached
* regular (non-Zuse) identity drive.
*
* No-op if sig->cert_offset is 0 (a genesis-mode Zuse drive has no cert
* region -- that's capsule_zuse_boot_try_attach()'s own job, not this
* one's) or if mama_vm has no installed Zuse cert yet (nothing to verify
* the attached cert against). Otherwise: reads the cert devblock(s),
* calls vm_identity_from_cert() against mama_vm's own zuse_cert_pubkey
* and sig->drive_uuid. On success, reads the drive's own
* user_identity_seed_t for its username and births a console VM +
* RUNCAP-born user VM pair (idempotent -- no-ops if that username is
* already live this session), installs the verified VMIdentity onto the
* user VM, and registers the "<username>~user" pairing
* (sk_repl_dispatch_line(), repl.c, looks for this). Does NOT USE the
* new console automatically -- that stays an explicit, later,
* ACL-gated step (BINDSTEP, §F.9), not something a bare attach should
* trigger silently.
*
* @param dev The just-attached, already-open block device.
* @param sig Its already-checked homeblocks_sig_t.
* @param mama_vm Hera's own VM (the verifier -- her zuse_cert_pubkey is
* the trust root regular user certs are checked against).
*/
void capsule_wirebind_try_attach(struct blkio_dev *dev,
const homeblocks_sig_t *sig,
VM *mama_vm);
/**
* capsule_wirebind_eject - Graceful detach of whatever VM is currently
* attached via the home-blocks USB path (FABRIC-2.md §F.10, decision 1).
* The drive is still physically present when this runs.
*
* Sequence: resolve the tracked attached-VM id to a live registry entry
* (no-op, returns -1, if nothing is tracked or the entry is already
* dead/gone -- capsule_vm_kill()'s own idempotency covers a VM already
* killed by some other path); blk_vm_flush_all() while the VM is still
* alive; if the console's active VM is this same VM, reset it to Hera
* (sk_repl_set_active_vm(NULL)) *before* teardown -- required, not
* optional, to avoid a dangling console pointer; capsule_vm_kill() by
* name; clear the tracked state.
*
* Single-USB-device constraint (§F.8) means there is never more than one
* candidate, so this always targets "whatever's currently attached" --
* no name argument.
*
* @return 0 on success, -1 if nothing was attached to eject.
*/
int capsule_wirebind_eject(void);
/**
* capsule_wirebind_unclean_detach - Abrupt-path counterpart to
* capsule_wirebind_eject() (FABRIC-2.md §F.10, decision 2 -- the UNCLEAN
* node, closed alongside EJECT). Called from the existing
* bot_msc_detach_pending hot-unplug signal (repl.c) -- the device is
* already gone by the time this runs, so no flush is attempted; data
* since the last flush is lost, which is correct unclean-removal
* semantics. Otherwise identical to capsule_wirebind_eject(): same
* active-VM reset-before-kill step, same tracked-state clear.
*
* FABRIC-3.md §VII follow-on, 2026-09-06: now requires the departing
* device to actually be the one tracked as this WIREBIND user's own
* (g_wirebind_attached_dev) -- a real bug otherwise, found live once
* genuine multi-device attach made a *different* device's detach
* reachable while a WIREBIND user's own stayed attached.
*
* @param dev The device that just detached; every other value is a no-op.
*/
void capsule_wirebind_unclean_detach(struct blkio_dev *dev);
/**
* capsule_wirebind_attached_username - The plain username (no "~user"
* registry-name suffix) of whichever identity is currently tracked as
* attached, or NULL if none is (FABRIC-2.md §I.1/4.4s -- the `(user)`
* console prompt segment reads this). Points into WIREBIND's own
* internal storage; valid only until the next attach/eject/detach, same
* caveat as console_get_vm_name().
*/
const char *capsule_wirebind_attached_username(void);
/**
* capsule_wirebind_overflow_idle_check - FABRIC-2.md §I.2's own "overflow
* trigger," decided and built 2026-09-05. Called once per idle tick
* (sk_repl_idle(), repl.c, alongside blk_migration_idle_check() -- same
* ~1 Hz cadence), same as that function's own convention.
*
* No-op if nothing is attached via WIREBIND. Otherwise reads the attached
* drive's own free/total via blk_get_device_free_blocks() (real numbers:
* a WIREBIND-attached drive is always HOMEBLOCKS_SIG_OK, i.e. already
* STFR/v2-formatted, by the time blk_subsys_attach_device() runs on the
* same dev pointer right after WIREBIND itself -- not PROVISIONAL, not
* raw). If free space is below a fixed threshold AND the attached
* identity does not already own a claim (blk_owner_has_claim() -- a disk
* scan, not a RAM flag, so this decision survives reboot/reattach for
* free), claims a fixed number of additional devblocks on Artemis's own
* device via blk_firsttouch_claim() -- a one-time-per-identity extension,
* not a growth loop, deliberately: this does not free space on the
* user's own drive, it only extends their pool onto system-resident
* space, so re-claiming every tick once already extended would walk
* Artemis's device to exhaustion for no benefit.
*/
void capsule_wirebind_overflow_idle_check(void);
#endif /* __STARKERNEL__ */
#endif /* STARKERNEL_CAPSULE_WIREBIND_H */