89d814f9f3cdd0a95162abeeda105463bca72e9d
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5689c397fc |
Bug-fix sweep: repl reentrancy, virtio/blocksys bounds, identity CRCs, LOG_LINE_MAX
Code review fixes, all compile clean (hosted gcc + aarch64/riscv64 kernel flags):
- repl.c (H1): reentrancy guards on the MSG-TICK idle pump. sk_repl_idle()
now defers when Hera is mid-interpret (g_mama_interpreting) or when its
own vm_interpret is on the stack (g_idle_pump_active), so a blocking
KEY/EXPECT/QUERY inside a dispatched line can no longer re-enter the
interpreter and clobber the in-flight input buffer.
- virtio_rng.c: clamp device-returned used_len to VRNG_BUF_SIZE before the
caller's data_buf copy, closing a device-controlled OOB read.
- block_subsystem.c: first-write path now keys off created_time==0 instead
of dead magic==0 so fresh blocks get a real created_time stamp; first_free/
last_allocated fixed to absolute Forth LBNs (set in blk_compute_fresh_geometry
from slot->start_lbn, no longer the wrong physical-BAM-index values from
compute_totals_from_B); physical-bounds guard on blk_meta_zone_read/write
prevents unsigned underflow on a corrupt fence >= device size.
- capsule_zuse_boot.c / capsule_wirebind.c: identity seed validated magic ->
version -> CRC-64 (compute_crc64 over offsetof(crc)) before trusting it,
so a corrupt/format-mismatched record is refused, never loaded.
- log.h / starkernel/log.h: unused LOG_LINE_MAX 256 renamed LOG_MSG_LINE_MAX
to lift the include-order collision with vm.h's LOG_LINE_MAX 64; stale
include-order comments dropped (kernel_main.c, shim.c, capsule_birth.c).
- FABRIC-3.md: three stale-doc carry-forward items closed [x] with
|
||
|
|
6e9c1d3bc2 |
Phase 8 C (4/n): metadata fence allocator integration + zone I/O
Corrected meta_fence_blocks units from "Forth 1 KiB blocks" to 4 KiB devblocks (matching bam_devblocks/reloc_devblocks) before anything depended on the original meaning -- a clean fix, not a migration. This let the fence fold directly into compute_totals_from_B()'s existing payload4k formula (total_devblocks - 1 - B - R - F) instead of a separate user_blocks subtraction: total_blocks/user_blocks/free_blocks all shrink correctly for free, in both the fresh-format and reload paths, from one formula change. New blk_meta_zone_read()/blk_meta_zone_write() -- raw, unpacked 4 KiB devblock I/O, same shape as the header/BAM/reloc-table regions, addressed by devblock_from_top counting down from the device's last physical devblock. Refuses rather than clamps if the index exceeds the on-disk meta_fence_blocks. C-only, no FORTH word wraps either -- same discipline as vm_zuse_cert_install(), which will be this zone's first real tenant. Verified independently at every step, never trusting the kernel's own report: capacity math cross-checked against a from-scratch Python recomputation of the same formula (exact match); accessor correctness via a temporary probe (written/run/captured/reverted) that wrote a known pattern and read it back, then independently confirmed via a raw read of the disk image at the exact expected physical byte offset. Full 3-arch acceptance boot against the real, untouched disk/artemis.img, probe code fully reverted -- clean, conservation intact. Still open: wiring vm_zuse_cert_install() to actually persist through these accessors, and the MINT word itself. Documented in FABRIC-3.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U14ET9CWAtbQMbYqomKgXd |
||
|
|
2035ebeac0 |
Phase 8 C (3/n): top-of-device metadata fence, step 1 (field round-trip)
Corrected substrate: this OS is anti-POSIX, anti-file by design -- the prior "dedicated system-identity disk" framing was wrong vocabulary, caught before any code was written (saved as feedback_no_files_anti_posix.md). The real primitives are content-addressed capsules and raw LBN blocks, never a filesystem. Design (agreed on request): a growable metadata fence at the TOP of a device's block space, mirroring block_subsystem.c's existing bottom BAM reservation from the opposite end -- the two grow toward each other, never colliding, same shape as a stack/heap. Starts at BLK_META_FENCE_INIT (128 blocks), explicitly never RAM-backed. Reuses Artemis's already-attached, already-proven virtio-blk device -- no new device. Rejected reusing BAM's own reserved zone directly: those blocks are fully claimed by BAM bookkeeping, not free space. Step 1 only: new meta_fence_blocks field in blk_volume_meta_t, appended after reloc_devblocks and carved from _pad[] -- identical graceful- default technique reloc_devblocks already established (a pre-existing volume reads it back as 0, not a format break). Added a compile-time _Static_assert on the struct's total size, same discipline homeblocks_sig.h uses -- caught a real bug immediately: the hand-summed _pad[] formula was off by 4 bytes (a compiler alignment gap the manual count missed), found via offsetof() rather than re-deriving by hand. Worked against disposable clones throughout, never the real disk/artemis.img (ARTDISK is ?=-overridable) -- artemis-metafence-fresh.img (blank, fresh-format path) and artemis-metafence-test.img (copy of the pre-existing artemis.img, graceful-default-on-reload path), kept as regression fixtures matching disk/README.md's existing convention. Verified independently via direct byte reads of the disk image, not the kernel's own log output (log_message(LOG_INFO,...) doesn't reach serial in this build -- unrelated pre-existing gap): fresh format writes 128 at header offset 184, a reboot without reformatting preserves it, the old pre-fence image reads back 0. Full 3-arch acceptance boot against the real, untouched disk/artemis.img also clean. Allocator (user_blocks math) and zone read/write accessors both still open -- next steps, documented in FABRIC-3.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U14ET9CWAtbQMbYqomKgXd |
||
|
|
2c45744995 |
Implement homeblocks_sig_check(): the drive signature check (Phase 8)
Real, complete verification logic -- not yet wired to any write path. homeblocks_sig_check(dev, sig_start_fblock, out_sig) reads the 4 consecutive 1KB blkio forth-blocks the 4KB header spans, verifies magic -> version -> CRC-64 in order, returns HOMEBLOCKS_SIG_OK/_BLANK/ _BAD_VERSION/_BAD_CRC/_READ_ERROR. Reuses block_subsystem.c's existing CRC-64/ISO (compute_crc64, previously static/file-local, now exposed via block_subsystem.h) rather than a second CRC implementation -- same algorithm already proven via per-block checksums. Takes the header's starting block as a plain parameter rather than resolving it internally: verifies a signature given a location, finding that location (GPT-partition-relative) stays the caller's job. Verified against the actual shipped code, not a reimplementation: a standalone host test links the real homeblocks_sig.c against a fake in-memory blkio_dev and exercises all four outcomes -- blank media, a correctly-minted header (round-trips drive_uuid/minted_time_ns), a flipped CRC, an unrecognized version. All four pass. A full QEMU-hotplug live test isn't proportionate yet since nothing calls this function from the live kernel path -- wiring it into the attach path is the next punch-list item. Clean zero-warning compile and clean boot on all three architectures confirms no build/link regression from exposing compute_crc64 and adding the new source file to every kernel build. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn |
||
|
|
36d832ff47 |
Single-block relocation: RELOCATE-BLOCK, resolve_lbn(), persisted exception table
Implements the full design from the prior commit in one pass. resolve_lbn()
is the single choke point threaded through the ten public LBN-consuming
entry points (blk_get_buffer, blk_update, blk_flush, blk_is_allocated,
blk_mark_allocated, blk_mark_free, blk_is_valid, blk_get_meta, blk_set_meta,
plus blk_get_empty_buffer covered via delegation) -- an LBN->LBN redirect,
not a new storage allocator, since the LBN space is already unified across
every attached blkio_dev backend. VM window cache staleness across a
relocation reuses the existing blk_vm_check_epoch() mechanism from
Milestone 2h's hot-detach fix for free -- g.epoch bumps on relocation too.
Persistence lands in the same pass: two new uint32_t fields
(reloc_start/reloc_devblocks) appended after hdr_crc in blk_volume_meta_t,
carved from existing padding without moving any earlier field's byte
offset -- an old formatted volume's zeroed padding reads back as
reloc_devblocks=0 ("no reloc capacity"), gracefully, not a format-breaking
change. compute_totals_from_B() generalized to account for the new
reserved region. reloc_flush_to_disk()/reloc_load_from_disk() mirror the
BAM I/O functions' own absolute-devblock-addressing shape; the persisted
copy's owner is first_disk_slot() (already existed, already used for this
exact "which device is canonical" question by blk_get_volume_meta()).
blk_subsys_relocate_block() is a mechanical primitive only -- copies
content (staged through a local buffer, since obtaining the target's
blk_get_buffer() result can evict and invalidate the source's cache
pointer if they share a device), frees the source BAM entry, appends the
exception entry, bumps the epoch, flushes to disk. RELOCATE-BLOCK exposes
it to FORTH, no policy of its own (ACL's job, per this session's direction).
A first live-test attempt gave a false negative against disk/artemis.img
(predates reloc capacity, so relocation only ever existed in memory that
boot) -- traced to the test's own setup before being mistaken for a bug,
then re-verified correctly against a fresh volume (new fixture,
disk/artemis-reloc-test.img): relocated a RAMDRIVE block to the fresh
disk, confirmed live resolution through the redirect, then confirmed both
the redirect and the relocated content survived an abrupt QEMU kill and
full reboot. Also fixed three lingering "glibc" doc-comment
misattributions from Milestone 2h (the actual allocator is this kernel's
own kmalloc) that survived an earlier FABRIC-2.md-only correction. All
three architectures re-verified clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
|
||
|
|
af267a52a6 |
Artemis Milestone 2h: hot-detach -- 2h complete
blk_subsys_detach_device() (block_subsystem.c) walks the device chain, refuses removal of anything but the current tail (a mid-chain removal would corrupt every later slot's start_lbn -- this architecture's own doc already argues USB stays last specifically to avoid that), unlinks, shrinks total_user_lbn, closes and frees the slot. Discards rather than flushes dirty state -- the device is physically gone by the time this runs (PORTSC disconnect only). Trigger wiring mirrors the attach path: bot_msc_attached (set only once attach actually succeeds) gates a new bot_msc_detach_pending flag set at PORTSC disconnect (not Disable Slot completion, which is conditionally skipped and would miss concurrent connect/disconnect pairs), consumed in sk_repl_idle(). Advisor flagged the real hazard ahead of time: block_words.c's VM block window (blk_vm_lbn[]/blk_vm_cbuf[]) can go stale across a detach then a same-LBN re-attach, and suggested a pointer-identity re-check in blk_vm_load() as a minimal fix. That fix was implemented, then directly falsified by its own designed-for-this test: attach a blank device, read a block (populating the cache), detach, re-attach a device with distinct content at the identical LBN, read again -- served stale content from the first device. Root cause, confirmed live: glibc's allocator hands free(slot) straight back to the very next same-size calloc(), so the "fresh" and stale pointers were bitwise identical despite being two different devices. Fixed properly with a monotonic blk_subsys_epoch() counter (bumped on every attach/detach) checked by a new blk_vm_check_epoch() helper at the one choke point (blk_vm_find(), plus blk_vm_flush_all() which reads the same arrays directly) that covers every path touching the window cache -- unfooled by address reuse. Verified live with a new disk/usb-thumbdrive-test2.img fixture (distinct content from the existing blank test image): attach A, read (cache hit populated), detach, re-attach B at the same LBN, read again -- correctly ran a fresh device read and returned B's real content, not A's stale cached zeros. The failing pointer-comparison attempt's own capture log kept as evidence, not deleted. All three architectures re-verified clean. FABRIC-2.md Section X 2h marked complete -- enumeration through hot-detach all live and verified; only WRITE(10) (2g's own still-open item) remains unimplemented in the driver, not blocking anything here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn |
||
|
|
148c4aa12c |
Fix silent disk overwrite of unrecognized Artemis disks
The generic block subsystem (blk_format_or_load_disk) auto-reformatted any disk lacking its own low-level 'STFR' header at attach time, before Artemis's Forth-level BLANK/LithosAnanke/Unrecognized classification ever ran -- so ART-HALT-UNRECOG's "Disk preserved" message was false. Split detection from commit: an unrecognized/blank disk is now left PROVISIONAL (geometry computed in memory only, all writes refused) until explicitly confirmed via the new blk_subsys_confirm_format() / BLK-CONFIRM-FORMAT primitive. Artemis calls it from ART-FORMAT and ART-RESUME, never from ART-HALT-UNRECOG. Verified on amd64/aarch64/riscv64: parity intact (identical dict_hash), normal recognized-disk resume + persist-read unaffected, and a regenerated disk/artemis-unrecognized-test.img (the old copy had itself been silently corrupted by this exact bug) now stays byte-for-byte identical across a halted boot on amd64 and riscv64. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a5ed8c3d87 | Initial commit — LithosAnanke kernel |