aarch64 BYE crash: add heap-address diagnostics, rule out gdb debugging on this target

Continued investigating the aarch64 BYE cold-restart exception (FABRIC-2.md
Section I). Added permanent boot diagnostics: kmalloc_heap_base_addr()/
kmalloc_heap_end_addr() now print in print_heap_stats(), confirming the
fault address is provably inside the kmalloc heap (not kernel code, not
firmware). Bumped aarch64 QEMU RAM to 4096MB to test heap-placement
sensitivity (no effect -- heap size is a fixed 2GiB default, independent
of total RAM once "enough" exists).

Three separate live gdb debugging attempts (software breakpoint, hardware
breakpoint on arch_cold_reset, hardware breakpoint on mama_word_bye's
entry) all silently failed to fire despite disassembly-confirmed-correct
addresses and confirmed execution reaching those points. A sanity check
(hbreak on console_println, called thousands of times per boot) also never
fired even 8802 lines into a serial log -- conclusively a gdbstub/QEMU
tooling limitation for this aarch64 target, not a kernel-side finding.
Live single-stepping is not currently viable here; documented so it isn't
re-attempted the same way.

Root cause still open. Full trail in FABRIC-2.md Section I.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Robert Allan James
2026-08-18 17:06:58 -04:00
co-authored by Claude Sonnet 5
parent 8d90538801
commit bc31916461
11 changed files with 308066 additions and 32 deletions
+77 -30
View File
@@ -1096,34 +1096,81 @@ now-confirmed-ineffective interrupt mask was removed from `mama_word_bye()`, res
its pre-hypothesis form) rather than carry forward a change that both doesn't work and its pre-hypothesis form) rather than carry forward a change that both doesn't work and
correctly bothered Captain Bob on design grounds. correctly bothered Captain Bob on design grounds.
**Correction: the "nowhere near `arch_cold_reset`" argument above does not actually hold.** **False alarm, then corrected back: the kernel does NOT relocate to a UEFI-chosen address.**
That 4-byte shift, tracking a code-size change exactly, is what exposed the flaw: it means The 4-byte shift briefly looked like it disproved "nowhere near `arch_cold_reset`," and led
the fault address is coupled to kernel binary layout, which sent us back to check how the to a (wrong) detour: `AllocatePages(AllocateAnyPages, ...)` in `uefi_loader.c` was mistaken
kernel is actually loaded. `src/starkernel/boot/uefi_loader.c` calls UEFI's for the address the *executable code* loads at. Checked directly against
`AllocatePages(AllocateAnyPages, ...)` — the kernel is loaded at an address UEFI's own page `src/starkernel/boot/elf_loader.c`'s `elf_load_kernel()`: that `AllocateAnyPages` call only
allocator chooses at boot, not a fixed base — and `elf_loader.c`'s `elf_apply_relocations()` allocates a scratch buffer to hold the raw ELF *file bytes* before parsing. The actual
then applies real PIE-style relocations against that chosen base. `nm`'s addresses (like segment-load address (`load_base`) is either `0` for `ET_EXEC` or a fixed `0x400000` for
`arch_cold_reset`'s `0x40ef40`) are link-time addresses, assuming the ELF's own default base; `ET_DYN` — never UEFI-chosen. `readelf -h` on the built kernel confirms `Type: EXEC`, so
they say nothing about where the code actually lands at runtime, which is wherever `load_base=0`: the kernel really does run at exactly the addresses its own link-time ELF
`AllocateAnyPages` happened to place it — plausibly right in the `0xbe0xxxxx` neighborhood symbol table (and `nm`) report. The original "`0xbe03e81c` is nowhere near
this crash keeps landing in. Comparing the two earlier builds (with/without the PSCI fix) `arch_cold_reset`'s real code" conclusion was correct after all; the mid-session "correction"
and seeing the address stay bit-for-bit identical proved nothing either, in hindsight — that above was itself the mistake, now itself corrected. (The 4-byte shift is still real and still
fix only changed an immediate value, not instruction count, so nothing in the image's size needs an explanation — see the heap finding below, which supplies one.)
or layout changed between those two builds regardless of where the real fault was. **Net
effect: the original "this must be garbage heap memory, not real code" conclusion is
unproven, not confirmed.** It remains plausible, but so does "this is legitimate relocated
kernel code that a proper address-translation would identify," and nothing done in this pass
distinguishes the two. No load-base address is printed anywhere in the current boot log, so
there was no data available to settle it further.
**Where this stands:** root cause not found. Two competing signatures observed across **Decisive finding: the fault address is provably inside the kmalloc heap.** Added console
history (`ESR_EL1=0x02000000`/EC=0 "Unknown reason" here and in the 2026-08-08 keyboard prints of `kmalloc_heap_base_addr()`/`kmalloc_heap_end_addr()` (both already existed as
test; `ESR_EL1=0x9600004f`/EC=0x25 genuine data-abort with a poisoned-looking accessors, just never logged) right after `kmalloc_init()`, and bumped the aarch64 QEMU
`FAR_EL1=0x000055bd` in the June/July 2026 crashes) may or may not be the same underlying `-m` from 2048 to 4096 (`Makefile.starkernel`) to see whether more physical RAM shifts
bug. The interrupt-race hypothesis is refuted. The "wild jump into garbage RAM" framing is anything. Rebuilt and booted: `Heap base addr: 1207959552` (`0x48000000`), `Heap end addr:
unproven, undermined by not accounting for `AllocateAnyPages`-based relocation. Stopped here 3355443200` (`0xc8000000`). Both recorded fault addresses, `0xbe03e81c` and `0xbe03e820`,
deliberately, per Captain Bob's direction, rather than continuing to dig live. Next session fall squarely inside that range (`0x48000000 ≤ 0xbe03e81c ≤ 0xc8000000`, about 1.97 GiB into
should start by printing the actual UEFI-chosen load base and kernel entry point at boot the 2 GiB heap, ~167 MB short of its end). **This settles the "garbage RAM vs. real code"
(nothing currently does), so `ELR_EL1` values can be translated back to real source question the earlier back-and-forth couldn't: it is heap-internal, not kernel code, not
locations instead of link-time guesses — that is the missing piece every path above kept firmware.** It also explains the 4-byte shift cleanly: the fault isn't landing at a fixed
running into. Genuinely open, unlike everything else this document tracks as closed. absolute address, it's the result of whatever computation goes wrong reading a value that
depends on the heap's or kernel's own layout — a computation that necessarily moves by a few
bytes when the kernel binary's size changes by a few bytes. The bump to 4096 MB RAM did not
change heap placement (heap size is a fixed 2 GiB default via `KARGS_DEFAULT_HEAP_SIZE`,
independent of total RAM once "enough" exists), so that specific change didn't add further
data, but confirms heap placement isn't RAM-size-sensitive at this configuration.
**Live gdb debugging attempted three times against the full acceptance-harness device set —
conclusively ruled out as a viable path in this environment, for a reason unrelated to the
kernel itself.** Bumped aarch64 QEMU to `-m 4096` (`Makefile.starkernel`) and added
`kmalloc_heap_base_addr()`/`kmalloc_heap_end_addr()` console prints (`kernel_main.c`,
`print_heap_stats()`) to get the heap-bracketing evidence above. Then, with the same full
device set (`artdisk`, keyboard, `ramfb`) plus `-s -S` and `gdb-multiarch` attached via a
long-lived FIFO-fed session (so gdb could sit waiting through the full ~30-minute real-time
boot without a `-batch` timeout cutting it off):
1. Software breakpoint (`break arch_cold_reset`) at its correctly-identified address
(`0x40ef60`, confirmed against `nm` for that exact build) — did not fire. Crash occurred
normally, gdb reported `[Inferior 1 (process 1) exited normally]`.
2. Hardware breakpoint (`hbreak arch_cold_reset`) at the same address, confirmed "Hardware
assisted breakpoint 1 at 0x40ef60" by gdb — also did not fire. Same outcome.
3. Hardware breakpoint at `mama_word_bye`'s own entry (`0x40bf80`) — unconditionally reached
(confirmed by "BYE: reaping children" printing, the function's first statement) — also
did not fire.
Disassembly (`aarch64-linux-gnu-objdump`) confirmed the call site itself is correct and
unremarkable: `mama_word_bye` ends with `bl 40ef60 <arch_cold_reset>`, a plain direct branch
to the right address — ruling out a bad relocation or corrupted call instruction as the
reason the breakpoints didn't fire.
**Sanity check, decisive:** set `hbreak console_println` — a function called thousands of
times from the very first moment of kernel boot — on a fresh boot. After the serial log had
already accumulated **8,802 lines** (deep into Artemis's stress campaign, `console_println`
necessarily called thousands of times to produce that output), gdb still reported only
`Continuing.` — zero breakpoint hits, ever. **This means gdb breakpoints (software and
hardware alike) do not function at all against this QEMU aarch64 (`cortex-a57`, TCG) target
in this environment** — not an icache-coherency property of our kernel, a limitation of the
debugging setup itself. The icache-corruption hypothesis this session was pursuing is
therefore neither confirmed nor refuted; it's simply unreachable with this tooling as
configured. All gdb sessions and their QEMU instances were killed; no code changes came out
of this thread beyond the (kept) heap-address prints and the `-m 4096` bump, both harmless
diagnostics worth keeping regardless.
**Where this stands:** root cause narrowed but not found. Confirmed heap-internal (not
kernel code, not firmware, not truly random/unmapped memory — see the decisive finding
above). Live single-stepping is not currently viable against this target; before attempting
it again, the QEMU aarch64 gdbstub setup itself needs to be validated independently (try a
different QEMU version, `-accel tcg,thread=single`, or confirm hardware breakpoint support
against a trivial known-working aarch64 QEMU target first, outside this kernel entirely) —
don't repeat this exact approach expecting a different result. Two competing crash
signatures still unreconciled: `ESR_EL1=0x02000000`/EC=0 "Unknown reason" (this crash, and
the 2026-08-08 keyboard-input crash) vs. `ESR_EL1=0x9600004f`/EC=0x25 genuine data-abort with
a poisoned-looking `FAR_EL1=0x000055bd` (June/July 2026 crashes) — may or may not be the same
bug. Genuinely open.
+1 -1
View File
@@ -815,7 +815,7 @@ else ifeq ($(ARCH),aarch64)
-chardev socket,id=cserial,path=$$SERIAL_SOCK,server=on,wait=off,logfile=$$LOG \ -chardev socket,id=cserial,path=$$SERIAL_SOCK,server=on,wait=off,logfile=$$LOG \
-serial chardev:cserial \ -serial chardev:cserial \
-display $(QEMU_DISPLAY) \ -display $(QEMU_DISPLAY) \
-m 2048 \ -m 4096 \
-no-reboot \ -no-reboot \
-d guest_errors; \ -d guest_errors; \
kill $$TAILPID 2>/dev/null; wait $$TAILPID 2>/dev/null || true; \ kill $$TAILPID 2>/dev/null; wait $$TAILPID 2>/dev/null || true; \
+1 -1
View File
@@ -1,5 +1,5 @@
# Capsule Block Manifest — Auto-generated # Capsule Block Manifest — Auto-generated
<!-- Generated by mkcapsule --manifest 2026-08-18T18:07:32Z --> <!-- Generated by mkcapsule --manifest 2026-08-18T19:02:24Z -->
<!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. --> <!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. -->
<!-- Hand-written justifications and immutability notes live --> <!-- Hand-written justifications and immutability notes live -->
<!-- in MANIFEST.md alongside this auto-generated index. --> <!-- in MANIFEST.md alongside this auto-generated index. -->
BIN
View File
Binary file not shown.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+2
View File
@@ -306,6 +306,8 @@ static void print_heap_stats(void) {
print_uint(" Free bytes: ", stats.free_bytes); print_uint(" Free bytes: ", stats.free_bytes);
print_uint(" Used bytes: ", stats.used_bytes); print_uint(" Used bytes: ", stats.used_bytes);
print_uint(" Peak bytes: ", stats.peak_bytes); print_uint(" Peak bytes: ", stats.peak_bytes);
print_uint(" Heap base addr: ", (uint64_t)kmalloc_heap_base_addr());
print_uint(" Heap end addr: ", (uint64_t)kmalloc_heap_end_addr());
console_println(""); console_println("");
} }