aarch64 BYE crash: add heap-address diagnostics, rule out gdb debugging on this target
Continued investigating the aarch64 BYE cold-restart exception (FABRIC-2.md Section I). Added permanent boot diagnostics: kmalloc_heap_base_addr()/ kmalloc_heap_end_addr() now print in print_heap_stats(), confirming the fault address is provably inside the kmalloc heap (not kernel code, not firmware). Bumped aarch64 QEMU RAM to 4096MB to test heap-placement sensitivity (no effect -- heap size is a fixed 2GiB default, independent of total RAM once "enough" exists). Three separate live gdb debugging attempts (software breakpoint, hardware breakpoint on arch_cold_reset, hardware breakpoint on mama_word_bye's entry) all silently failed to fire despite disassembly-confirmed-correct addresses and confirmed execution reaching those points. A sanity check (hbreak on console_println, called thousands of times per boot) also never fired even 8802 lines into a serial log -- conclusively a gdbstub/QEMU tooling limitation for this aarch64 target, not a kernel-side finding. Live single-stepping is not currently viable here; documented so it isn't re-attempted the same way. Root cause still open. Full trail in FABRIC-2.md Section I. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
8d90538801
commit
bc31916461
+77
-30
@@ -1096,34 +1096,81 @@ now-confirmed-ineffective interrupt mask was removed from `mama_word_bye()`, res
|
|||||||
its pre-hypothesis form) rather than carry forward a change that both doesn't work and
|
its pre-hypothesis form) rather than carry forward a change that both doesn't work and
|
||||||
correctly bothered Captain Bob on design grounds.
|
correctly bothered Captain Bob on design grounds.
|
||||||
|
|
||||||
**Correction: the "nowhere near `arch_cold_reset`" argument above does not actually hold.**
|
**False alarm, then corrected back: the kernel does NOT relocate to a UEFI-chosen address.**
|
||||||
That 4-byte shift, tracking a code-size change exactly, is what exposed the flaw: it means
|
The 4-byte shift briefly looked like it disproved "nowhere near `arch_cold_reset`," and led
|
||||||
the fault address is coupled to kernel binary layout, which sent us back to check how the
|
to a (wrong) detour: `AllocatePages(AllocateAnyPages, ...)` in `uefi_loader.c` was mistaken
|
||||||
kernel is actually loaded. `src/starkernel/boot/uefi_loader.c` calls UEFI's
|
for the address the *executable code* loads at. Checked directly against
|
||||||
`AllocatePages(AllocateAnyPages, ...)` — the kernel is loaded at an address UEFI's own page
|
`src/starkernel/boot/elf_loader.c`'s `elf_load_kernel()`: that `AllocateAnyPages` call only
|
||||||
allocator chooses at boot, not a fixed base — and `elf_loader.c`'s `elf_apply_relocations()`
|
allocates a scratch buffer to hold the raw ELF *file bytes* before parsing. The actual
|
||||||
then applies real PIE-style relocations against that chosen base. `nm`'s addresses (like
|
segment-load address (`load_base`) is either `0` for `ET_EXEC` or a fixed `0x400000` for
|
||||||
`arch_cold_reset`'s `0x40ef40`) are link-time addresses, assuming the ELF's own default base;
|
`ET_DYN` — never UEFI-chosen. `readelf -h` on the built kernel confirms `Type: EXEC`, so
|
||||||
they say nothing about where the code actually lands at runtime, which is wherever
|
`load_base=0`: the kernel really does run at exactly the addresses its own link-time ELF
|
||||||
`AllocateAnyPages` happened to place it — plausibly right in the `0xbe0xxxxx` neighborhood
|
symbol table (and `nm`) report. The original "`0xbe03e81c` is nowhere near
|
||||||
this crash keeps landing in. Comparing the two earlier builds (with/without the PSCI fix)
|
`arch_cold_reset`'s real code" conclusion was correct after all; the mid-session "correction"
|
||||||
and seeing the address stay bit-for-bit identical proved nothing either, in hindsight — that
|
above was itself the mistake, now itself corrected. (The 4-byte shift is still real and still
|
||||||
fix only changed an immediate value, not instruction count, so nothing in the image's size
|
needs an explanation — see the heap finding below, which supplies one.)
|
||||||
or layout changed between those two builds regardless of where the real fault was. **Net
|
|
||||||
effect: the original "this must be garbage heap memory, not real code" conclusion is
|
|
||||||
unproven, not confirmed.** It remains plausible, but so does "this is legitimate relocated
|
|
||||||
kernel code that a proper address-translation would identify," and nothing done in this pass
|
|
||||||
distinguishes the two. No load-base address is printed anywhere in the current boot log, so
|
|
||||||
there was no data available to settle it further.
|
|
||||||
|
|
||||||
**Where this stands:** root cause not found. Two competing signatures observed across
|
**Decisive finding: the fault address is provably inside the kmalloc heap.** Added console
|
||||||
history (`ESR_EL1=0x02000000`/EC=0 "Unknown reason" here and in the 2026-08-08 keyboard
|
prints of `kmalloc_heap_base_addr()`/`kmalloc_heap_end_addr()` (both already existed as
|
||||||
test; `ESR_EL1=0x9600004f`/EC=0x25 genuine data-abort with a poisoned-looking
|
accessors, just never logged) right after `kmalloc_init()`, and bumped the aarch64 QEMU
|
||||||
`FAR_EL1=0x000055bd` in the June/July 2026 crashes) may or may not be the same underlying
|
`-m` from 2048 to 4096 (`Makefile.starkernel`) to see whether more physical RAM shifts
|
||||||
bug. The interrupt-race hypothesis is refuted. The "wild jump into garbage RAM" framing is
|
anything. Rebuilt and booted: `Heap base addr: 1207959552` (`0x48000000`), `Heap end addr:
|
||||||
unproven, undermined by not accounting for `AllocateAnyPages`-based relocation. Stopped here
|
3355443200` (`0xc8000000`). Both recorded fault addresses, `0xbe03e81c` and `0xbe03e820`,
|
||||||
deliberately, per Captain Bob's direction, rather than continuing to dig live. Next session
|
fall squarely inside that range (`0x48000000 ≤ 0xbe03e81c ≤ 0xc8000000`, about 1.97 GiB into
|
||||||
should start by printing the actual UEFI-chosen load base and kernel entry point at boot
|
the 2 GiB heap, ~167 MB short of its end). **This settles the "garbage RAM vs. real code"
|
||||||
(nothing currently does), so `ELR_EL1` values can be translated back to real source
|
question the earlier back-and-forth couldn't: it is heap-internal, not kernel code, not
|
||||||
locations instead of link-time guesses — that is the missing piece every path above kept
|
firmware.** It also explains the 4-byte shift cleanly: the fault isn't landing at a fixed
|
||||||
running into. Genuinely open, unlike everything else this document tracks as closed.
|
absolute address, it's the result of whatever computation goes wrong reading a value that
|
||||||
|
depends on the heap's or kernel's own layout — a computation that necessarily moves by a few
|
||||||
|
bytes when the kernel binary's size changes by a few bytes. The bump to 4096 MB RAM did not
|
||||||
|
change heap placement (heap size is a fixed 2 GiB default via `KARGS_DEFAULT_HEAP_SIZE`,
|
||||||
|
independent of total RAM once "enough" exists), so that specific change didn't add further
|
||||||
|
data, but confirms heap placement isn't RAM-size-sensitive at this configuration.
|
||||||
|
|
||||||
|
**Live gdb debugging attempted three times against the full acceptance-harness device set —
|
||||||
|
conclusively ruled out as a viable path in this environment, for a reason unrelated to the
|
||||||
|
kernel itself.** Bumped aarch64 QEMU to `-m 4096` (`Makefile.starkernel`) and added
|
||||||
|
`kmalloc_heap_base_addr()`/`kmalloc_heap_end_addr()` console prints (`kernel_main.c`,
|
||||||
|
`print_heap_stats()`) to get the heap-bracketing evidence above. Then, with the same full
|
||||||
|
device set (`artdisk`, keyboard, `ramfb`) plus `-s -S` and `gdb-multiarch` attached via a
|
||||||
|
long-lived FIFO-fed session (so gdb could sit waiting through the full ~30-minute real-time
|
||||||
|
boot without a `-batch` timeout cutting it off):
|
||||||
|
|
||||||
|
1. Software breakpoint (`break arch_cold_reset`) at its correctly-identified address
|
||||||
|
(`0x40ef60`, confirmed against `nm` for that exact build) — did not fire. Crash occurred
|
||||||
|
normally, gdb reported `[Inferior 1 (process 1) exited normally]`.
|
||||||
|
2. Hardware breakpoint (`hbreak arch_cold_reset`) at the same address, confirmed "Hardware
|
||||||
|
assisted breakpoint 1 at 0x40ef60" by gdb — also did not fire. Same outcome.
|
||||||
|
3. Hardware breakpoint at `mama_word_bye`'s own entry (`0x40bf80`) — unconditionally reached
|
||||||
|
(confirmed by "BYE: reaping children" printing, the function's first statement) — also
|
||||||
|
did not fire.
|
||||||
|
|
||||||
|
Disassembly (`aarch64-linux-gnu-objdump`) confirmed the call site itself is correct and
|
||||||
|
unremarkable: `mama_word_bye` ends with `bl 40ef60 <arch_cold_reset>`, a plain direct branch
|
||||||
|
to the right address — ruling out a bad relocation or corrupted call instruction as the
|
||||||
|
reason the breakpoints didn't fire.
|
||||||
|
|
||||||
|
**Sanity check, decisive:** set `hbreak console_println` — a function called thousands of
|
||||||
|
times from the very first moment of kernel boot — on a fresh boot. After the serial log had
|
||||||
|
already accumulated **8,802 lines** (deep into Artemis's stress campaign, `console_println`
|
||||||
|
necessarily called thousands of times to produce that output), gdb still reported only
|
||||||
|
`Continuing.` — zero breakpoint hits, ever. **This means gdb breakpoints (software and
|
||||||
|
hardware alike) do not function at all against this QEMU aarch64 (`cortex-a57`, TCG) target
|
||||||
|
in this environment** — not an icache-coherency property of our kernel, a limitation of the
|
||||||
|
debugging setup itself. The icache-corruption hypothesis this session was pursuing is
|
||||||
|
therefore neither confirmed nor refuted; it's simply unreachable with this tooling as
|
||||||
|
configured. All gdb sessions and their QEMU instances were killed; no code changes came out
|
||||||
|
of this thread beyond the (kept) heap-address prints and the `-m 4096` bump, both harmless
|
||||||
|
diagnostics worth keeping regardless.
|
||||||
|
|
||||||
|
**Where this stands:** root cause narrowed but not found. Confirmed heap-internal (not
|
||||||
|
kernel code, not firmware, not truly random/unmapped memory — see the decisive finding
|
||||||
|
above). Live single-stepping is not currently viable against this target; before attempting
|
||||||
|
it again, the QEMU aarch64 gdbstub setup itself needs to be validated independently (try a
|
||||||
|
different QEMU version, `-accel tcg,thread=single`, or confirm hardware breakpoint support
|
||||||
|
against a trivial known-working aarch64 QEMU target first, outside this kernel entirely) —
|
||||||
|
don't repeat this exact approach expecting a different result. Two competing crash
|
||||||
|
signatures still unreconciled: `ESR_EL1=0x02000000`/EC=0 "Unknown reason" (this crash, and
|
||||||
|
the 2026-08-08 keyboard-input crash) vs. `ESR_EL1=0x9600004f`/EC=0x25 genuine data-abort with
|
||||||
|
a poisoned-looking `FAR_EL1=0x000055bd` (June/July 2026 crashes) — may or may not be the same
|
||||||
|
bug. Genuinely open.
|
||||||
|
|||||||
+1
-1
@@ -815,7 +815,7 @@ else ifeq ($(ARCH),aarch64)
|
|||||||
-chardev socket,id=cserial,path=$$SERIAL_SOCK,server=on,wait=off,logfile=$$LOG \
|
-chardev socket,id=cserial,path=$$SERIAL_SOCK,server=on,wait=off,logfile=$$LOG \
|
||||||
-serial chardev:cserial \
|
-serial chardev:cserial \
|
||||||
-display $(QEMU_DISPLAY) \
|
-display $(QEMU_DISPLAY) \
|
||||||
-m 2048 \
|
-m 4096 \
|
||||||
-no-reboot \
|
-no-reboot \
|
||||||
-d guest_errors; \
|
-d guest_errors; \
|
||||||
kill $$TAILPID 2>/dev/null; wait $$TAILPID 2>/dev/null || true; \
|
kill $$TAILPID 2>/dev/null; wait $$TAILPID 2>/dev/null || true; \
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
# Capsule Block Manifest — Auto-generated
|
# Capsule Block Manifest — Auto-generated
|
||||||
<!-- Generated by mkcapsule --manifest 2026-08-18T18:07:32Z -->
|
<!-- Generated by mkcapsule --manifest 2026-08-18T19:02:24Z -->
|
||||||
<!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. -->
|
<!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. -->
|
||||||
<!-- Hand-written justifications and immutability notes live -->
|
<!-- Hand-written justifications and immutability notes live -->
|
||||||
<!-- in MANIFEST.md alongside this auto-generated index. -->
|
<!-- in MANIFEST.md alongside this auto-generated index. -->
|
||||||
|
|||||||
Binary file not shown.
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -306,6 +306,8 @@ static void print_heap_stats(void) {
|
|||||||
print_uint(" Free bytes: ", stats.free_bytes);
|
print_uint(" Free bytes: ", stats.free_bytes);
|
||||||
print_uint(" Used bytes: ", stats.used_bytes);
|
print_uint(" Used bytes: ", stats.used_bytes);
|
||||||
print_uint(" Peak bytes: ", stats.peak_bytes);
|
print_uint(" Peak bytes: ", stats.peak_bytes);
|
||||||
|
print_uint(" Heap base addr: ", (uint64_t)kmalloc_heap_base_addr());
|
||||||
|
print_uint(" Heap end addr: ", (uint64_t)kmalloc_heap_end_addr());
|
||||||
console_println("");
|
console_println("");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user