17 KiB
Hermes Message/Channel Heat Physics — Design Doc
Date: 2026-07-08 (rev 2026-07-08a)
Branch: lithosananke
Status: Design only. Not implemented. Depends on a sibling
architectural doc: HERMES-MESSAGE-BLOCK-STORAGE-DESIGN-20260708.md
(rev c), which plans to move Hermes's message area off its current
in-memory CREATE...ALLOT arena (MSG-ARENA/CH-ARENA, MSG-MAX=32/
CH-MAX=16) onto a dedicated, LBN-backed block range within the kernel
ramdrive — a Hermes-owned physical BAM over Hermes's own blocks, built
independently of and with zero coupling to Artemis (Artemis's own
disk-block model is precedent for the allocator shape, not something
Hermes's storage shares or depends on). That redesign is out of scope for
this doc — this doc's window-sizing reasoning (section 2) is provisional
against the current in-memory scale and will need revisiting once the
block-backed redesign lands, since block-backed storage will likely
change both the arena's real capacity and its churn characteristics.
Sibling doc to
VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md's "Fold-in question" section,
which settled that messages and blocks get their own independent decay
mechanisms — not merged into VMPhysics, not derived from its slope —
each fit from its own rolling window, in separate follow-on docs. This is
that doc for Hermes (messages + channels). See
ARTEMIS-BLOCK-PHYSICS-DESIGN-20260708.md for the block-heat sibling.
Author: Captain Bob / Claude Code
Correction (2026-07-09): this doc's original Goal section (rev a)
claimed message/channel heat has no conservation invariant, citing a
"settled fold-in decision" from VM-PHYSICS-DYNAMIC-FLEET-DESIGN- 20260705.md. That claim was wrong — it was derived without first reading
.claude/HERMES.md, which is authoritative and states plainly: "Every
object Hermes manages — messages and channels — participates in K≡1.0.
Fleet K includes the thermal contribution of in-flight messages and open
channels." A full Tripod-docs audit (TRIPOD.md/HERMES.md/
ARTEMIS.md) confirmed this directly contradicts what's below. The Goal
section is corrected in place. The rest of this doc's mechanism — an
independently-fit decay rate for messages and channels, its own
HermesHeatWindow struct, its own rolling window, separate from
word-level and VM-fleet physics — is unaffected: HERMES.md requires
message/channel heat to participate in K≡1.0, not that it share
VM-fleet's inference mechanism. Those are separable, and only the
conservation claim was wrong.
Problem
capsules/hermes/init.4th hardcodes 65208 CONSTANT Q-DECAY (Block 4100)
and applies it identically to two different entities via straight
multiplicative shrink:
MSG-COOL-ONE/MSG-COOL-ALL(Block 4108):heat = heat * Q-DECAYfor every live message in the 32-slotMSG-ARENA.CH-COOL-ALL(Block 4114): same formula, walked over the live channel linked list (16-slotCH-ARENA).
This is the exact "before" picture VM-PHYSICS-DYNAMIC-FLEET-DESIGN
already diagnosed for VM heat: 65208 was independently re-declared as
an identical hand-picked constant in three separate files
(compudynamics.4th, hermes/init.4th, artemis/init.4th) — copy-pasted,
uncoordinated, with no regression anywhere. The VM-heat instance of this
was fixed (capsule_vm_physics.c). Hermes's and Artemis's instances
weren't — this doc is Hermes's fix.
No evidence exists that 65208 (≈0.9950 per cool-call, i.e. a message or
channel's heat halves after roughly 138 MSG-COOL-ALL calls) was ever
tuned against real message/channel churn. It's the same number VM heat
used, which was itself never tuned against VM touch frequency either.
Goal
Replace the hardcoded Q-DECAY with a msg_decay_slope_q48 inferred from
observed cooling behavior, mirroring the shape of the word-level and
VM-fleet physics engines — same mechanics (rolling window, inferred rate,
skip-rather-than-substitute-a-default when unwarmed) applied to a third
kind of entity — without merging into either existing engine's state or
inference mechanism.
Message and channel heat DOES participate in K≡1.0 — per
.claude/HERMES.md's Compudynamic Invariant: "Fleet K includes the
thermal contribution of in-flight messages and open channels... Hermes
reaping a message or channel must rebalance K correctly." This doc does
not design that redistribution mechanism — HERMES.md already establishes
it conceptually (ACK → clean death, K redistributed; reap → K rebalanced),
and HERMES.md's own G8 gap tracks the one known incompleteness (HERMES- K exists and totals live message/channel heat, but wiring it into Hera's
K-FLEET is deferred pending an async inter-VM return-value model). What
this doc does design is orthogonal to that: the rate at which a
live message or channel's heat decays between creation and reap. Whether
that decay is later credited back to the fleet correctly is HERMES.md's
concern and G8's tracked gap, not this doc's.
Why this can't just copy either existing mechanism verbatim
Not word-level's shape (rolling window of word_ids + trajectory
replay against current heat): words are permanent dictionary entries —
extract_heat_trajectory (inference_engine.c) replays a window of
word_ids and looks up each one's current heat, exploiting that a word
never stops existing. Messages and channels are transient: MSG-REAP
frees a message's arena slot back to the free list once its heat cools
past whatever threshold or it's explicitly consumed, and a freed slot's
"current heat" is meaningless — it may already belong to an unrelated
message by the time anyone would replay a window containing its old ID.
Per-entity trajectory replay does not work here.
Not VM-fleet's shape either, but structurally the closest analog:
capsule_vm_physics.c (rev o) doesn't curve-fit a trajectory — it
recovers the transfer rate directly, because the transfer law is known
exactly (amount = elapsed_us * slope_q48 >> 16) and any unclamped touch
can be inverted for an exact per-sample rate; the fleet-wide
fleet_transfer_slope_q48 is the median of recovered per-sample rates
over a 64-deep window of aggregate touch samples, not per-VM. The same
"known law, recover the rate, aggregate not per-entity" shape applies
here: heat_new = heat_old * Q-DECAY is also a fully known multiplicative
law, so there is nothing to curve-fit — what's actually being inferred is
not "what is the decay law" (known) but "is the current decay rate
tuned right for observed churn," which is a different question than either
existing engine answers, addressed below.
The mechanism
1. What "inferring the rate" actually means here
Because the decay law itself is known and fixed-shape (heat *= rate),
there's no unknown functional relationship to recover the way VM-fleet
recovers fleet_transfer_slope_q48 from touch samples, or the way Loop #6
recovers decay_slope_q48 from a heat trajectory. What can legitimately
vary and be worth observing: how fast messages/channels are actually
being reaped relative to how fast they're being created. If churn is
high (many MSG-ALLOC/MSG-REAP cycles per cool-sweep), a slow decay
rate lets dead-weight messages linger, consuming arena slots that new
messages need. If churn is low, an aggressive decay rate discards heat
signal (e.g. NACK-retry priority, per hermes/init.4th's own
MSG-REDELIVER-NACKED) before it's useful.
msg_decay_slope_q48 should therefore be a function of observed reap
throughput relative to arena occupancy, not a re-derivation of the
Q-DECAY law itself. Concretely: a rolling window of samples taken once
per MSG-COOL-ALL call, each recording (live_count, reaped_since_last);
periodically re-fit msg_decay_slope_q48 so that, at the observed churn
rate, an average message crosses the reap threshold within a target
window (e.g. "most messages should be reaped within N cool-sweeps of
going idle" — the specific target is an open question, see Resolved
Items).
This is a genuinely different inference problem than either existing
engine solves, which is why it doesn't reduce to copying infer_decay_slope_q48
or vm_physics_tick's median-of-rates — it's closer in spirit to Loop #5's
window-width inference (which also asks "is the current parameter well-
matched to observed dynamics," not "what is the true decay curve").
2. HermesHeatWindow — new struct, two instances (message, channel)
typedef struct {
uint32_t live_count; /* live entities at this sample */
uint32_t reaped_since_last; /* entities freed since the prior sample */
} HermesHeatSample;
typedef struct {
HermesHeatSample samples[HERMES_HEAT_WINDOW_DEPTH];
uint32_t head;
uint32_t count; /* saturates at HERMES_HEAT_WINDOW_DEPTH */
int is_warm; /* count >= HERMES_HEAT_WINDOW_DEPTH */
} HermesHeatWindow;
Two instances — one for messages, one for channels — since they have
independent arenas, independent churn rates, and are proposed to have
independent decay slopes as this doc's own conservative default (see
Open questions). Not one shared window; MSG-MAX=32
and CH-MAX=16 are different enough scales that conflating them risks
exactly the kind of unjustified shared-constant problem this doc exists
to fix.
HERMES_HEAT_WINDOW_DEPTH: proposed 16, not yet settled, and provisional
against a scale that's about to change. MSG-MAX=32 and CH-MAX=16 are
today's in-memory arena limits — far smaller than word-level's millions
of executions or even VM-fleet's VM_FLEET_WINDOW_DEPTH=64, so a window
deeper than the arena itself provides no additional information at
current scale. But per this doc's Status header, the message area is
planned to move onto a dedicated block range, which will likely raise
that capacity substantially (Artemis's equivalent, ART-DATA-BLKS, is
22,998 — three orders of magnitude larger than MSG-MAX=32). 16 is a
starting guess against today's scale, not derived from anything, and
should be treated as provisional until the block-backed redesign's real
capacity is known (see Resolved Items).
3. Recording — passive, hooked at the existing cool-sweep call sites
MSG-COOL-ALL and CH-COOL-ALL already sweep their entire arena once per
call; a new hermes_heat_sample(HermesHeatWindow*, uint32_t live_count, uint32_t reaped_since_last) call at the end of each sweep is the only new
call site needed — no new dispatch points, matching both existing engines'
"passive observer, hooked at what's already there" discipline.
reaped_since_last bookkeeping is the one real implementation question
this doc leaves open. MSG-REAP and channel reaping aren't currently
instrumented to report a count to the cool-sweep — that's new plumbing,
not free. Whether it's a shared counter incremented at each reap site and
read-and-reset by the cool-sweep, or something else, is unresolved (see
Resolved Items) — flagged rather than guessed at.
4. Inference — heartbeat-gated, same skip-don't-substitute discipline
hermes_heat_tick(), called from HERMES-TICK's existing heartbeat path
(mirroring vm_physics_heartbeat_tick's hook into each VM's own tick):
when a window is warm, re-fits that window's decay slope; when unwarmed,
leaves the current slope untouched rather than substituting a default —
same philosophy both existing engines already use.
The actual fit function (how (live_count, reaped_since_last) samples
become a new msg_decay_slope_q48) is intentionally not specified in
concrete formula form here — see Resolved Items. Writing a specific
formula now, without the real observed-churn data either existing engine's
design had before it settled on its final shape (VM-fleet went through
revisions b through o before landing on median-of-rates), would be
inventing content rather than designing it.
5. Application — replaces Q-DECAY Q.* at both cool-sweep sites
MSG-COOL-ONE and CH-COOL-ALL's Q-DECAY Q.* becomes
msg_decay_slope_q48 Q.* / ch_decay_slope_q48 Q.* (or a single
msg_decay_slope_q48 if the message/channel split above turns out
unwarranted after real observation — see Resolved Items). No change to
the multiplicative decay shape itself, only to where the rate comes
from.
Consumer code migration
hermes/init.4thBlock 4100:65208 CONSTANT Q-DECAYdeleted.hermes/init.4thBlock 4108 (MSG-COOL-ONE/MSG-COOL-ALL):Q-DECAYreference replaced with a call to fetch the currentmsg_decay_slope_q48(new C primitive, e.g.MSG-DECAY-SLOPE@);MSG-COOL-ALLgains the sample-recording call described in section 3.hermes/init.4thBlock 4114 (CH-COOL-ALL): same pattern for channel heat.- New kernel-only file pair, matching
capsule_vm_physics.c's own scope (Hermes-specific state, no hosted-build equivalent since the hosted build has no Hermes):src/starkernel/capsule/hermes_heat_physics.cinclude/starkernel/hermes_heat_physics.h
- New FORTH-visible primitives (backed by the new file):
MSG-DECAY-SLOPE@,CH-DECAY-SLOPE@(or one shared accessor if sections above's split is resolved differently), plus a status word mirroringVM-PHYSICS-STATUS's shape for diagnostics.
What gets deleted
hermes/init.4thBlock 4100:65208 CONSTANT Q-DECAY.
Nothing else — MSG-COOL-ONE/MSG-COOL-ALL/CH-COOL-ALL keep their
existing structure, only the decay-rate source changes.
Verification approach
- Three-arch acceptance as usual (kernel-only code,
#ifdef __STARKERNEL__gates per.claude/CLAUDE.md). - A targeted test: drive real message/channel churn (repeated
SEND-BROADCAST-TEST-style traffic) at two different rates across separate boot runs, confirmmsg_decay_slope_q48converges to different values reflecting the different observed churn — the direct analog of confirming VM-fleet'sfleet_transfer_slope_q48actually responds to real touch frequency rather than sitting at its seed value. - Confirm
HermesHeatWindow's dead-entity handling (a message reaped mid-window) doesn't fault or corrupt the sample count — same class of test as VM-fleet's "kill-during-warm-up" test (VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md).
Explicitly out of scope
- Moving Hermes's message/channel storage onto a dedicated LBN-backed
block range (mirroring Artemis's
ART-DATA-BLKS). Flagged in this doc's Status header as a real dependency, but the redesign itself — which blocks, what the on-disk layout looks like, how it interacts withMSG-ARENA/CH-ARENA's current in-memory shape — is a separate, not-yet-written architectural doc, not part of this heat-decay design. - Artemis block heat — sibling doc,
ARTEMIS-BLOCK-PHYSICS-DESIGN-20260708.md. - Designing how K is credited/debited at message and channel alloc/reap
— already governed conceptually by
HERMES.md's Compudynamic Invariant section; the one known incompleteness (wiringHERMES-Kinto Hera'sK-FLEET) isHERMES.md's own tracked G8 gap, not something this doc redesigns. This doc's decay-rate inference operates on top of that redistribution, not in place of it. - The multi-level DoE rewrite (word/VM-fleet/message/block as four independent metric spaces) — this doc gives message/channel heat its mechanism; observing it as part of a real experiment is separate, larger work already on the punch list.
- Implementation — this is a design doc. Writing
hermes_heat_physics.cis follow-on work, not part of this pass.
Resolved items
None yet — this is a first-pass design doc, not an iterated one. Unlike
VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md's extensive revision history
(rev a through t, each resolving a question found during real
implementation and three-arch verification), this doc has had no
implementation pass yet to surface and resolve open questions against.
Open questions (explicitly not resolved — flagged, not guessed at)
HERMES_HEAT_WINDOW_DEPTHvalue. Proposed 16 (matchingCH-MAX), not derived from real observed churn data. Needs a real implementation pass with instrumented boot logs (same discovery method VM-fleet used:doe_log.c's per-VM heat CSV columns exposed the dead-flat-trajectory bug that led to VM-fleet's rev-f seed-value fix) before treating this as settled.- The concrete fit function for turning
(live_count, reaped_since_last)samples intomsg_decay_slope_q48. Section 4 deliberately stops short of a formula. A first candidate worth prototyping: target reap-latency-in-cool-sweeps, adjust the slope proportionally to the ratio of observed to target latency — but this is a starting hypothesis to implement and test against, not a specified design. reaped_since_lastbookkeeping mechanism — where the reap-count plumbing lives and how it's reset per sample window.- Whether messages and channels genuinely need independent slopes,
or whether observed data shows one shared
msg_decay_slope_q48suffices — section 2 assumes independence as this doc's own conservative default (matching the pattern word-level and VM-fleet physics each use their own independently-fit rate), but this should be revisited once real data exists.