Files
LithosAnanake/docs/working/architecture/HERMES-MESSAGE-CHANNEL-PHYSICS-DESIGN-20260708.md

17 KiB

Hermes Message/Channel Heat Physics — Design Doc

Date: 2026-07-08 (rev 2026-07-08a) Branch: lithosananke Status: Design only. Not implemented. Depends on a sibling architectural doc: HERMES-MESSAGE-BLOCK-STORAGE-DESIGN-20260708.md (rev c), which plans to move Hermes's message area off its current in-memory CREATE...ALLOT arena (MSG-ARENA/CH-ARENA, MSG-MAX=32/ CH-MAX=16) onto a dedicated, LBN-backed block range within the kernel ramdrive — a Hermes-owned physical BAM over Hermes's own blocks, built independently of and with zero coupling to Artemis (Artemis's own disk-block model is precedent for the allocator shape, not something Hermes's storage shares or depends on). That redesign is out of scope for this doc — this doc's window-sizing reasoning (section 2) is provisional against the current in-memory scale and will need revisiting once the block-backed redesign lands, since block-backed storage will likely change both the arena's real capacity and its churn characteristics. Sibling doc to VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md's "Fold-in question" section, which settled that messages and blocks get their own independent decay mechanisms — not merged into VMPhysics, not derived from its slope — each fit from its own rolling window, in separate follow-on docs. This is that doc for Hermes (messages + channels). See ARTEMIS-BLOCK-PHYSICS-DESIGN-20260708.md for the block-heat sibling. Author: Captain Bob / Claude Code

Correction (2026-07-09): this doc's original Goal section (rev a) claimed message/channel heat has no conservation invariant, citing a "settled fold-in decision" from VM-PHYSICS-DYNAMIC-FLEET-DESIGN- 20260705.md. That claim was wrong — it was derived without first reading .claude/HERMES.md, which is authoritative and states plainly: "Every object Hermes manages — messages and channels — participates in K≡1.0. Fleet K includes the thermal contribution of in-flight messages and open channels." A full Tripod-docs audit (TRIPOD.md/HERMES.md/ ARTEMIS.md) confirmed this directly contradicts what's below. The Goal section is corrected in place. The rest of this doc's mechanism — an independently-fit decay rate for messages and channels, its own HermesHeatWindow struct, its own rolling window, separate from word-level and VM-fleet physics — is unaffected: HERMES.md requires message/channel heat to participate in K≡1.0, not that it share VM-fleet's inference mechanism. Those are separable, and only the conservation claim was wrong.


Problem

capsules/hermes/init.4th hardcodes 65208 CONSTANT Q-DECAY (Block 4100) and applies it identically to two different entities via straight multiplicative shrink:

  • MSG-COOL-ONE/MSG-COOL-ALL (Block 4108): heat = heat * Q-DECAY for every live message in the 32-slot MSG-ARENA.
  • CH-COOL-ALL (Block 4114): same formula, walked over the live channel linked list (16-slot CH-ARENA).

This is the exact "before" picture VM-PHYSICS-DYNAMIC-FLEET-DESIGN already diagnosed for VM heat: 65208 was independently re-declared as an identical hand-picked constant in three separate files (compudynamics.4th, hermes/init.4th, artemis/init.4th) — copy-pasted, uncoordinated, with no regression anywhere. The VM-heat instance of this was fixed (capsule_vm_physics.c). Hermes's and Artemis's instances weren't — this doc is Hermes's fix.

No evidence exists that 65208 (≈0.9950 per cool-call, i.e. a message or channel's heat halves after roughly 138 MSG-COOL-ALL calls) was ever tuned against real message/channel churn. It's the same number VM heat used, which was itself never tuned against VM touch frequency either.

Goal

Replace the hardcoded Q-DECAY with a msg_decay_slope_q48 inferred from observed cooling behavior, mirroring the shape of the word-level and VM-fleet physics engines — same mechanics (rolling window, inferred rate, skip-rather-than-substitute-a-default when unwarmed) applied to a third kind of entity — without merging into either existing engine's state or inference mechanism.

Message and channel heat DOES participate in K≡1.0 — per .claude/HERMES.md's Compudynamic Invariant: "Fleet K includes the thermal contribution of in-flight messages and open channels... Hermes reaping a message or channel must rebalance K correctly." This doc does not design that redistribution mechanism — HERMES.md already establishes it conceptually (ACK → clean death, K redistributed; reap → K rebalanced), and HERMES.md's own G8 gap tracks the one known incompleteness (HERMES- K exists and totals live message/channel heat, but wiring it into Hera's K-FLEET is deferred pending an async inter-VM return-value model). What this doc does design is orthogonal to that: the rate at which a live message or channel's heat decays between creation and reap. Whether that decay is later credited back to the fleet correctly is HERMES.md's concern and G8's tracked gap, not this doc's.

Why this can't just copy either existing mechanism verbatim

Not word-level's shape (rolling window of word_ids + trajectory replay against current heat): words are permanent dictionary entries — extract_heat_trajectory (inference_engine.c) replays a window of word_ids and looks up each one's current heat, exploiting that a word never stops existing. Messages and channels are transient: MSG-REAP frees a message's arena slot back to the free list once its heat cools past whatever threshold or it's explicitly consumed, and a freed slot's "current heat" is meaningless — it may already belong to an unrelated message by the time anyone would replay a window containing its old ID. Per-entity trajectory replay does not work here.

Not VM-fleet's shape either, but structurally the closest analog: capsule_vm_physics.c (rev o) doesn't curve-fit a trajectory — it recovers the transfer rate directly, because the transfer law is known exactly (amount = elapsed_us * slope_q48 >> 16) and any unclamped touch can be inverted for an exact per-sample rate; the fleet-wide fleet_transfer_slope_q48 is the median of recovered per-sample rates over a 64-deep window of aggregate touch samples, not per-VM. The same "known law, recover the rate, aggregate not per-entity" shape applies here: heat_new = heat_old * Q-DECAY is also a fully known multiplicative law, so there is nothing to curve-fit — what's actually being inferred is not "what is the decay law" (known) but "is the current decay rate tuned right for observed churn," which is a different question than either existing engine answers, addressed below.

The mechanism

1. What "inferring the rate" actually means here

Because the decay law itself is known and fixed-shape (heat *= rate), there's no unknown functional relationship to recover the way VM-fleet recovers fleet_transfer_slope_q48 from touch samples, or the way Loop #6 recovers decay_slope_q48 from a heat trajectory. What can legitimately vary and be worth observing: how fast messages/channels are actually being reaped relative to how fast they're being created. If churn is high (many MSG-ALLOC/MSG-REAP cycles per cool-sweep), a slow decay rate lets dead-weight messages linger, consuming arena slots that new messages need. If churn is low, an aggressive decay rate discards heat signal (e.g. NACK-retry priority, per hermes/init.4th's own MSG-REDELIVER-NACKED) before it's useful.

msg_decay_slope_q48 should therefore be a function of observed reap throughput relative to arena occupancy, not a re-derivation of the Q-DECAY law itself. Concretely: a rolling window of samples taken once per MSG-COOL-ALL call, each recording (live_count, reaped_since_last); periodically re-fit msg_decay_slope_q48 so that, at the observed churn rate, an average message crosses the reap threshold within a target window (e.g. "most messages should be reaped within N cool-sweeps of going idle" — the specific target is an open question, see Resolved Items).

This is a genuinely different inference problem than either existing engine solves, which is why it doesn't reduce to copying infer_decay_slope_q48 or vm_physics_tick's median-of-rates — it's closer in spirit to Loop #5's window-width inference (which also asks "is the current parameter well- matched to observed dynamics," not "what is the true decay curve").

2. HermesHeatWindow — new struct, two instances (message, channel)

typedef struct {
    uint32_t live_count;       /* live entities at this sample */
    uint32_t reaped_since_last; /* entities freed since the prior sample */
} HermesHeatSample;

typedef struct {
    HermesHeatSample samples[HERMES_HEAT_WINDOW_DEPTH];
    uint32_t head;
    uint32_t count;   /* saturates at HERMES_HEAT_WINDOW_DEPTH */
    int      is_warm; /* count >= HERMES_HEAT_WINDOW_DEPTH */
} HermesHeatWindow;

Two instances — one for messages, one for channels — since they have independent arenas, independent churn rates, and are proposed to have independent decay slopes as this doc's own conservative default (see Open questions). Not one shared window; MSG-MAX=32 and CH-MAX=16 are different enough scales that conflating them risks exactly the kind of unjustified shared-constant problem this doc exists to fix.

HERMES_HEAT_WINDOW_DEPTH: proposed 16, not yet settled, and provisional against a scale that's about to change. MSG-MAX=32 and CH-MAX=16 are today's in-memory arena limits — far smaller than word-level's millions of executions or even VM-fleet's VM_FLEET_WINDOW_DEPTH=64, so a window deeper than the arena itself provides no additional information at current scale. But per this doc's Status header, the message area is planned to move onto a dedicated block range, which will likely raise that capacity substantially (Artemis's equivalent, ART-DATA-BLKS, is 22,998 — three orders of magnitude larger than MSG-MAX=32). 16 is a starting guess against today's scale, not derived from anything, and should be treated as provisional until the block-backed redesign's real capacity is known (see Resolved Items).

3. Recording — passive, hooked at the existing cool-sweep call sites

MSG-COOL-ALL and CH-COOL-ALL already sweep their entire arena once per call; a new hermes_heat_sample(HermesHeatWindow*, uint32_t live_count, uint32_t reaped_since_last) call at the end of each sweep is the only new call site needed — no new dispatch points, matching both existing engines' "passive observer, hooked at what's already there" discipline.

reaped_since_last bookkeeping is the one real implementation question this doc leaves open. MSG-REAP and channel reaping aren't currently instrumented to report a count to the cool-sweep — that's new plumbing, not free. Whether it's a shared counter incremented at each reap site and read-and-reset by the cool-sweep, or something else, is unresolved (see Resolved Items) — flagged rather than guessed at.

4. Inference — heartbeat-gated, same skip-don't-substitute discipline

hermes_heat_tick(), called from HERMES-TICK's existing heartbeat path (mirroring vm_physics_heartbeat_tick's hook into each VM's own tick): when a window is warm, re-fits that window's decay slope; when unwarmed, leaves the current slope untouched rather than substituting a default — same philosophy both existing engines already use.

The actual fit function (how (live_count, reaped_since_last) samples become a new msg_decay_slope_q48) is intentionally not specified in concrete formula form here — see Resolved Items. Writing a specific formula now, without the real observed-churn data either existing engine's design had before it settled on its final shape (VM-fleet went through revisions b through o before landing on median-of-rates), would be inventing content rather than designing it.

5. Application — replaces Q-DECAY Q.* at both cool-sweep sites

MSG-COOL-ONE and CH-COOL-ALL's Q-DECAY Q.* becomes msg_decay_slope_q48 Q.* / ch_decay_slope_q48 Q.* (or a single msg_decay_slope_q48 if the message/channel split above turns out unwarranted after real observation — see Resolved Items). No change to the multiplicative decay shape itself, only to where the rate comes from.

Consumer code migration

  • hermes/init.4th Block 4100: 65208 CONSTANT Q-DECAY deleted.
  • hermes/init.4th Block 4108 (MSG-COOL-ONE/MSG-COOL-ALL): Q-DECAY reference replaced with a call to fetch the current msg_decay_slope_q48 (new C primitive, e.g. MSG-DECAY-SLOPE@); MSG-COOL-ALL gains the sample-recording call described in section 3.
  • hermes/init.4th Block 4114 (CH-COOL-ALL): same pattern for channel heat.
  • New kernel-only file pair, matching capsule_vm_physics.c's own scope (Hermes-specific state, no hosted-build equivalent since the hosted build has no Hermes):
    • src/starkernel/capsule/hermes_heat_physics.c
    • include/starkernel/hermes_heat_physics.h
  • New FORTH-visible primitives (backed by the new file): MSG-DECAY-SLOPE@, CH-DECAY-SLOPE@ (or one shared accessor if sections above's split is resolved differently), plus a status word mirroring VM-PHYSICS-STATUS's shape for diagnostics.

What gets deleted

  • hermes/init.4th Block 4100: 65208 CONSTANT Q-DECAY.

Nothing else — MSG-COOL-ONE/MSG-COOL-ALL/CH-COOL-ALL keep their existing structure, only the decay-rate source changes.

Verification approach

  • Three-arch acceptance as usual (kernel-only code, #ifdef __STARKERNEL__ gates per .claude/CLAUDE.md).
  • A targeted test: drive real message/channel churn (repeated SEND-BROADCAST-TEST-style traffic) at two different rates across separate boot runs, confirm msg_decay_slope_q48 converges to different values reflecting the different observed churn — the direct analog of confirming VM-fleet's fleet_transfer_slope_q48 actually responds to real touch frequency rather than sitting at its seed value.
  • Confirm HermesHeatWindow's dead-entity handling (a message reaped mid-window) doesn't fault or corrupt the sample count — same class of test as VM-fleet's "kill-during-warm-up" test (VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md).

Explicitly out of scope

  • Moving Hermes's message/channel storage onto a dedicated LBN-backed block range (mirroring Artemis's ART-DATA-BLKS). Flagged in this doc's Status header as a real dependency, but the redesign itself — which blocks, what the on-disk layout looks like, how it interacts with MSG-ARENA/CH-ARENA's current in-memory shape — is a separate, not-yet-written architectural doc, not part of this heat-decay design.
  • Artemis block heat — sibling doc, ARTEMIS-BLOCK-PHYSICS-DESIGN-20260708.md.
  • Designing how K is credited/debited at message and channel alloc/reap — already governed conceptually by HERMES.md's Compudynamic Invariant section; the one known incompleteness (wiring HERMES-K into Hera's K-FLEET) is HERMES.md's own tracked G8 gap, not something this doc redesigns. This doc's decay-rate inference operates on top of that redistribution, not in place of it.
  • The multi-level DoE rewrite (word/VM-fleet/message/block as four independent metric spaces) — this doc gives message/channel heat its mechanism; observing it as part of a real experiment is separate, larger work already on the punch list.
  • Implementation — this is a design doc. Writing hermes_heat_physics.c is follow-on work, not part of this pass.

Resolved items

None yet — this is a first-pass design doc, not an iterated one. Unlike VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md's extensive revision history (rev a through t, each resolving a question found during real implementation and three-arch verification), this doc has had no implementation pass yet to surface and resolve open questions against.

Open questions (explicitly not resolved — flagged, not guessed at)

  1. HERMES_HEAT_WINDOW_DEPTH value. Proposed 16 (matching CH-MAX), not derived from real observed churn data. Needs a real implementation pass with instrumented boot logs (same discovery method VM-fleet used: doe_log.c's per-VM heat CSV columns exposed the dead-flat-trajectory bug that led to VM-fleet's rev-f seed-value fix) before treating this as settled.
  2. The concrete fit function for turning (live_count, reaped_since_last) samples into msg_decay_slope_q48. Section 4 deliberately stops short of a formula. A first candidate worth prototyping: target reap-latency-in-cool-sweeps, adjust the slope proportionally to the ratio of observed to target latency — but this is a starting hypothesis to implement and test against, not a specified design.
  3. reaped_since_last bookkeeping mechanism — where the reap-count plumbing lives and how it's reset per sample window.
  4. Whether messages and channels genuinely need independent slopes, or whether observed data shows one shared msg_decay_slope_q48 suffices — section 2 assumes independence as this doc's own conservative default (matching the pattern word-level and VM-fleet physics each use their own independently-fit rate), but this should be revisited once real data exists.