The Cost of a Bit in a World: Landauer Meets the Family Law at a Pure Number
1. One Number a Reader Can Check
This note reports one exact identity, one quarantined near-miss, and a set of leads for somebody who works on information systems professionally. The identity: combine Landauer's 1961 result that erasing one bit costs
Three ingredients, none of them new. From 1961: erasing one bit of information costs at least
E_bit =
Erasing one bit inside a world costs eleven percent of that world's door, exactly, on every rung. No temperature, no material, no scale. A bit is a fixed fraction of the price of admission. That is the whole result, and it takes a calculator: ln2 = 0.693147, divided by 2π = 6.283185, gives 0.110318. The rest of this note is context and leads.
2The Chain, and Its One Loose Link
The author's proposition runs: information is energy, energy is mass, mass is
3Rung by Rung — and One Landing That Must Be Quarantined
Applying the identity across
4What the Framework Finds When It Looks at a Neural Network
The organising question behind these notes was whether the framework's claim — one geometric object at many magnifications — reaches an information system. Taking the framework's own four terms in turn, and reporting honestly: Algebra: yes, and non-commutativity is load-bearing. The core operation is matrix multiplication, associative but order-sensitive, and attention makes the asymmetry explicit — a head computes Q·Kᵀ, so position i attending to j is not j attending to i. Order matters for the reason it matters in the quaternions: the operation is a directed relation. But the algebra is associative and has no division; it is nowhere on the Cayley–Dickson ladder, and Cayley's theorem has no purchase on it. Differential equations: genuinely, not metaphorically. The residual stream updates as x(ℓ+1) = x(ℓ) + F(x(ℓ)), which is Euler integration of dx/dℓ = F(x,ℓ) at unit step. Depth is an integration variable and the residual stream is a field on layer-time and token-position. (Lead: the Neural ODE literature.) Spaces: yes, and one echoes P2. Features appear to be directions in a shared vector space of order 10³–10⁴ dimensions, and there are far more features than dimensions, packed at near-orthogonal rather than orthogonal angles — 'superposition'. The logic is P2's affordability argument: bending the open directions is expensive, so the structure is carried by packing. More things than room, made affordable by geometry. Things interacting: attention heads, but the interaction is non-local. Any position reaches any other in one step — no delay, no neighbourhood, and no seal. That is the sharpest structural difference from a framework in which exit cones and horizons are foundational.
5What It Does Not Find — Three Absences, One Fatal
No conservation law, and this is the fatal one. This framework's predictive power comes almost entirely from books that must balance: the Noether ledger, redshift as transfer and never write-off, the winding number that cannot unwind, eternity as topological protection. An information system has no such invariant — layer normalisation destroys scale information at every step, there is no variational principle, no continuous symmetry of a dynamics, therefore no Noether theorem. Things are computed and discarded. No metric that responds to content. The activation space has an inner product but it does not change because something is present. Nothing bends. (Section 6 qualifies this: there is a responsive metric, but it lives in parameter space.) Time is not a symmetry — with one exception worth recording. Every layer carries its own weights, so layer forty is not layer four advanced; it is a different apparatus, and the autoregressive mask gives a degenerate light cone, unbounded toward the past and sealed toward the future. But at inference the weights do not move while the residual stream advances: the stage is fixed, the play ticks. That is P3, exactly. It fails during training, when the stage itself moves. Which suggests the useful statement is that learning is precisely the regime in which the stage is not eternal.
6Two Places the Subjects Already Touch
The first is a curvature theory of learning, roughly forty years old, that nobody appears to have set beside P2. The parameter space of a statistical model carries the Fisher information metric g_ij — equivalently, the second derivative of the Kullback–Leibler divergence between nearby models, which is to say the curvature of the statistical manifold, a genuine Riemannian metric with genuine curvature. And natural gradient descent, the update g⁻¹·∇L, is the statement that what looks like a force in coordinates is curvature in the metric. Learning is not a push down a landscape; it is flow in a curved space, and the apparent force is an artefact of the wrong coordinates. That is P2, reached independently, in a domain with no gravity in it. (Lead: Amari.) The second runs the other way, and is the more important of the two. The expectation was that this framework might be implemented in information science. But
7The Compute Conjecture, and Its Honest Ledger
The author's conjecture: the computing centres now being built will prove unnecessary once one can take a second derivative of information flow and find the downhill, the natural curvature, faster than brute force. What is right about it: the second derivative of information flow is the Fisher metric of Section 6, and using it to find the downhill is the natural gradient. This is already the frontier rather than a speculation — the optimizer that trained most current systems is a diagonal preconditioner, which is a crude curvature approximation, and better ones have been taking ground. (Leads: K-FAC, Shampoo, SOAP, Muon.) The obstruction is dimensional, not conceptual. Full curvature is an
8Small State Spaces: the Sixty-Four, the Beiwerk, the Encounter
The author's preference for the I Ging's sixty-four as a manageable engine has a historical hook: Leibniz, corresponding with the Jesuit Bouvet around 1703, recognised that the sixty-four hexagrams in the Fu Xi arrangement are the binary numbers 0 to 63 — six lines, broken or unbroken, six bits. One distinction must be kept, though: sixty-four states is a catalogue, while sixty-four weighted variables is a 64-dimensional continuum, incomparably larger. And what makes the I Ging an engine rather than a list is not the sixty-four but the changing lines — a transition rule. States plus a rule. (The engine the Chinese trusted with time was the sexagenary sixty, ten stems meshed against twelve branches.) The smallness is defensible rather than merely appealing: Rule 110 has two states and eight rules and is Turing complete; Conway's Life has two states. Tiny kernels iterated with interaction produce unbounded complexity. And a modern instance of exactly the proposal exists — vector-quantised models compress to a discrete codebook of a few hundred entries and generate by transitions among codes. A learned I Ging, with the hexagrams found rather than inherited. The author's Beiwerk claim — reduce to a small core and the remainder is trimming — is measurable and roughly right. The intrinsic dimension of a task, the smallest random subspace within which training still succeeds, comes out in the hundreds to low thousands for large models; low-rank adaptation works at rank eight, sometimes rank one. The structure underneath really is small. Not sixty-four, but on a logarithmic scale sixty-four and a thousand are neighbours and both are nowhere near 10¹². And his sketch of an encounter — two people on a path resolving sex, then friend-or-foe, then uneasy trust, then a signal of harmlessness, then relaxation, modelled as two state vectors sharing a space with a Taktfrequenz — has more empirical backing than a sketch deserves. The sequence has been filmed: cross-cultural work found the eyebrow flash, a raise of order a sixth of a second, apparently universal, and the smile appears to descend from the primate fear grimace. Interactional synchrony is measured — postural sway, breathing, speech rhythm entraining, rising with rapport. And the programme has already succeeded in its simplest case: the social force model treats pedestrians as particles with repulsive terms and a clock, and reproduces lane formation and bottleneck arching without being told to. Humans as vectors in a space with forces is not fantasy. It works, for movement. What nobody has is a state vector carrying the noble, the drunkard, the one who means harm.
9The Two Missing Things
The algebra. Nobody has it. What is observed is that features add — the old word-vector arithmetic, king minus man plus woman — which would make the algebra abelian: no order-dependence whatever. That is almost certainly incomplete; it holds approximately and fails in ways nobody can characterise. P1 is precisely the objection: if the space describes anything that composes, order must matter somewhere. Finding the non-commutative structure underneath the observed addition is the sharpest open problem in these notes, and there are no good candidates. The weight scale — the author's question, one noble and one peasant, how do they interact. This probably has an answer and it is a power law: word frequencies follow Zipf's inverse rank over many decades, and feature activations look similarly heavy-tailed, a few directions firing constantly and most almost never. The space is steeply aristocratic, and the exponent is measured and unexplained. The cleanest example of the field's pre-theoretic condition is the scaling laws: loss falls as a power law in parameters over many orders of magnitude, with no derivation whatsoever. In this series' six columns that is a measured anchor awaiting its theorem — the condition physics was in before somebody found the law under the regularity. And one obstruction recurs, worth naming once. The I Ging has three thousand years of commentary on what a hexagram means when it changes and not one row of what actually happened afterwards. The encounter has an enormous descriptive ethology and almost no record of how particular encounters went. A dynamics cannot be fitted to labels that were never recorded. That may be the honest reason the I Ging stayed an oracle rather than becoming a physics, and it is the first thing to ask a specialist: if a corpus with an outcome column exists anywhere, several questions below become tractable at once.
10Carnot's Order
The author's remark that the mathematics of these systems is undiscovered but already in use describes the normal order rather than an anomaly, and the precedent is closer to home than it looks. Newcomen's engine ran in 1712; Carnot's analysis came in 1824 and Clausius's formulation in the 1850s. A hundred and forty years of working machines before the law — and Carnot did not derive it from first principles. He derived it from studying the engines. The theory came out of the machine. That is a methodological hint, and it points at taking these systems apart rather than at theory first. It is also where the two subjects meet a second time:
11Open Questions, Named
For whoever picks this up, in the order they seem tractable. One: the motif count. Domain-independent recurring circuits are real — induction heads, which implement 'if A was followed by B, expect B after A' and care not at all whether A and B are amino acids or ingredients, form abruptly in training and turn on in-context learning; and the same motifs recur across models trained on different data. How many are there? An empirical question with a number at the end. If small, the compute conjecture strengthens sharply; if it runs to millions, it dies. Two: the non-commutative structure of Section 9. Three: whether the Fisher metric has exploitable structure. Four: whether the 10⁴ sample-efficiency gap is a metric. Five: whether somebody who counts bits professionally recognises the gearbox. Six: whether anything at all is conserved. Seven: whether an outcome column exists. Note that questions three and six are the same question. Find the potential and the geometry becomes cheap, because one stops having to measure the landscape everywhere. Find the potential and something is conserved along its level sets — which is the invariant Section 5 says is missing. The author's word for it was intent.
12The Invitation, and the Flags
The invitation is the same one the chemistry notes make, and in the same spirit. Section 1 is a two-line calculation: ln2 divided by 2π. Section 3 is four multiplications. Anyone who works on information systems can check both in five minutes and then judge whether a bit being eleven percent of a world's door is a coincidence of unit conventions or a statement about what information is. This note does not know. It reports the number and its one near-miss, with the near-miss quarantined and the reasons for quarantining it written down. The flags, without softening. Every reference here is a lead rather than a source: the whole note was written from memory without a literature search, so it fails
13The Sentence
Erasing one bit inside any world of this ladder costs eleven percent of that world's door — Landauer's heat and P2's admission price differing by ln2 over 2π and by nothing else — which is the tightest form the author's chain has yet taken, and which suggests that the borrowing between these two subjects runs the way nobody expected: not a field theory implemented in information, but a field theory whose engine already counts bits on a wall, and whose one outstanding calculation is a bit-counting problem wearing a physicist's coat.
References
R. Landauer (1961), irreversibility and heat generation; L. Szilard (1929); J. D. Bekenstein, Phys. Rev. D 7, 2333 (1973); R. Clausius (1865); S. Carnot (1824); G. W. Leibniz, correspondence with J. Bouvet (~1703). S. Amari, information geometry and the natural gradient. Neural ODEs: Chen and co-authors (~2018). Superposition: Elhage and co-authors (~2022). Induction heads and universality: Olsson, Elhage and co-authors. Scaling laws: Kaplan and co-authors (~2020); Hoffmann and co-authors (~2022). Intrinsic dimension: Li and co-authors (~2018); Aghajanyan and co-authors (~2020). Low-rank adaptation: Hu and co-authors. Curvature-aware optimizers: Martens and Grosse (K-FAC), Shampoo, SOAP, Muon. The bitter lesson: Sutton (2019). Cellular automata: Wolfram; Cook on Rule 110. Vector quantisation: van den Oord and co-authors. Zipf's law: Zipf. Ethology of greeting: Eibl-Eibesfeldt. Interactional synchrony: Condon. Social force model: Helbing and Molnár (~1995). The Yijing, Wilhelm and Legge translations. And the papers of this series: the Allgemeine Feldtheorie (P1–P3, §4