← The living map spec / 12-information-budget

12. Information budget: tolerances from what the data can detect

Status: Design, 2026-10-07. Nothing in this document has been implemented or measured. Every speedup is a feasibility hypothesis with a testing milestone (§8), like the numbers in 06-compute.md §10.

Most error controls in the compute design end in a tolerance that has no stated basis: the reduction tolerance of the fidelity lattice, the residual of the regulation solve, the length of averaging windows, the number of replicas per behavioral state, the Markov-state construction, the shooting-window length for chaotic gradients, and the precision of messages between neurons. This document sets all of them from one principle:

An approximation is good enough when the frozen data cannot tell it apart from the exact model, and compute is spent where it buys the most detectable information.

Detectability is measured in nats, in observation space, through the same forward models (indicator, behavior) that the benchmark scores. The log score frozen at M0 (11-prediction-and-history.md) uses the same currency.

1. The information cost of an approximation

Let M be the exact model (reference fidelity, exact solves, infinite sampling) and A an approximation of it: a reduction, a finite average, an inexact solve or a compressed message. For a benchmark or falsification task T with frozen dataset design D (which neurons are imaged, at what rate and noise, how many animals and trials, which interventions), define the information cost

Δ_T(A) = KL( p_M(D) ‖ p_A(D) )

the divergence between the distributions that M and A assign to a whole dataset of that design. It needs only the design and the observation model, not the measured outcomes, so it can be computed for held-out tasks without leakage (03-evaluation.md, leakage audit).

Why it is the right tolerance. By Pinsker's inequality, for any test φ on the dataset,

| P_M(φ rejects) − P_A(φ rejects) |  ≤  TV(p_M, p_A)  ≤  √(Δ_T / 2)

and the same bound limits the change in the expectation of any statistic with values in [0, 1]. A per-task budget τ_T therefore bounds how much the approximation can change any pass/fail outcome, any falsification verdict and any bounded score on that task. With the illustrative τ_T = 0.02 nats, no test's rejection probability moves by more than 0.1. τ_T is set at M0, per task, well inside the task's equivalence margin δ_t.

Local form. For an approximation that perturbs some quantity u (parameters, slice coordinates, sensor averages, a model output) by a small δ,

Δ_T ≈ ½ δᵀ F_T δ,      F_T = Fisher information of the dataset design with respect to u

Scope. For distributional tasks Δ is computed for the frozen task statistics. By the data-processing inequality this bounds what those tests can detect, not every possible analysis of the raw data. An approximation that passes the budget for one benchmark version must be re-checked when a later version adds data, statistics or interventions; cached artifacts record the Δ and the benchmark version they were certified for (§7). Undetectable on a dataset is not the same as scientifically unnecessary: the Necessity Map (03-evaluation.md) remains the test for that.

2. Combining approximations

A configuration uses many approximations at once. To first order:

The budget constraint for task T is

( Σ_i √Δ_i^bias )²  +  Σ_j E[Δ_j^noise]  ≤  τ_T

with τ_T split between the bias and noise terms at M0. Because the bias bound is conservative, the coupled comparisons of 06-compute.md §9 measure the total directly on validation runs and calibrate the local estimates.

3. Allocating compute: Fisher-weighted Neyman allocation

Many knobs are separable: replicas per behavioral state, averaging length per single-cell solve, number of shadow neurons per type, initial-state samples per type. For knob j with cost c_j per unit (replica, simulated second), error variance s_j per unit, and Fisher weight λ_j (the F_T-weighted size of its error), choosing units n_j to minimize total cost Σ c_j n_j subject to Σ ½ λ_j s_j / n_j ≤ τ_noise gives

n_j  ∝  √( λ_j s_j / c_j ),      minimum cost  =  ( Σ_j √(c_j λ_j s_j) )² / (2 τ_noise)

This is Neyman's optimal stratified allocation with Fisher weights. Uniform allocation costs (Σ c_j)(Σ λ_j s_j) / (2 τ_noise), which by the Cauchy–Schwarz inequality is never less. The gain is large when λ is heavy-tailed. That is expected in the fly, where most neurons are not imaged and affect observables only through propagation, but it must be measured (§8). Knobs with λ_j s_j below the budget share get no compute beyond the minimum that the solver's own validity checks require.

4. Regulation solve

4.1 Fisher weights for the slow loop

Monte Carlo error in the regulation solve enters as error δe_i in each neuron's averaged sensors. It moves the slice coordinates by δa = (J G)⁻¹ δe, including network coupling, and from there the observables. The relevant Fisher matrix is therefore with respect to the sensor errors,

F_e = (J G)⁻ᵀ F_a (J G)⁻¹

computed with the implicit-function adjoint of 06-compute.md §5.3 (per-neuron k×k solves plus a few Krylov iterations for coupling). Neuron i's weight is the k×k block F_e,ii. For independent zero-mean errors across neurons, E[Δ_noise] = ½ Σ_i tr(F_e,ii C_i), so only diagonal blocks are needed.

4.2 Where the weights act

4.3 Adaptive precision across iterations

Far from the root, an iteration needs only enough precision to point in the right direction. Pseudo-transient continuation (06-compute.md §5.2) therefore uses increasing precision: at iteration m, the Monte Carlo standard error of e is held at θ ‖e_m‖ (θ ≈ 0.3–0.5), so the sample size grows as ‖e_m‖⁻². If the residual contracts by a factor r per iteration, the total cost over all iterations is at most 1/(1 − r²) times the cost of the final one, instead of the number of iterations times that cost. Under quadratic (Newton) convergence the ratio is smaller still.

Two parts of the solve are not relaxed by Fisher weighting, because they protect equilibrium selection, not observable accuracy:

  1. The precision schedule during continuation is measured in the solver's own per-neuron metric (scaled by the loop gain), not in F_e. A direction the data cannot see can still carry a trajectory across a separatrix to another branch, which would change observables nonlocally. Early-iteration noise is also kept below a fraction of the local stability margin λ_min.
  2. Certification checks 2 (loop-gain stability) and 3 (basin check) run at their own required precision for every type (06-compute.md §5.2). Only the final stopping tolerance, check 1, is set by the budget: the residual is accepted when its contribution ½ eᵀ F_e e plus the expected noise term fits within the solve's share of τ.

5. Sensor-targeted Markov state models

Stratified averaging (06-compute.md §5.4) is valid only if the state decomposition is Markovian at the chosen lag. Clustering general activity and then testing Markovianity can fail and lengthens averaging when it does. The states are needed for one purpose: estimating sensor averages and their sensitivities. So build them for that purpose.

Construction. Learn a coarse state s_t = f(recent past) by a predictive information bottleneck: maximize I(s_t ; future of the target process) − β⁻¹ I(s_t ; past). The target process is the Fisher-weighted sensor vector W σ_t, where W keeps the leading directions of Σ_i F_e,ii (the sensor combinations that matter for observables), together with s_{t+τ} itself. Variational estimators of this objective and kinetic-variance objectives used for molecular Markov models (VAMPnets) are the implementation starting points.

Why it helps. The exact minimal sufficient statistic of the past for the future of a stationary process (its causal states) evolves as a Markov chain. An approximate bottleneck is therefore Markov to the degree that it is predictively sufficient, and it can be tested on that basis. Behavioral regimes with identical weighted sensor statistics are merged, so fewer states are needed.

Checks. The implied-timescale test is retained. Added: the residual conditional mutual information I(W σ_future ; past | s_t), estimated as the held-out log-score gain of a richer history model over the state model, must be below the noise share of τ after scaling by the effective data size. The law of total expectation, E[σ] = Σ π_s E[σ | s], holds for any partition; the Markov property is what makes π estimable from short replicas.

6. Data-rate gate for trajectory-level tasks

6.1 Necessary condition

Tracking a chaotic trajectory requires observations that deliver information at least as fast as the dynamics create it. For deterministic dynamics, estimating the state with bounded error over a channel requires capacity at least the restoration entropy of the system (Matveev and Pogromsky), which is at least the topological entropy. By the variational principle and Pesin's formula for SRB measures, this is at least the sum of the positive Lyapunov exponents. In bits per second:

C_obs  ≥  Σ_{λ_i > 0} λ_i / ln 2          (necessary for bounded-error trajectory tracking)

The imaging setup is a fixed encoder, not an optimally designed one, so it can do no better than the coding bound. Its capacity is bounded above by summing Gaussian-channel capacities over imaged neurons and behavioral channels:

C_obs  ≤  Σ_imaged i  B_i log₂(1 + SNR_i)  +  C_behavior

where B_i is the effective bandwidth (the smaller of half the frame rate and the indicator's kinetic bandwidth) and SNR_i the signal-to-noise ratio of neuron i's trace. Intrinsic noise (channel gating, stochastic release) adds entropy and makes tracking harder, so the condition stays necessary in the stochastic case.

6.2 Consequences

The Lyapunov spectrum is already measured on the worm at M2 (10-risks.md). The gate adds the capacity estimate for each dataset design.

7. Messages between neurons: bandlimited, at the noise floor

State stays fp32 (06-compute.md §4.6). The waveforms exchanged between neurons at window boundaries (temporal blocking, §4.3 there) and between devices (partitioned runs, §4.5 there) are messages, not state, and carry less information than their raw size suggests:

8. Budget, milestones and limits

Item Estimate Tested in Plan changes if
Budget τ_T and its split per task Set per task, inside δ_t M0 (frozen with the task specification) Budgets too tight for any affordable configuration: revisit task design or dataset size, not the budget
Agreement of local Δ estimates with coupled comparisons Within a factor of 2 A0 extension, M2b, M4 Worse: use coupled comparisons as the primary measure, local estimates only for ranking
Saving from Fisher-weighted allocation in the regulation solve Unknown; depends on the tail of λ M4 (worm), P0.5 (larva) < 2×: keep uniform allocation for simplicity
Saving from adaptive precision across iterations ~(number of iterations) × (1 − r²) A0 extension, M2b Branch switches in basin checks rise: lower θ or keep fixed precision during continuation
States needed by sensor-targeted models vs. activity clustering Fewer M4 No reduction, or Markov tests fail more often: revert to activity clustering
Data-rate gate outcome per task Unknown M2 (Lyapunov spectrum and imaging capacity on the worm) Fails for E3 trajectory tasks: those tasks are frozen as distributional at M0 or in the next benchmark version
Message compression ratio and its Δ > 10× fewer exchanged samples for graded sources M1 profiler Δ exceeds its share: raise f_c or resolution per class

Limits.

9. What is borrowed and what is new

Borrowed: Pinsker's inequality and Stein-type detectability; local Fisher approximations to KL divergence; Neyman allocation in stratified sampling; dynamic sample sizes in stochastic Newton-type methods; the predictive information bottleneck and causal states; VAMPnets for molecular Markov models; data-rate theorems for estimation and control of unstable and chaotic systems; the sampling theorem.

New in CAKE, and untested:

  1. A single observation-space information budget sets every tolerance in the emulator: reductions, solves, sampling, state construction and messages. It is computed from the frozen dataset design, so it applies to held-out tasks without leakage.
  2. A data-rate feasibility gate decides, before fitting, which tasks can be scored as trajectories and which only as distributions.
  3. Markov state models for regulation are built to predict Fisher-weighted sensor futures, not general activity, with a predictive-sufficiency check in the same currency.

Decisions D50–D54 in decisions.md.