12. Information budget: tolerances from what the data can detect
Status: Design, 2026-10-07. Nothing in this document has been implemented or measured. Every speedup is a feasibility hypothesis with a testing milestone (§8), like the numbers in 06-compute.md §10.
Most error controls in the compute design end in a tolerance that has no stated basis: the reduction tolerance of the fidelity lattice, the residual of the regulation solve, the length of averaging windows, the number of replicas per behavioral state, the Markov-state construction, the shooting-window length for chaotic gradients, and the precision of messages between neurons. This document sets all of them from one principle:
An approximation is good enough when the frozen data cannot tell it apart from the exact model, and compute is spent where it buys the most detectable information.
Detectability is measured in nats, in observation space, through the same forward models (indicator, behavior) that the benchmark scores. The log score frozen at M0 (11-prediction-and-history.md) uses the same currency.
1. The information cost of an approximation
Let M be the exact model (reference fidelity, exact solves, infinite sampling) and A an approximation of it: a reduction, a finite average, an inexact solve or a compressed message. For a benchmark or falsification task T with frozen dataset design D (which neurons are imaged, at what rate and noise, how many animals and trials, which interventions), define the information cost
Δ_T(A) = KL( p_M(D) ‖ p_A(D) )
the divergence between the distributions that M and A assign to a whole dataset of that design. It needs only the design and the observation model, not the measured outcomes, so it can be computed for held-out tasks without leakage (03-evaluation.md, leakage audit).
Why it is the right tolerance. By Pinsker's inequality, for any test φ on the dataset,
| P_M(φ rejects) − P_A(φ rejects) | ≤ TV(p_M, p_A) ≤ √(Δ_T / 2)
and the same bound limits the change in the expectation of any statistic with values in [0, 1]. A per-task budget τ_T therefore bounds how much the approximation can change any pass/fail outcome, any falsification verdict and any bounded score on that task. With the illustrative τ_T = 0.02 nats, no test's rejection probability moves by more than 0.1. τ_T is set at M0, per task, well inside the task's equivalence margin δ_t.
Local form. For an approximation that perturbs some quantity u (parameters, slice coordinates, sensor averages, a model output) by a small δ,
Δ_T ≈ ½ δᵀ F_T δ, F_T = Fisher information of the dataset design with respect to u
- For Gaussian observation noise with covariance Σ_obs on predicted means μ(u), F_T = (∂μ/∂u)ᵀ Σ_obs⁻¹ (∂μ/∂u), and Δ_T = ½ δμᵀ Σ_obs⁻¹ δμ exactly for a mean shift.
- For distributional tasks (E3, E4) scored by a statistic s with sampling covariance V_T at the frozen dataset size, Δ_T ≈ ½ δsᵀ V_T⁻¹ δs: half the squared shift in standard errors. τ_T = 0.02 corresponds to a shift of 0.2 standard errors.
- Diagonal blocks of F_T are estimated from score samples: draw synthetic datasets from the model, take one vector–Jacobian product of the log-likelihood per sample, and accumulate outer products of the per-neuron blocks. m samples give all diagonal blocks at the cost of m backward passes.
Scope. For distributional tasks Δ is computed for the frozen task statistics. By the data-processing inequality this bounds what those tests can detect, not every possible analysis of the raw data. An approximation that passes the budget for one benchmark version must be re-checked when a later version adds data, statistics or interventions; cached artifacts record the Δ and the benchmark version they were certified for (§7). Undetectable on a dataset is not the same as scientifically unnecessary: the Necessity Map (03-evaluation.md) remains the test for that.
2. Combining approximations
A configuration uses many approximations at once. To first order:
- Deterministic errors (biases) from reductions, inexact solves and compressed messages add as vectors. √Δ is a seminorm in the F_T metric, so the triangle inequality gives √Δ_bias ≤ Σ_i √Δ_i.
- Zero-mean independent noise from finite sampling adds in expectation: E[Δ_noise] = Σ_j ½ tr(F_T C_j), where C_j is the covariance of error source j. The cross term with the bias vanishes.
The budget constraint for task T is
( Σ_i √Δ_i^bias )² + Σ_j E[Δ_j^noise] ≤ τ_T
with τ_T split between the bias and noise terms at M0. Because the bias bound is conservative, the coupled comparisons of 06-compute.md §9 measure the total directly on validation runs and calibrate the local estimates.
3. Allocating compute: Fisher-weighted Neyman allocation
Many knobs are separable: replicas per behavioral state, averaging length per single-cell solve, number of shadow neurons per type, initial-state samples per type. For knob j with cost c_j per unit (replica, simulated second), error variance s_j per unit, and Fisher weight λ_j (the F_T-weighted size of its error), choosing units n_j to minimize total cost Σ c_j n_j subject to Σ ½ λ_j s_j / n_j ≤ τ_noise gives
n_j ∝ √( λ_j s_j / c_j ), minimum cost = ( Σ_j √(c_j λ_j s_j) )² / (2 τ_noise)
This is Neyman's optimal stratified allocation with Fisher weights. Uniform allocation costs (Σ c_j)(Σ λ_j s_j) / (2 τ_noise), which by the Cauchy–Schwarz inequality is never less. The gain is large when λ is heavy-tailed. That is expected in the fly, where most neurons are not imaged and affect observables only through propagation, but it must be measured (§8). Knobs with λ_j s_j below the budget share get no compute beyond the minimum that the solver's own validity checks require.
4. Regulation solve
4.1 Fisher weights for the slow loop
Monte Carlo error in the regulation solve enters as error δe_i in each neuron's averaged sensors. It moves the slice coordinates by δa = (J G)⁻¹ δe, including network coupling, and from there the observables. The relevant Fisher matrix is therefore with respect to the sensor errors,
F_e = (J G)⁻ᵀ F_a (J G)⁻¹
computed with the implicit-function adjoint of 06-compute.md §5.3 (per-neuron k×k solves plus a few Krylov iterations for coupling). Neuron i's weight is the k×k block F_e,ii. For independent zero-mean errors across neurons, E[Δ_noise] = ½ Σ_i tr(F_e,ii C_i), so only diagonal blocks are needed.
4.2 Where the weights act
- Stratified averaging. The state-conditioned average E[σ] = Σ_s π_s σ̄_s has noise contribution π_s² S_s / n_s from state s, with S_s the long-run (integrated-autocorrelation) covariance of the sensors in that state. Replicas per state then follow §3 with λ s = π_s² Σ_i tr(F_e,ii S_s,i). This replaces the unweighted rule "rare states get extra replicas" (06-compute.md §5.4).
- Shared network window. The network run is shared by all neurons, so its length cannot be chosen per neuron. Its required length is set by the Fisher-weighted aggregate Σ_i tr(F_e,ii S_i), not by the worst-converged neuron in an unweighted norm. Neurons that the data barely see no longer dictate the window.
- Per-neuron single-cell solves (step 2 of the decomposition) are separable and follow §3 directly.
4.3 Adaptive precision across iterations
Far from the root, an iteration needs only enough precision to point in the right direction. Pseudo-transient continuation (06-compute.md §5.2) therefore uses increasing precision: at iteration m, the Monte Carlo standard error of e is held at θ ‖e_m‖ (θ ≈ 0.3–0.5), so the sample size grows as ‖e_m‖⁻². If the residual contracts by a factor r per iteration, the total cost over all iterations is at most 1/(1 − r²) times the cost of the final one, instead of the number of iterations times that cost. Under quadratic (Newton) convergence the ratio is smaller still.
Two parts of the solve are not relaxed by Fisher weighting, because they protect equilibrium selection, not observable accuracy:
- The precision schedule during continuation is measured in the solver's own per-neuron metric (scaled by the loop gain), not in F_e. A direction the data cannot see can still carry a trajectory across a separatrix to another branch, which would change observables nonlocally. Early-iteration noise is also kept below a fraction of the local stability margin λ_min.
- Certification checks 2 (loop-gain stability) and 3 (basin check) run at their own required precision for every type (06-compute.md §5.2). Only the final stopping tolerance, check 1, is set by the budget: the residual is accepted when its contribution ½ eᵀ F_e e plus the expected noise term fits within the solve's share of τ.
5. Sensor-targeted Markov state models
Stratified averaging (06-compute.md §5.4) is valid only if the state decomposition is Markovian at the chosen lag. Clustering general activity and then testing Markovianity can fail and lengthens averaging when it does. The states are needed for one purpose: estimating sensor averages and their sensitivities. So build them for that purpose.
Construction. Learn a coarse state s_t = f(recent past) by a predictive information bottleneck: maximize I(s_t ; future of the target process) − β⁻¹ I(s_t ; past). The target process is the Fisher-weighted sensor vector W σ_t, where W keeps the leading directions of Σ_i F_e,ii (the sensor combinations that matter for observables), together with s_{t+τ} itself. Variational estimators of this objective and kinetic-variance objectives used for molecular Markov models (VAMPnets) are the implementation starting points.
Why it helps. The exact minimal sufficient statistic of the past for the future of a stationary process (its causal states) evolves as a Markov chain. An approximate bottleneck is therefore Markov to the degree that it is predictively sufficient, and it can be tested on that basis. Behavioral regimes with identical weighted sensor statistics are merged, so fewer states are needed.
Checks. The implied-timescale test is retained. Added: the residual conditional mutual information I(W σ_future ; past | s_t), estimated as the held-out log-score gain of a richer history model over the state model, must be below the noise share of τ after scaling by the effective data size. The law of total expectation, E[σ] = Σ π_s E[σ | s], holds for any partition; the Markov property is what makes π estimable from short replicas.
6. Data-rate gate for trajectory-level tasks
6.1 Necessary condition
Tracking a chaotic trajectory requires observations that deliver information at least as fast as the dynamics create it. For deterministic dynamics, estimating the state with bounded error over a channel requires capacity at least the restoration entropy of the system (Matveev and Pogromsky), which is at least the topological entropy. By the variational principle and Pesin's formula for SRB measures, this is at least the sum of the positive Lyapunov exponents. In bits per second:
C_obs ≥ Σ_{λ_i > 0} λ_i / ln 2 (necessary for bounded-error trajectory tracking)
The imaging setup is a fixed encoder, not an optimally designed one, so it can do no better than the coding bound. Its capacity is bounded above by summing Gaussian-channel capacities over imaged neurons and behavioral channels:
C_obs ≤ Σ_imaged i B_i log₂(1 + SNR_i) + C_behavior
where B_i is the effective bandwidth (the smaller of half the frame rate and the indicator's kinetic bandwidth) and SNR_i the signal-to-noise ratio of neuron i's trace. Intrinsic noise (channel gating, stochastic release) adds entropy and makes tracking harder, so the condition stays necessary in the stochastic case.
6.2 Consequences
- If the bound fails (C_obs upper bound < Σλ⁺ / ln 2), no estimator can track trajectories with bounded error for that dataset design. Trajectory losses beyond the short-horizon regime are then ill posed. The affected tasks are specified as distributional (E3, E4 statistics; forecasts scored by proper scores on distributions) and fitted with shadowing or likelihood-ratio gradients. Teacher forcing over long horizons is not used for them. The gate is evaluated per task design, before fitting.
- If the bound holds, assimilation and trajectory fitting are possible in principle but not guaranteed (the condition is necessary, not sufficient).
- Shooting windows. Gradients through a window of length T_w grow by up to e^{λ_max T_w}. Capping that growth at κ_grad gives T_w ≤ ln(κ_grad) / λ_max, and the number of shooting nodes per simulated second is λ_max / ln κ_grad. This replaces the open-ended curriculum of 06-compute.md §6 with a computed window, refined by the measured finite-time exponents.
- Predictability horizon. A forecast that starts with state uncertainty ε₀ (from the assimilation posterior) and tolerates error ε_tol has expected trajectory skill up to T_pred ≈ ln(ε_tol / ε₀) / λ_max. This gives model-based guidance for choosing the horizons each forecasting task freezes, and for where trajectory scoring should give way to distributional scoring (11-prediction-and-history.md). It is not a universal neural horizon.
The Lyapunov spectrum is already measured on the worm at M2 (10-risks.md). The gate adds the capacity estimate for each dataset design.
7. Messages between neurons: bandlimited, at the noise floor
State stays fp32 (06-compute.md §4.6). The waveforms exchanged between neurons at window boundaries (temporal blocking, §4.3 there) and between devices (partitioned runs, §4.5 there) are messages, not state, and carry less information than their raw size suggests:
- Bandwidth. Graded transmission is low-pass: release kinetics, receptor kinetics and the postsynaptic membrane filter it. A graded source's output waveform is represented by samples at a rate of at least twice the class's transmission bandwidth f_c (with a guard band) and reconstructed by bandlimited interpolation, instead of one value per electrical step. f_c per class comes from B1-style transfer impedances and the receptor kinetics. At 25 µs electrical steps and f_c of a few hundred Hz, this is a reduction of more than an order of magnitude in exchanged samples. Spiking sources remain event-based (spike times).
- Resolution. Message quantization noise is set below a fraction of the intrinsic noise referred to the same signal (stochastic release, channel noise) at G1/S1 fidelity. In deterministic configurations, where no intrinsic noise exists, it is set by the information budget. Block floating point with 8–16 bits per sample is the expected range.
- Where it helps. This reduces waveform-buffer traffic and inter-device communication. It does not reduce per-step state traffic, which temporal blocking and replica batching address.
- Validation. Exactness is lost, so compressed messages are a fidelity level like any other: their Δ is measured by coupled comparisons against uncompressed messages, and CI checks the nonlinearity of presynaptic release (Hill-type), which can move energy above f_c.
8. Budget, milestones and limits
| Item | Estimate | Tested in | Plan changes if |
|---|---|---|---|
| Budget τ_T and its split per task | Set per task, inside δ_t | M0 (frozen with the task specification) | Budgets too tight for any affordable configuration: revisit task design or dataset size, not the budget |
| Agreement of local Δ estimates with coupled comparisons | Within a factor of 2 | A0 extension, M2b, M4 | Worse: use coupled comparisons as the primary measure, local estimates only for ranking |
| Saving from Fisher-weighted allocation in the regulation solve | Unknown; depends on the tail of λ | M4 (worm), P0.5 (larva) | < 2×: keep uniform allocation for simplicity |
| Saving from adaptive precision across iterations | ~(number of iterations) × (1 − r²) | A0 extension, M2b | Branch switches in basin checks rise: lower θ or keep fixed precision during continuation |
| States needed by sensor-targeted models vs. activity clustering | Fewer | M4 | No reduction, or Markov tests fail more often: revert to activity clustering |
| Data-rate gate outcome per task | Unknown | M2 (Lyapunov spectrum and imaging capacity on the worm) | Fails for E3 trajectory tasks: those tasks are frozen as distributional at M0 or in the next benchmark version |
| Message compression ratio and its Δ | > 10× fewer exchanged samples for graded sources | M1 profiler | Δ exceeds its share: raise f_c or resolution per class |
Limits.
- Local estimates use the Fisher matrix at the current parameters. Large errors can change the regime (a branch switch, a bifurcation), where Δ is nonlocal. The coupled comparisons and certification checks remain the backstop, and the budget does not replace them.
- The budget depends on dataset size. Larger future datasets shrink tolerances, and configurations must be re-certified for each benchmark version.
- The data-rate bound assumes an SRB invariant measure for the identification with Lyapunov exponents. It is a necessary condition only and says nothing about the difficulty of the estimation problem when the condition holds.
- Mutual-information estimates in high dimensions are biased. CMI is estimated as a difference of held-out log scores between fitted models, which is a bound conditional on those models, not an exact value.
9. What is borrowed and what is new
Borrowed: Pinsker's inequality and Stein-type detectability; local Fisher approximations to KL divergence; Neyman allocation in stratified sampling; dynamic sample sizes in stochastic Newton-type methods; the predictive information bottleneck and causal states; VAMPnets for molecular Markov models; data-rate theorems for estimation and control of unstable and chaotic systems; the sampling theorem.
New in CAKE, and untested:
- A single observation-space information budget sets every tolerance in the emulator: reductions, solves, sampling, state construction and messages. It is computed from the frozen dataset design, so it applies to held-out tasks without leakage.
- A data-rate feasibility gate decides, before fitting, which tasks can be scored as trajectories and which only as distributions.
- Markov state models for regulation are built to predict Fisher-weighted sensor futures, not general activity, with a predictive-sufficiency check in the same currency.
Decisions D50–D54 in decisions.md.