6. Computational design
This document describes the current design only. The reasons for each choice, and the alternatives rejected, are in decisions.md. All numbers are order-of-magnitude estimates; the milestone that measures each one is listed in §10. Tolerances for reductions, solves, sampling and messages come from the information budget in 12-information-budget.md.
1. Cost drivers
A naive implementation fails for these reasons. Every section below exists to remove one of them.
| Driver | Naive cost (fly CNS, full detail) | Addressed in |
|---|---|---|
| Memory traffic | ~15 GB of state traversed per 25 µs step ≈ 600 TB per simulated second; bandwidth-bound | §2, §4; messages: 12 §7 |
| Serial depth of regulation burn-in | ~10⁵ simulated seconds ≈ 4×10⁹ sequential steps per candidate Φ | §5; adaptive precision: 12 §4.3 |
| Burn-in nested in optimizer and ensemble | ~10²⁶–10²⁷ FLOP | §5, §6, §7; Fisher-weighted allocation: 12 §3–4 |
| Gradients of chaotic long-time statistics | Exponentially growing, meaningless gradients | §6; data-rate gate: 12 §6 |
| Gradients through body contact | Non-smooth, biased | §6 |
| Extracellular diffusion on EM geometry | Petavoxel grids | §3 (homogenization) |
| Recomputation across experiments and team members | Same burn-ins and reductions recomputed many times | §8 |
2. Data model and memory layout
The core abstraction separates what never changes during a run from what does, as MuJoCo separates mjModel from mjData:
| Object | Contents | Lifetime |
|---|---|---|
Organism |
Connectome, morphologies, cell types, transcriptome priors, provenance | Organism release (immutable, content-hashed) |
Model |
Compiled arrays: tree topology, compartment geometry, synapse index tables (CSR), mechanism groups, fidelity level per cell type | Built once per (Organism, fidelity configuration) |
Params |
Φ (per-type regulatory programs), latent edges, global physics; a JAX pytree | Changes every optimizer step |
State |
Voltages, gates, concentrations, vesicle pools, body state, RNG counters | Per replica, per time step |
Layout rules:
- Group neurons by mechanism signature (the set of CKL mechanisms they carry), then by compartment count. Each group is stored struct-of-arrays and updated by one kernel launch. This avoids warp divergence: different cell types carry different channel sets, so mixing them in one launch would serialize their code paths.
Modelis shared across replicas; onlyStateis per replica. Replicas per GPU ≈ (GPU memory −Model) /State.
3. Fidelity lattice
Every component has named fidelity levels. A configuration picks a level per component per cell type (or region).
| Component | Levels (cheap → reference) |
|---|---|
| Neuron electrical | N0 graded point neuron → N1 reduced, impedance-preserving (~10–50 compartments) → N2 full EM morphology |
| Channel gating | G0 deterministic → G1 stochastic (Markov / Langevin) |
| Chemical synapses | S0 lumped deterministic → S1 lumped stochastic → S2 per-site stochastic |
| Extracellular diffusion | D0 well-mixed neuropil compartments → D1 homogenized coarse grid → D2 EM-geometry voxels |
| Body | B0 recorded replay → B1 reduced mechanics (resistive force theory, MJX) → B2 continuum body + CFD |
Three mechanisms use the same lattice:
- Validation of reductions: each lower level is checked against the level above (on the worm, against full reference). "Within tolerance" means within the reduction's share of the task's information budget Δ ≤ τ_T (12-information-budget.md §1–2).
- Fidelity probes (§9): online error estimates in production runs.
- Multilevel Monte Carlo (§7): most samples from cheap levels, corrections from expensive ones.
Reduction is not ablation. Moving down the lattice approximates a physical process numerically. The Necessity Map (03-evaluation.md) removes a process entirely to test whether it is scientifically needed. Reductions may be used whenever their measured error is within tolerance; they do not wait for the Necessity Map.
Reduction details:
-
N1 neurons: transfer impedances depend on membrane conductances (which regulation and inference change), capacitance, axial resistivity, and, in the quasi-active regime, on the operating point at which active channels are linearized. They are therefore not fixed per organism release. Each N1 model is a parametric reduction valid over a declared domain Θ_dom of these quantities:
- Build: compute full-morphology transfer impedances at sample points of Θ_dom, then fit a 10–50 compartment model whose compartment parameters are functions of those quantities and which preserves impedances between synapse groups, integration sites and output sites (parametric model-order reduction).
- Certify: report sup over held-out points of Θ_dom of the relative impedance error, over the frequency band of interest. The reduction is used only where this bound is within tolerance.
- Leaving the domain (detected from the current parameters at each burn-in or fit) triggers a rebuild, and the fidelity probes (§9) check it online.
Transfer impedance is the quantity hypothesis H2 relies on.
-
Synapse lumping. Merging a group of synapses into one aggregate state is exact only if the aggregate dynamics close. Sufficient conditions:
- same presynaptic compartment (or presynaptic sites that are isopotential to tolerance), so all sites see the same presynaptic voltage and Ca²⁺ drive;
- same bouton parameter class: release machinery, facilitation and depression parameters, vesicle-pool size and replenishment rate;
- same receptor type and kinetic parameters, and the same postsynaptic compartment, hence the same driving force;
- no interaction between sites (transmitter spillover or shared receptor pools).
Under these conditions:
- Deterministic (mean-field) release: every site has the same transmitter waveform, and receptor Markov models are linear in their state given that waveform, so the merged model is exact. CI checks lumped against per-site equality.
- Stochastic release: given the pool states, the sum of independent binomials with a common release probability is binomial, Bin(Σ_k N_k, p). With identical, independent per-site replenishment, the total pool is a closed birth–death process. Release counts are therefore exact in distribution, not pathwise.
- Postsynaptic receptors under stochastic release: different sites receive transmitter at different times. Receptor binding is nonlinear in transmitter concentration, so the aggregate receptor state does not close. It is exact only in the low-occupancy regime, where binding is linear in transmitter. Otherwise it is an approximation.
When any condition fails (different boutons with separate Ca²⁺ nanodomains, depletion histories that diverge stochastically, different transmitter waveforms), lumping is an approximation. It is then a lattice level (S0/S1) like any other, validated against S2 per-site models. CI uses distributional tests (moments and two-sample tests of release counts and postsynaptic conductance) on worm circuits and on sampled fly synapse groups.
-
D1 diffusion: EM geometry is used offline to compute effective diffusivity tensors (tortuosity, volume fraction) per neuropil region; simulation runs on a ~1–5 µm grid with spectral or multigrid solvers.
4. Execution model
4.1 Whole brain per device
At the production fidelity (N1/G0/S1/D1), a fly CNS instance is small:
| Component | Count | State | Size (fp32) |
|---|---|---|---|
| Compartments | ~1.7×10⁵ neurons × ~30 ≈ 5×10⁶ | ~40 variables each | ~0.8 GB |
| Lumped synapse groups | ~10⁷ (estimate) | ~5–10 states each | ~0.2–0.4 GB |
| Diffusion grid, indicator, body | — | — | < 0.5 GB |
| Total State | ~1.5–2 GB |
The unit of parallelism is the replica, not a slice of the brain. Each GPU runs whole brains; replicas, candidate Φs and experimental conditions spread across GPUs with no communication during the forward pass.
4.2 Replica-batched kernels
Synaptic scatter and tree solves are dominated by reading index and topology arrays from Model, not by arithmetic. When R replicas share a GPU, each kernel loads those arrays once and applies them to all R replicas' State. Index traffic is divided by R and arithmetic intensity rises accordingly. Replicas that need different Models (different fidelity configurations) are batched separately.
4.3 Temporal blocking
Neurons interact only through synapses and gap junctions, so one neuron can be integrated for many steps in fast on-chip memory:
- A thread block loads one N1 neuron (~30 compartments × ~40 variables ≈ 5 KB) into shared memory.
- It integrates K steps (e.g. a ~1 ms window), reading presynaptic waveforms from the previous window iteration, and writes its output waveform at the window boundary.
- Spiking synapses are exact when the window is shorter than the minimum axonal plus synaptic delay. Graded synapses and gap junctions use waveform relaxation, which converges quickly for windows short relative to synaptic time constants.
- Memory traffic drops by ~K / (relaxation iterations); expected ~5–20×.
- Exchanged waveforms of graded sources are bandlimited and quantized at the noise floor; state stays fp32 (12-information-budget.md §7).
4.4 Time stepping
- Multirate operator splitting: electrical ~10–25 µs (spiking) or ~0.1 ms (graded neurons, chosen per neuron); synaptic/chemical ~0.1–1 ms; diffusion ~1–10 ms; mechanics ~0.1–1 ms.
- Gating variables use exponential (Rush–Larsen) integration, so step size is set by accuracy, not stability; checked by step-halving tests.
- Cable equations use implicit Euler or Crank–Nicolson with the Hines solver.
4.5 Partitioned runs (reference fidelity only)
N2 fly runs don't fit on one device. They partition the connectome with a min-cut on synaptic coupling. Spiking synapses synchronize at the minimum axonal plus synaptic delay (exact, as in NEST). Graded synapses and gap junctions use waveform relaxation across devices, communicating once per window, with the same message compression as §4.3.
4.6 Kernels and numerics
- Tree solves: one warp or thread block per N1 neuron, state in shared memory. N2 trees use parallel scheduling of the Hines elimination (dendritic hierarchical scheduling).
- Synaptic scatter: CSR sorted by postsynaptic target, segmented reductions; event-driven for spiking sources, dense for graded sources.
- Stack: JAX for composition and autodiff; Pallas/Triton kernels generated by the CKL compiler for solve, scatter and gating hot loops; CUDA and HIP behind one kernel interface; MJX for rigid bodies.
- Precision: fp32 state; compensated summation for conservation-critical accumulators (charge, ion mass). No reduced-precision state.
- Randomness: counter-based RNG (Philox) keyed by (replica, neuron, synapse, step). The random draws are therefore identical regardless of batching, partitioning or GPU count. This alone does not make trajectories identical: floating-point reductions depend on summation order, and chaotic dynamics amplify the differences.
- Reproducibility has two levels:
- Bitwise reproducibility holds within a reproducibility class: code version, compiler and flags, device architecture, partition and replica-batch layout. Kernels use fixed-order segmented reductions (no floating-point atomics) so that this holds. CI re-runs reference cases within each supported class and requires bitwise equality.
- Statistical reproducibility holds across classes: every Ladder statistic and every cached artifact agrees within its recorded tolerance. CI checks this across classes with two-sample tests on reference workloads. Results are never compared trajectory by trajectory across classes.
- Gradient memory: √T checkpointing within windows; adjoint state stored only at window boundaries.
5. Regulation solver (developmental burn-in)
5.1 What is computed
The developmental outcome is defined in 04-physics.md, L7: the adult regulatory state is the endpoint of the regulatory dynamics, started from z(t₀) ~ p₀,c and run along the developmental schedule in the closed loop. The solver never redefines that outcome. It computes it in one of two modes:
- Endpoint mode solves directly for the selected equilibrium. It is used only when the validity conditions of §5.5 hold.
- Transient mode integrates the averaged slow dynamics along the schedule. It is used whenever endpoint mode is invalid, and whenever the data are themselves transient (F3 and F7 time courses).
Direct simulation of days of regulation, with activity resolved, is the reference used to validate both modes on the worm (§5.9).
5.2 Endpoint mode: the slice equation
For rules with constant gain and no leak (L7 Case 1), each neuron's state stays on its slice z_i = z_i(t₀) + G_c a_i. The unknowns are the slice coordinates a ∈ ℝ^{Nk}: k per neuron, not n. The equation is
e(a) = s* − E_closed-loop[ σ( z(t₀) + G a ) ] = 0
This replaces the ill-posed equation E[sensor(θ)] = setpoint, which has a manifold of solutions per neuron and a singular Jacobian. The equivalent square system in z is
F(z) = [ s* − E[σ(z)] ; Nᵀ (z − z(t₀)) ] = 0, N spans null(Gᵀ)
Its Jacobian is nonsingular if and only if the loop gain J G is nonsingular, where J = ∂E[σ]/∂z. (If [−J; Nᵀ] v = 0, then v = G b for some b, and J G b = 0.)
Other rule forms:
- Leak. Solve G e(z) − D (z − z^ref) = 0 in full z, with Jacobian −(G J + D). Use the Nᵀ rows only for the directions in the left null space of [G D].
- Involutive learned rules. These use a flat parameterization: the network learns an invertible map w = φ(z) in which the dynamics are dw/dt = [B(w); 0] e. The last n − k flat coordinates are then conserved exactly, and the slice equation is solved over the first k.
- Non-involutive learned rules always use transient mode.
Algorithm. Pseudo-transient continuation on a, starting at a = 0 (the initial state):
a_{m+1} = a_m + (I/Δt_m + Ĵ G)⁻¹ e(a_m)
Δt_m grows as the residual falls (switched evolution relaxation)
For small Δt each step follows the regulatory flow, and for large Δt it becomes Newton's method. The solver therefore tends to converge to the equilibrium the trajectory reaches from the initial state, not to an arbitrary root. Ĵ G is a k×k block per neuron, estimated with k Jacobian–vector products using common random numbers.
Precision schedule. The Monte Carlo error of e(a_m) is held at a fixed fraction θ of the residual, so sampling grows as the residual falls and early iterations are cheap. This schedule is measured in the solver's own loop-gain-scaled metric and is kept below a fraction of the stability margin, so that noise cannot carry the iteration to another branch (12-information-budget.md §4.3).
Certification. An endpoint is accepted only if all of the following hold; otherwise that type switches to transient mode:
- The residual is within tolerance relative to the Monte Carlo error of E[σ]. The tolerance is the solve's share of the information budget, ½ eᵀ F_e e plus the expected noise term, with F_e the Fisher information of the dataset with respect to sensor errors (12-information-budget.md §4.1). Checks 2 and 3 are not relaxed by Fisher weighting.
- Every eigenvalue of the estimated loop gain has real part above a margin λ_min > 0. This is checked per neuron, and for the network by Arnoldi iteration for the least stable modes.
- The basin check passes: for a random ≥1% of neurons per type, and for every worm run, the endpoint agrees with transient-mode integration from a = 0.
A root that fails check 2 satisfies the sensor equations but is not a developmental outcome, because the dynamics move away from it.
5.3 Per-neuron decomposition (default)
- Run the closed-loop network briefly; record every neuron's input waveforms and sensory traces.
- Solve each neuron's slice equation (k unknowns) against its recorded inputs, holding the network fixed. These are ~1.7×10⁵ independent single-cell problems, batched by cell type.
- Re-run the network with the new states to refresh inputs. Repeat until the inputs stop changing.
Convergence condition. Split J = J_d + J_o, where J_d is block-diagonal (each neuron's sensors with respect to its own state, inputs held fixed) and J_o carries network coupling. To first order the outer iteration is block Jacobi on the slice equation. It converges if and only if
ρ( (J_d G)⁻¹ J_o G ) < 1
where ρ is the spectral radius, estimated during the run by power iteration with Jacobian–vector products. If the estimate exceeds 0.9, the solver switches to the joint solver (§5.6) without waiting for stagnation. Implicit-function gradients use the same structure: per-neuron k×k solves plus a few Krylov iterations for the coupling.
5.4 Stratified averaging with Markov state models
Sensor averages must cover the animal's behavioral states (e.g. worm forward runs, reversals, turns), which switch over tens of seconds. Long time-averaging windows would make this the most expensive step. Borrowing from molecular dynamics, where Markov state models (as used in Folding@home) avoid long trajectories:
- Build discrete behavioral/brain states by a predictive information bottleneck targeted at the Fisher-weighted sensor futures, not by clustering general activity (12-information-budget.md §5).
- Estimate a state-transition matrix T from many short replicas started across states.
- Compute state-conditioned sensor averages σ̄_s from short windows in each state, and weight them by the stationary occupancy π of T: E[σ] = Σ_s π_s σ̄_s.
- Replicas are allocated across states by Fisher-weighted Neyman allocation (weighted-ensemble sampling), so each state's contribution to the information budget is controlled, not merely its variance (12-information-budget.md §3, §4.2).
This replaces "run long enough to visit every state many times" with "run short windows in every state", which parallelizes. It is valid if the state decomposition is Markovian at the chosen lag time, which is checked with standard implied-timescale tests and with a residual conditional-mutual-information test of predictive sufficiency.
Sensitivities. The loop gain needs the derivative of the average, and regulation can shift behavior as well as activity within each state:
∂E[σ]/∂z = Σ_s π_s ∂σ̄_s/∂z + Σ_s σ̄_s ∂π_s/∂z
∂π = π (∂T) Z, Z = (I − T + 1π)⁻¹ (fundamental matrix)
The second term, from shifts in state occupancy, is estimated with likelihood-ratio weights on the transition counts.
5.5 Validity conditions: when endpoint mode and implicit differentiation apply
| # | Condition | Diagnostic | If violated |
|---|---|---|---|
| V1 | Loop gain well-conditioned | σ_min(Ĵ G) above its Monte Carlo error by a factor ≥ 10; condition number below κ_max (set at M4) | Transient mode; report proximity to a bifurcation |
| V2 | Selected equilibrium stable | All eigenvalues of Ĵ G have real part > λ_min | Reject root; transient mode |
| V3 | Equilibrium is the one reached from the initial state | Basin check (§5.2) | Transient mode for that type; cache results by trajectory |
| V4 | Regulation slow relative to activity mixing | ε = τ_mix/τ_reg < ε_max (default 0.1) | Simulate the offending (fast-path) coordinates in time together with activity |
| V5 | No active bounds | No regulated quantity at its saturation limit at the endpoint | Transient mode with complementarity at the bounds |
| V6 | Conservation structure known | Rule is Case 1, leak with a known left null space, or flat-parameterized involutive | Transient mode |
| V7 | Leak memory short relative to age | 1/λ_min(D) < 0.1 × developmental duration | Transient mode (the adult state is not an equilibrium) |
| V8 | Quantity of interest is an endpoint | Not a trajectory (F3, F7 time courses) | Transient mode |
Implicit differentiation (§6) is valid exactly when V1–V7 hold at the endpoint. Each run records which conditions held, for each type.
5.6 Joint solver (fallback)
If decomposition fails its convergence condition (§5.3), or input changes stop decreasing between outer iterations, switch to a joint solver on all slice coordinates together: averaged sensors from parallel replicas, stochastic pseudo-transient Newton–Krylov steps (heterogeneous multiscale method), common random numbers and control variates across iterations.
5.7 Transient mode
Integrate the averaged slow dynamics dz/dt = G(z) e(z) − D(z − z^ref) along the developmental schedule (wiring, body and rearing environment as functions of age), from z(t₀) ~ p₀,c.
- Integrator. Linearly implicit (Rosenbrock) steps sized by the slow dynamics. Each step needs one stratified average (§5.4) and one loop-gain estimate.
- Cost. Slow steps ≈ (developmental duration / τ_reg) × steps per τ_reg, typically 10²–10³ per neuron population. Each step costs one averaging window. This is measured at M4.
- Gradients. By the adjoint of the averaged slow dynamics, with checkpointing at slow steps. This is the only route to gradients for F3/F7 time courses and for non-involutive rules.
- Fast path with V4 violated. Integrated together with activity over the window in which it is fast (minutes), then handed to the slow integrator.
5.8 Amortized surrogate
A per-cell-type surrogate (neural operator) maps (Φ_c, initial state, input-waveform statistics, morphology features) to the slice coordinates a*, trained on solves the pipeline already performs. It warm-starts or replaces inner solves. Every prediction is certified like any endpoint (§5.2: residual, stability and, on the sampled subset, basin); if a* moves beyond tolerance, the true solve is used and added to the training set. The surrogate touches only the slow inner loop, never network physics.
5.9 Ground truth on the worm
The worm is cheap enough to run direct day-long regulation simulations with activity resolved. Every solver mode above is validated against them before use on the fly:
- agreement of endpoints from the same initial states and schedules;
- agreement of the predicted within-type covariance (L7) with the spread of endpoints over samples of p₀,c;
- agreement of gradients with finite differences of directly simulated endpoints;
- agreement of downstream Emulation Ladder scores.
The A0 demonstrator (prototypes/a0-regulatory) is the first, synthetic, instance of these checks.
6. Gradients and optimization
-
Stages by timescale:
- Linear-response pretraining: around a resting operating point, the first-order response of every neuron to stimulating neuron i is one JVP. One JVP per stimulated neuron gives a full linearized pair-response matrix, a cheap initialization for E2 fitting.
- Short horizon: E1 and E2 data (single-cell ephys, optogenetic pair responses) use short windows that chaos doesn't reach; ordinary backpropagation through time, all conditions batched as replicas.
- Long horizon: E3 statistics and regulation targets use shadowing gradients (LSS/NILSS) or ensemble likelihood-ratio estimators. NILSS cost grows with the number of positive Lyapunov exponents, which is measured on the worm.
-
Data-rate gate and shooting windows. Before fitting, each task design is checked against the necessary condition C_obs ≥ Σλ⁺ / ln 2 (imaging and behavioral channel capacity vs. entropy production). Tasks that fail it are fitted and scored as distributions only. Multiple-shooting windows are T_w ≤ ln(κ_grad) / λ_max, which caps gradient growth at κ_grad (12-information-budget.md §6).
-
Implicit differentiation through the selected equilibrium (§5.2), valid when V1–V7 of §5.5 hold. With z* = z(t₀) + G a* and e(a*, φ) = 0, for any program parameter φ (set points, the gain subspace, rule parameters, parameters of p₀,c through the reparameterization z(t₀) = μ_c + L_c ξ):
da*/dφ = (J G)⁻¹ [ ∂s*/∂φ − ∂E[σ]/∂φ |_z − J ( ∂z(t₀)/∂φ + (∂G/∂φ) a* ) ]Only k×k systems per neuron and the network coupling (§5.3) are solved. Regulation trajectories are never unrolled. Where V1–V7 fail, and for all trajectory data (F3, F7), gradients come from the adjoint of transient mode (§5.7).
-
Gradient checks. On a random subset of parameters and types, every run compares implicit gradients with finite differences of the certified solve (common random numbers), and on the worm with finite differences of directly simulated endpoints. The measured gradient error is recorded with the fit.
-
Gauge. Program gains are parameterized as G_c = U_c C_c (L7). Fits to equilibrium data optimize U_c and hold C_c fixed, because its gradient is exactly zero there. C_c is fitted only when transient data enter.
-
Body replay (B0) inside the gradient loop: recorded kinematics and sensory input replace the body, so contacts are never differentiated. The closed loop is used for burn-in and evaluation.
-
Optimizer: Gauss–Newton / Levenberg–Marquardt with geodesic acceleration (Transtrum, Machta & Sethna), designed for sloppy models; matrix-free via JVP/VJP and conjugate gradients in the stiff subspace from the M3 Fisher analysis. The number of serial outer iterations (~10²–3×10², against ~10³ for first-order methods) is a feasibility hypothesis tested at M4 (§10).
-
Latent edges as continuous variables: gap-junction and peptide edges are inferred with continuous sparsity priors (e.g. horseshoe or relaxed spike-and-slab on edge conductance), so they are fitted by gradients along with everything else. Discrete model selection with simulation-based inference is used only for the final shortlist of contested edges.
7. Uncertainty at scale
- Worm: full posterior ensembles.
- Fly: posterior mode, stiff subspace from Hessian/Fisher-vector products, sampling within it and a Laplace approximation outside.
- Initial-state distribution: predictions marginalize over p₀,c by sampling initial states (reparameterized, so gradients flow through the samples). For small Σ₀,c the linear covariance formula of L7 is used and checked against samples.
- Multilevel Monte Carlo over the fidelity lattice (§3) for posterior expectations, with samples per level set by Fisher-weighted Neyman allocation (12-information-budget.md §3).
- Experiment design: rank candidates by the linearized (Fisher / D-optimal) criterion; run full simulation-based information gain only for the top ~1%.
8. Content-addressed compute cache
Expensive intermediate results are cached and reused across the team, keyed by a hash of every input that affects them:
| Artifact | Key |
|---|---|
| N1 reductions, transfer impedances | Organism release, neuron ID, reduction settings, declared parameter domain Θ_dom (passive membrane, capacitance, axial resistivity, linearization operating points), code version |
| Burn-in endpoints | Model hash, Params hash (including p₀,c), initial-state samples or their seeds, developmental-schedule samples, rearing distribution, solver mode and settings, RNG seed, reproducibility class, certified information cost Δ and benchmark version |
| Conditioned state posteriors and forecasts (proposed) | Model and posterior hashes, conditioning-data hash and observation cutoff, history-state representation/version, particle or initial-state identities, assimilation policy, intervention, forecast horizon, observation model, solver settings, RNG seed, reproducibility class; contract in 11-prediction-and-history.md |
| Markov state models | Model hash, Params hash, clustering settings, reproducibility class |
| Surrogate training pairs | Cell type, Φ_c, initial-state statistics, input statistics hash |
| Benchmark scores | Model version, benchmark version (task specifications of 03-evaluation.md) |
Exactness. A cache hit is bitwise identical to recomputation only within one reproducibility class (§4.6). Across classes, a hit is statistically equivalent: it agrees within the tolerances recorded for that artifact type, and is labeled as such. The workflow layer (07-platform.md) consults the cache before scheduling any job.
Warm starts from nearby cache entries (e.g. the closest previous Φ) can change the answer when the regulation dynamics have several stable equilibria: a solver started near another branch can converge to it. A warm start is therefore allowed only if the result is certified as the endpoint reached from the specified initial state (§5.2, checks 1–3). Otherwise the solve restarts from a = 0. When certification holds, warm starts change only convergence speed, not the answer.
9. Error control
All error estimates below are reported in the information currency of 12-information-budget.md: the change in the predicted dataset distribution, Δ in nats, compared with the task's budget τ_T.
Local probes. Every production run embeds shadow neurons: higher-fidelity copies of a random ~1% of neurons, driven by the same inputs as their twins, with outputs not fed back. Divergence between shadow and twin gives an online estimate of local reduction error per cell type, under the inputs that the reduced network supplied. Types whose error exceeds tolerance are moved up the lattice (§3) in the next run.
Local error is not the error of the result: recurrence can amplify or suppress it, and closed-loop behavior can shift. Two further checks propagate local error to the quantities reported:
- Goal-oriented propagation. The adjoint of the observable of interest (a Ladder score, a pair response, a behavioral statistic) weights each type's local error, giving a first-order estimate of the error in that observable (dual-weighted residual). This identifies the types whose reduction matters for a given result, not merely those with large local error. The estimated shift δs is converted to Δ ≈ ½ δsᵀ V_T⁻¹ δs, using the statistic's sampling covariance at the frozen dataset size.
- Coupled comparisons. On validation runs, paired replicas with common random numbers are run at production fidelity and with a probed subnetwork (or the whole network, on the worm) raised to higher fidelity with feedback on. Differences are compared on the same Ladder statistics used for scoring (E2–E5), including closed-loop behavior. The goal-oriented estimate is calibrated against these comparisons.
Every published result carries its measured local error, the propagated error estimate, and the date and outcome of the latest coupled comparison for that configuration.
10. Budget and assumptions
All numbers in this section are feasibility hypotheses, not guarantees. Each has a milestone that tests it and a criterion that would change the plan. Benchmarks measure complete inference workloads (burn-in, gradients, ensembles, closed-loop evaluation) end to end, not isolated kernels.
Worm (reference fidelity throughout): ~10¹²–10¹³ FLOP per simulated second (~10¹³–10¹⁴ with continuum body and full Stokes fluid). Direct burn-in and full ensembles are expected to be affordable; this is tested at M4.
Fly CNS (production fidelity), illustrative:
| Item | Estimate | Tested in | Plan changes if |
|---|---|---|---|
| State per brain | ~1.5–2 GB | M0 (reduction sizes, lumped synapse counts) | > 1 device memory: partitioned production runs (§4.5), revisit D9 |
| Forward wall time per simulated second per GPU | ~0.3–1 s with temporal blocking and replica batching | M1 (roofline profiler), then end-to-end at P0.5 | > 5 s: re-plan fly fidelity configuration |
| Network simulated time per outer iteration | ~10³ s (decomposed regulation, stratified averaging) | M4 (worm), P0.5 (larva) | > 10⁴ s: surrogate-first burn-in or per-type direct fits for the fly |
| Serial outer iterations | ~10²–3×10² | M4 | > 10³: revisit D11 |
| Share of types needing transient mode (§5.5) | Unknown; assumed small | M2b, M4 | > 20%: budget transient mode as the default |
| Saving from information-budget allocation and adaptive precision in the regulation solve | Unknown; depends on the tail of the Fisher weights | A0 extension, M2b, M4 | < 2×: keep uniform allocation (12-information-budget.md §8) |
| Imaging capacity vs. entropy production (data-rate gate) | Unknown | M2 | Gate fails for trajectory tasks: freeze them as distributional |
| Hardware for a full fit | ~64–256 GPUs | End-to-end benchmark at P0.5 | — |
| Wall time for a full fit | ~days | End-to-end benchmark at P0.5 | > 1 month: reduce scope before P2 |
Assumptions that could break the budget:
- Regulation decomposition satisfies its convergence condition (§5.3) for most types (else the joint solver, §5.6, ~5–10× slower).
- Endpoint mode is valid (§5.5) for most types (else transient mode, whose cost scales with developmental duration over the regulatory timescale).
- Brain states are Markovian at a usable lag time (else averaging windows lengthen to cover state mixing).
- The number of positive Lyapunov exponents is modest (else long-horizon gradients use likelihood-ratio estimators with higher variance).
- Local Fisher estimates of information cost agree with coupled comparisons (else coupled comparisons become the primary error measure, at higher validation cost).
- N1 reductions stay within tolerance over the parameter domain that inference explores (else more types run at N2, or reductions are rebuilt more often, raising memory and cost).