Decision log
Each entry records a design decision, why it was made, what was rejected, and what would make us revisit it. The other spec documents describe only the current design; this log explains how it got there. New entries for interface, file-format or benchmark-protocol changes go through a CAKE Enhancement Proposal (08-community.md).
Spec versions
| Version | Date | Change |
|---|---|---|
| v0.1 | 2026-10-07 | Initial spec: hypotheses H*, H2, H3; physics layers; Emulation Ladder; roadmap |
| v0.2 | 2026-10-07 | Computational review: D4–D6 |
| v0.3 | 2026-10-07 | Identifiability gate, rearing history, F1 fair-baseline protocol: D7, D8 |
| v0.4 | 2026-10-07 | Structural speedups: D9–D11 |
| v0.5 | 2026-10-07 | Platform plan modeled on large scientific codes: D12, D13 |
| v0.6 | 2026-10-07 | Spec split into documents; consolidated compute design; D14–D20 |
| v0.7 | 2026-10-07 | Scientific and strategic review: D21–D26 |
| v0.8 | 2026-10-07 | Evaluation controls, missing biology, novelty check: D27–D35 |
| v0.9 | 2026-10-07 | First Track B1 results; F5 refined: D36 |
| v0.10 | 2026-10-07 | Mathematical review: regulation as integral control with explicit equilibrium selection; three-way test outcomes; executable benchmark; milestone reordering; history-dependent prediction and forecast contract: D37–D49 |
| v0.11 | 2026-10-07 | Information budget: one observation-space tolerance for reductions, solves, sampling and messages; Fisher-weighted allocation; sensor-targeted Markov state models; data-rate gate: D50–D54 |
Decisions
D1. Start with C. elegans; the fly waits for a verdict on H*.
- Why: best ground truth (causal atlas, whole-brain imaging, transcriptome, multiple connectomes), and cheap enough to run every approximation against an unapproximated reference.
- Rejected: starting with the fly because its connectome is newer and larger.
- Revisit if: worm data turn out insufficient to decide H*.
D2. Fit per-type regulatory programs, not per-neuron conductances (H*).
- Why: shrinks unknowns by orders of magnitude and makes slow compensation predictable.
- Rejected: direct per-neuron fitting (kept as the F1 baseline).
- Revisit if: F1–F4 reject H* or the M3 gate fails after reparameterization.
D3. Compare against raw fluorescence; the indicator is part of the model.
- Why: deconvolution adds its own errors and hides indicator buffering of Ca²⁺.
- Revisit if: never for calcium data; voltage imaging gets its own forward model.
D4. Solve regulation burn-in as a fixed point, not by simulating days of regulation. (v0.2)
- Why: direct burn-in has ~4×10⁹ sequential steps per candidate, which no amount of hardware shortens.
- Revisit if: fixed-point solutions disagree with direct burn-in on the worm.
- Refined in v0.10 (D37, D38): the fixed point is a shortcut for a dynamically defined outcome, valid only under stated conditions.
D5. Replay recorded body kinematics inside the gradient loop. (v0.2)
- Why: gradients through contact mechanics are unreliable.
- Rejected: smoothed contact models (change the physics).
- Revisit if: fits depend on body–brain feedback not captured by replay.
D6. Shadowing gradients for long-time statistics of chaotic dynamics. (v0.2)
- Why: backpropagation through time gives meaningless gradients for these targets.
- Revisit if: too many positive Lyapunov exponents make NILSS too expensive (fall back to likelihood-ratio estimators).
D7. Identifiability gate (M3) before building the regulation pipeline. (v0.3)
- Why: fewer parameters don't guarantee identifiable ones; the gate is the cheapest way to keep or kill H*.
D8. Rearing history is an explicit model input. (v0.3)
- Why: under H*, parameters depend on experienced activity during development.
D9. Whole brain per device; replicas are the unit of parallelism. (v0.4)
- Why: a reduced fly brain is ~1.5–2 GB of state; partitioning one brain across devices adds communication for no benefit.
- Rejected: domain-decomposed brains as the default (kept for reference-fidelity runs).
- Revisit if: production fidelity grows beyond one device's memory.
D10. Per-neuron regulation decomposition, with the joint solver as fallback. (v0.4)
- Why: turns most burn-in work into independent single-cell problems and makes the fixed-point Jacobian block-structured.
- Revisit if: outer iterations fail to converge for a large share of worm configurations.
- Refined in v0.10 (D38): unknowns are slice coordinates (k per neuron); convergence is predicted by the spectral radius of the outer iteration.
D11. Geodesic Levenberg–Marquardt as the outer optimizer. (v0.4)
- Why: serial outer iterations set wall-clock time; first-order methods crawl on sloppy models.
D12. Organize CAKE like an Earth system model: components, coupler, configurations, intercomparison. (v0.5)
- Why: the problem has the same structure (several coupled physical domains, many contributing groups, standard experiments).
D13. A mechanism language (CKL) compiled to kernels and derivatives. (v0.5)
- Why: biologists contribute mechanisms without writing GPU code; one source generates all backends and their gradients, so the differentiable and fast paths can't diverge.
- Rejected: hand-written kernels per mechanism; separate research (differentiable) and production (fast) simulators, which tend to drift apart.
D14. Split the spec into documents; separate current design from history. (v0.6)
- Why: the single file mixed science, physics, numerics, platform and governance, and its compute section had become a revision history with superseded budgets (and an orphaned table fragment from an editing error). Each document now has one purpose and an owning working group.
D15. A four-object core data model: Organism, Model, Params, State. (v0.6)
- Why: separating immutable structure from per-replica state makes replica batching, caching and differentiation straightforward (as
mjModel/mjDatain MuJoCo).
D16. A fidelity lattice shared by reduction validation, fidelity probes and multilevel Monte Carlo; reduction is distinct from ablation. (v0.6)
- Why: v0.1–v0.5 said reduced models waited for the Necessity Map while also using them from the start. Reductions are numerical approximations controlled by measured error; ablations test whether biology is needed. They answer different questions.
D17. Replica-batched kernels and mechanism-signature layout. (v0.6)
- Why: index and topology traffic dominates; sharing it across replicas raises arithmetic intensity, and grouping neurons by mechanism set avoids GPU warp divergence.
D18. Stratified averaging with Markov state models for regulation sensors. (v0.6)
- Why: the largest open cost uncertainty was the averaging window needed to cover slow behavioral state switching. Short windows per state, weighted by stationary occupancy, parallelize where long windows can't (as in molecular dynamics).
- Revisit if: implied-timescale tests show the state decomposition isn't Markovian at usable lag times.
D19. Content-addressed compute cache and a declarative workflow layer. (v0.6)
- Why: at fly scale, recomputing burn-ins, reductions and benchmarks across experiments and people would waste more compute than any single kernel optimization saves. Bit-reproducibility makes exact caching safe.
- Refined in v0.10 (D47): exactness holds within a reproducibility class; keys include initial-state and schedule samples; warm starts need certification.
D20. Latent edges as continuous variables with sparsity priors. (v0.6)
- Why: simulation-based inference over discrete edge sets scales poorly; continuous relaxation lets gradients do most of the work, with discrete model selection only for a contested shortlist.
D21. H* is defined by locality and activity dependence, not by one rule form. (v0.7)
- Why: a single assumed rule (calcium sensors) could fail while the hypothesis is true. A learned local rule family tests the claim itself.
D22. Named rival hypotheses (R1 genetic hardwiring, R2 idiosyncrasy) and a pre-registered decision rule. (v0.7)
- Why: acute data can't separate H* from R1, so F1 alone could not decide H*; and without a fixed rule, mixed test outcomes invite post-hoc interpretation.
D23. Commission chronic-manipulation and timed-depletion datasets (new F7; revised F3). (v0.7)
- Why: the decisive evidence for or against H* needs changed activity histories, which public data don't contain.
D24. Test H2 mainly in the fly, early, as an independent track. (v0.7)
- Why: most worm neurons are electrotonically compact, so transfer impedance barely differs from synapse counts there. The fly test needs only passive cable physics and existing data.
D25. Adopt existing tools first; extract the platform from working code. (v0.7)
- Why: building the compiler, coupler and kernels before the science fixes requirements too early, and delays every result by years.
- Rejected: platform-first build (v0.5 timeline).
D26. A focused core team, with the open consortium growing around it. (v0.7)
- Why: the first years are tightly coupled work that open consortia do slowly; consortia are better once interfaces are stable.
D27. Strong non-mechanistic baselines on every level. (v0.8)
- Why: without connectome-free and empirical baselines, a passed level can't show that anatomy and physics were needed. CAKE's distinctive claim is generalization to unseen perturbations, so that is where it must win.
D28. A positive control for regulation (AFD temperature set point) before testing H*. (v0.8)
- Why: a failed H* test is only informative if the machinery can reproduce a known case of experience-dependent set-point regulation.
D29. Continuous scoring normalized by the noise ceiling; pass/fail by pre-registered thresholds. (v0.8)
- Why: all-or-nothing levels hide progress and make baseline comparisons impossible.
D30. Burn-in along the developmental connectome series, not on the adult connectome alone. (v0.8)
- Why: under H*, regulation tracks wiring as it develops; the worm's developmental connectomes make this possible and add a stage-by-stage test.
D31. Associative plasticity is part of the model; learned behavior is part of E4. (v0.8)
- Why: an emulator with fixed synapses can't reproduce one of a nervous system's core functions. Homeostasis and learning are different processes and are modeled separately.
D32. Leakage audit between priors and validation data. (v0.8)
- Why: priors from expression and peptide maps and validation from functional data come from overlapping literatures; undetected overlap would inflate results.
D33. Regulation rules have fast and slow paths. (v0.8)
- Why: recent work shows immediate compensation through voltage-dependence shifts and persistent adaptation through density changes; a single slow path would mispredict F3 time courses.
D34. Set points may depend on neuromodulatory state. (v0.8)
- Why: homeostasis and neuromodulation cooperate; fixed set points would make regulation fight modulation (e.g. across fed and starved states).
D35. Add the Drosophila larva (P0.5) between worm and adult fly. (v0.8)
- Why: the jump from 302 to ~1.7×10⁵ neurons is too large to debug in one step. The larval brain (~3,000 neurons) tests the insect toolchain at a scale where reference fidelity is still cheap.
D36. Test H2 on fast responses in high-headroom cell types, with pathway-level weights and spike-initiation-zone targets. (v0.9)
- Why: the B1 panel run (26 neurons, 16 types, full FlyWire synapse table) found that for slow signals effective weights barely differ from synapse counts in most types, so a brain-wide, slow-signal test of H2 would be uninformative. Headroom is large and type-specific at 100 Hz. Averaging transfer over all outputs hides local computation (APL), and passive transfer to distant outputs overstates attenuation in spiking neurons.
- Revisit if: scaling up the panel shows large slow-signal headroom in cell types not yet sampled.
D37. The developmental outcome is defined by regulatory dynamics with an explicit equilibrium-selection rule. (v0.10)
- Why: the v0.9 burn-in equation
E[sensor(θ)] = setpointhas a manifold of solutions per neuron when sensors are fewer than regulated quantities, so it did not define what H* predicts. Its ordinary Jacobian is singular there, and the claim that warm starts cannot change the answer was false under multistability. Writing every rule in a canonical form (gains, leak, sensors, initial-state distribution, developmental schedule) makes the outcome well defined. For constant-gain integral control, the integrated error confines each neuron to the slice through its initial state, which selects the equilibrium. Path dependence then requires named mechanisms (multistability on the slice, non-involutive gains, higher-order error terms, saturation, plasticity), which gives F7 a sharper form (F7a/F7b). - Rejected: least-norm or pseudoinverse solutions of the sensor equation (they select by an arbitrary metric, not by the biology); treating the manifold of solutions as unresolvable degeneracy.
- Revisit if: the learned rule family fits only in forms with no conservation structure, so endpoint mode rarely applies (then transient mode becomes the default, D38).
D38. Endpoint mode (the slice equation) with stated validity conditions; transient mode otherwise. (v0.10)
- Why: solving the k-dimensional slice equation per neuron is better posed (its square form is nonsingular exactly when the loop gain J G is) and cheaper than solving for n parameters. Validity conditions V1–V8 say when implicit differentiation is valid; certification (residual, loop-gain stability, basin check) catches roots that are unstable or not reached from the initial state. A synthetic check during this review found a root of the sensor equations at which the dynamics diverge.
- Rejected: unconditional implicit differentiation through any root; Newton without continuation (it converges to arbitrary branches).
- Revisit if: a large share of worm types need transient mode (budget re-planned, §10 of 06-compute.md).
D39. Regulation predicts distributions: initial states are random, and their spread is part of Φ. (v0.10)
- Why: within-type and between-animal variability then become predictions, with a testable covariance structure: P Σ₀ Pᵀ, and controlled sensor combinations varying only as much as set points and inputs do. This turns F4 from "better than a shuffled null" into tests of specific statistics.
- Rejected: a single deterministic θ per neuron, which discards the main evidence that distinguishes H* from R2.
D40. Regulation gains are parameterized as subspace × rates; equilibrium data identify only the subspace. (v0.10)
- Why: in integral control every endpoint is invariant under G → G C, including a common rescaling of all rates. Making this gauge explicit stops M3 from reporting a structural fact as a fitting failure, and states which data (transients) identify rates.
D41. Three-way test outcomes, power requirements, and a decision rule that separates prediction from mechanism. (v0.10)
- Why: in v0.9 a per-neuron model winning was read as idiosyncrasy, and failed compensation as hardwiring, although missing mechanisms, weak interventions, insufficient measurement or optimization failure could explain either. There was no inconclusive outcome and no outcome for F1 losing only to the per-type baseline. Now: validity gates, a parsimony-based predictive selection, mechanistic claims that need their alternatives excluded, explicit handling of conflicts, and H*, R1 and R2 as limits of one nested family with interpretable estimates (κ, ρ, ι).
- Rejected: a single categorical verdict table; Bayes factors alone (they are sensitive to priors on rule parameters that nobody can justify yet).
D42. Executable benchmark: one normalized score, per-task losses, generalization axes, frozen specifications. (v0.10)
- Why: "performance divided by noise ceiling" was not defined for signs, kernels, distributions and mutant effects, and correlation-only scoring accepts wrong amplitudes and timings. The leakage rule also forbade a legitimate task (unseen pairs among observed neurons). Each task now has a loss, null, ceiling, interval and margin, scored per generalization axis, frozen and hashed at M0.
D43. Milestones reordered: A0 demonstrator and M2b (minimal regulation simulator, reduced closed loop) before M3; P0 exit includes E5 and F7. (v0.10)
- Why: M3 had to generate and recover regulatory programs before a regulation pipeline existed, and M4 needed closed-loop burn-in before embodiment (M6). The P0 exit checklist omitted F7 and E5, which M7 uses to decide progression.
D44. Synapse lumping exact only under stated conditions; N1 reductions certified over a parameter domain. (v0.10)
- Why: shared presynaptic neuron, receptor type and postsynaptic compartment do not make aggregate dynamics close when release histories, local Ca²⁺ or transmitter waveforms differ, and summed binomials are binomial only with a common release probability. Transfer impedances depend on conductances that inference and regulation change, so computing them once per release was inconsistent.
D45. Transcriptome presence as probabilistic priors; developmental connectomes through a probabilistic wiring model. (v0.10)
- Why: transcriptome thresholds have false negatives (CeNGEN sets sub-threshold expression to zero), so hard filtering can remove real mechanisms. The developmental connectome series comes from different animals, so a single interpolated path mixes age with individual variation.
D46. H3 by structured residual decomposition with a synthetic specificity check before F6. (v0.10)
- Why: structured residuals also arise from channel, stimulation, observation and morphology errors. Giving each source its algebraic signature and checking coherence makes attribution testable, and the planted-error check bounds the false-discovery rate before any enrichment claim.
D47. Statistical vs. bitwise reproducibility; error control propagated to reported quantities. (v0.10)
- Why: counter-based RNG reproduces random draws, not floating-point trajectories across batching, partitioning and hardware. Shadow neurons measure local error under supplied inputs, not its effect through recurrence and the closed loop. Adjoint-weighted propagation and coupled comparisons close that gap.
D48. Compute numbers are feasibility hypotheses with tests and re-plan criteria; H2 headroom is distinguished from validation. (v0.10)
- Why: the spec called its numbers estimates but stated stronger guarantees elsewhere. Each number now has the milestone that tests it, measured on complete inference workloads, and the result that would change the plan. Likewise, B1's changes in effective weights show where H2 could matter, not that it predicts better; F5 requires a named dataset with sufficient bandwidth, transmission-mode labels and an observation model, and pre-registers its test and control types from B1.
D49. Forecast a joint distribution over programs and history-dependent states, with an explicit observation cutoff and horizon. (v0.10)
- Why: A0's constructed local rule reaches different conductances at the same activity targets after different rearing histories. A compact branch coordinate selects its endpoint, but A0 obtains that coordinate from direct simulation rather than inferring it from recordings. The next question is which history summaries are both observable and sufficient for intervention forecasts. Scientific parallels, the posterior target, proposed tests, calibration requirements and evidence limits are documented in 11-prediction-and-history.md.
- Rejected: treating the per-type program or present activity alone as a complete predictive state; presenting retrospective assimilation as a forecast; inferring neural computational irreducibility from Turing/Rice results; treating broad predictive intervals as success without a proper score.
- Revisit if: discarded history predicts held-out residuals after conditioning on the proposed state, or calibration fails under the declared interventions and horizons. Expand the state or narrow its validated domain; do not change a frozen benchmark retrospectively.
D50. Every numerical tolerance is a share of a per-task information budget, measured in observation space. (v0.11)
- Why: reduction tolerances, solve residuals, averaging lengths and replica counts had no common basis, so they could not be traded against each other or tied to what a result can show. Δ_T = KL between the dataset distributions of the exact and approximate models bounds, by Pinsker's inequality, the change in any test's rejection probability by √(Δ_T/2). It needs only the dataset design and observation model, so it is leakage-free for held-out tasks, and it uses the same units as the log score (12-information-budget.md §1–2).
- Rejected: fixed relative tolerances per component (they ignore whether an error is visible in the data); tolerances on the normalized score S alone (they hide which approximation consumed the margin).
- Revisit if: local Fisher estimates of Δ disagree with coupled comparisons by more than a factor of 2 on the worm.
D51. Compute is allocated by Fisher-weighted Neyman allocation; the regulation solve uses adaptive precision. (v0.11)
- Why: the minimum cost of meeting a noise budget is (Σ √(c_j λ_j s_j))² / 2τ, never more than uniform allocation and much less when the Fisher weights are heavy-tailed, as expected when most fly neurons are not imaged. In pseudo-transient continuation, holding Monte Carlo error at a fixed fraction of the residual makes total cost a small multiple of the final iteration's (12-information-budget.md §3–4).
- Rejected: Fisher weighting of the continuation path or of the stability and basin checks: directions the data cannot see can still carry the solve to another branch.
- Revisit if: the measured saving is below 2× at M4, or basin-check failures rise under adaptive precision.
D52. Markov state models for stratified averaging target Fisher-weighted sensor futures. (v0.11)
- Why: the states exist only to estimate sensor averages and their sensitivities. A predictive information bottleneck on those targets keeps only distinctions that matter, and the minimal sufficient statistic of the past is Markov by construction, so Markovianity becomes a measurable approximation error rather than a hope (12-information-budget.md §5).
- Rejected: clustering general activity and testing Markovianity afterwards (kept as the fallback).
- Revisit if: sensor-targeted states need as many states as activity clustering, or fail implied-timescale tests more often.
D53. A data-rate gate decides, before fitting, which tasks can be trajectory-level. (v0.11)
- Why: bounded-error tracking of a chaotic trajectory requires observation capacity at least the entropy production of the dynamics (≥ Σλ⁺ under an SRB measure). If the imaging capacity is below that, trajectory losses are ill posed for any estimator, and fitting them wastes compute and produces meaningless gradients. The same Lyapunov measurement sets shooting windows, T_w ≤ ln κ_grad / λ_max (12-information-budget.md §6).
- Rejected: an open-ended curriculum of window lengths; treating trajectory failure as an optimization problem when it is an information limit.
- Revisit if: trajectory fitting succeeds on tasks the gate fails (which would indicate an error in the capacity or entropy estimate).
D54. Messages between neurons are bandlimited and quantized at the noise floor; state stays fp32. (v0.11)
- Why: graded transmission is low-pass and noisy, so per-step waveforms carry far less information than their size. Compressing the waveforms exchanged at window boundaries and between devices reduces buffer and communication traffic without reducing state precision (12-information-budget.md §7).
- Rejected: reduced-precision state (conflicts with conservation-critical accumulators, §4.6 of 06-compute.md).
- Revisit if: the Δ of compressed messages exceeds its share for classes with strongly nonlinear release.