3. Evaluation: what "solved" means
The model is evaluated against a fixed, versioned benchmark. Each task is scored continuously with a normalized score S, whose ceiling is set by variability between animals and between trials. A level counts as passed when the lower bound of S's interval reaches a pre-registered threshold on held-out data. Continuous scores show progress and allow comparison with baselines even before a level is passed. Every task has a frozen specification (below), so a score means the same thing for every model and every release.
| Level | Test | Example data (C. elegans → Drosophila) |
|---|---|---|
| E1 Cell | Reproduce held-out single-neuron electrophysiology per cell type (I–V curves, plateaus, graded or spiking responses) | Patch-clamp literature; fly whole-cell recordings |
| E2 Causal pairs | Predict sign, amplitude and kernel shape of responses to optogenetic stimulation of neuron i in neuron j | Signal propagation atlas (Randi et al. 2023, ~23k pairs); fly optogenetic + imaging datasets |
| E3 Spontaneous state | Predict whole-brain spontaneous activity statistics, manifold geometry and state transitions | Kato et al. 2015; Atanas et al. 2023; fly whole-brain spontaneous imaging |
| E4 Behavior | Closed-loop behavior distributions match (postures, eigenworms, ethograms, gait, chemotaxis/thermotaxis indices), including learned behavior: the change in behavior after a training protocol | Worm tracking; salt and temperature associative learning in worms; fly walking/flight datasets; fly olfactory conditioning |
| E5 Perturbation generalization | Predict unseen mutants, ablations and drugs, including slow compensation | unc-7/unc-9 (gap junctions), neuropeptide mutants, ablation literature |
| E6 Individual | Given one animal's connectome plus sparse recordings, predict that animal's specific activity | Same-animal functional + EM datasets (larval zebrafish CLEM, 2025) |
"Solved" for an organism = E1–E5 passed, with E6 attempted.
Scoring
For each task t, with loss ℓ_t:
S_t = (ℓ_null − ℓ_model) / (ℓ_null − ℓ_ceiling)
- ℓ_null is the loss of a declared null predictor for the task (table below).
- ℓ_ceiling is the loss of the replicate predictor: each held-out animal (or trial) predicted by the others under the same split, corrected for the finite number of replicates.
- S = 1 means the model predicts as well as one animal predicts another; S = 0 means no better than the null. Values above 1 are possible by chance and are reported as measured.
- Interval: 95% hierarchical bootstrap over the split unit, then the nested units (animals, then neurons or pairs, then trials).
- Comparisons between two models use the interval of the difference in S, with a pre-registered equivalence margin δ_t:
- "beats": lower bound > δ_t;
- "matches": the interval lies within [−δ_t, δ_t];
- "is non-inferior": lower bound > −δ_t;
- anything else is inconclusive.
Task specifications
Correlation is never the only loss: amplitude and timing errors must count.
| Level | Unit | Primary loss | Diagnostics (reported, not scored) | Null predictor | Ceiling |
|---|---|---|---|---|---|
| E1 | (cell type, protocol) | Squared error of electrophysiological features, each standardized by its within-type across-cell SD; plus trace MSE where voltage traces exist | Per-feature errors | Mean over all types | Leave-one-cell-out within type |
| E2 | (stimulated j, responder i) | MSE between predicted and measured response traces in raw fluorescence, through the observation model, normalized by trial variance. This scores sign, amplitude, kernel shape and latency jointly | Sign (Brier score of predicted sign probability); amplitude (|log ratio| for detected responses); shape (1 − correlation after amplitude normalization); latency error | Mean training response for the pair's class (or no response) | Across-animal replicate |
| E3 | Recording (animal) | Standardized distance between predicted and measured statistics: correlation matrix, log power spectra, state occupancy and transition rates, participation-ratio dimension. Each is standardized by its between-animal SD, plus an energy distance on activity windows | Each statistic separately | Phase-randomized surrogate preserving each neuron's spectrum | Between-animal |
| E4 | Animal track or trial | Energy distance on posture distributions (eigenworm coefficients, fly joint angles); KL divergence of ethogram transition matrices; absolute error of task indices (chemotaxis, thermotaxis). Learned behavior: difference-in-differences (after − before training, trained − control) | Each component | Measured distribution pooled across conditions (no condition specificity) | Between independent cohorts |
| E5 | (perturbation, readout) | Squared error of predicted vs. measured effect (perturbed − control) on E2–E4 statistics, divided by the effect's squared standard error. Compensation trajectories: the same over time points, so amplitude and timing count, plus |log| errors of κ and the half-time t½ | Trajectory correlation | No effect (perturbed = control) | Replicate cohorts of the same perturbation |
| E6 | Individual animal | E2/E3 loss on that animal. Score is the improvement of the individual model (its own connectome plus sparse recordings) over the population model | — | Population model (reference connectome) | Within-animal split-half |
Thresholds for passing and the margins δ_t are set per task at M0, with a written justification.
Generalization axes and split units
Each task is scored separately along each axis it supports. Scores are never pooled across axes, because they measure different abilities.
| Axis | Held out | May overlap with training | Leakage rule |
|---|---|---|---|
| G-pair | (j, i) pairs | Neuron identities, animals | The held-out pair, and its reciprocal for tests involving gap junctions, never appear in training |
| G-stim | Stimulated neurons | Responders, animals | No pair with a held-out source appears in training |
| G-identity | Neuron classes | Nothing involving those classes | No data of any kind involving the class, as source or responder; no prior fitted on its functional data |
| G-animal | Animals | Neuron identities, pairs | No data from held-out animals, including normalization statistics |
| G-intervention | Mutants, drugs, ablations | Everything else | No data under the intervention; no prior derived from it |
| G-environment | Environments and rearing conditions | Everything else | No data from the held-out condition |
| G-time | Later time points of perturbation trajectories | Earlier time points | Forecasting only; no smoothing across the split |
E2 prediction of unseen pairs among observed neurons (G-pair) is a legitimate task, distinct from generalizing to unseen identities (G-identity). Both are reported.
Frozen benchmark artifacts (M0)
Forecast contract. Each forecasting task also declares its observable, forecast horizons, observation cutoff, permitted conditioning data and modalities, assimilation/update policy, intervention schedule and observation model. Report skill by horizon and generalization axis. Future measured responses, state estimates smoothed using post-cutoff data and realized post-cutoff body feedback cannot enter a forecast issued at an earlier cutoff. Rolling assimilation forecasts and forecasts adapted using measurements under an intervention are distinct tasks from zero-shot intervention prediction. Comparisons use the same permitted information.
Predictive uncertainty. For probabilistic tasks, freeze a suitable proper score and report calibration with predictive interval width by horizon and condition. These assess animal-level predictive distributions; the hierarchical bootstrap interval on S assesses uncertainty in a benchmark score. Neither replaces the other. The rationale, A0 evidence and proposed history-state tests are in 11-prediction-and-history.md. These are prospective requirements, not claims that calibrated filtering or forecasting has already been implemented.
For each task and axis, M0 freezes and hashes a specification: data version, loss, null predictor, ceiling estimator (with its finite-replicate correction), interval method, δ_t, pass threshold, split unit and split seed, plus the applicable forecast contract, probabilistic score and calibration protocol above, and the task's information budget τ_T with its split between bias and noise (12-information-budget.md). The budget is computed from the dataset design and observation model only, never from held-out outcomes. The benchmark version is the hash of all specifications. A change to any of them is a new benchmark version, and results on different versions are not compared.
A key deliverable along the way is the Necessity Map. For each physical process (04-physics.md), remove it and measure how far the model drops on the ladder. This answers directly the open question "which biological details are needed for faithful emulation" raised by the State of Brain Emulation Report 2025.
Baselines
Passing a level means little unless CAKE beats simpler explanations. Every level is also scored for these baselines, trained on the same data and splits:
| Baseline | What it represents | What CAKE must show against it |
|---|---|---|
| Connectome-free data-driven models (latent dynamical models and foundation models of neural activity trained on the same recordings) | How far prediction goes with no anatomy at all | CAKE may tie on E2/E3 within the training distribution; it must win on E5 (unseen perturbations) and transfer to new conditions, or the anatomy and physics aren't earning their cost |
| The functional atlas as a model (measured pair responses used directly as a linear network, as Randi et al. 2023 did) | The best purely empirical connectivity | CAKE must predict pairs that were never measured, and nonlinear and slow effects |
| Simplified connectome models (leaky integrate-and-fire with uniform parameters; task-optimized point-neuron networks) | The current state of the art | CAKE must beat them on every level it claims, or the extra physics isn't justified |
| Direct-fit biophysical models (F1 baselines a and b) | CAKE without Regulatory Closure | As specified in F1 |
Controls
- Positive control for regulation. Before testing H* broadly, the regulation machinery must reproduce a case where experience-dependent set-point regulation is already established. In C. elegans, the temperature threshold of the AFD thermosensory neuron tracks the cultivation temperature over hours (Hedgecock & Russell 1975; Kimura et al. 2004). A regulation-closed AFD model reared at different temperatures must shift its threshold as observed. If the machinery can't reproduce a known case, a negative result for H* would be uninformative.
- Negative controls: shuffled set points and degree-preserving shuffled connectomes (F2).
- Leakage audit. Priors and validation data must be independent. Examples to check: peptide–receptor maps used as priors (L5) must not be derived from the functional data used to validate H3 (F6); the overlap rules of each generalization axis (above) hold, including for imaging sessions and normalization statistics. The audit runs at M0 and on every benchmark release.
Ablation vs. reduction
Ablation vs. reduction. The Necessity Map removes a process (e.g. no peptides, no gap junctions) to test whether the biology is needed. A reduction keeps the process but computes it more cheaply (e.g. fewer compartments). Reductions are numerical approximations, validated by measured error against higher fidelity (06-compute.md §3, §9), and don't wait for the Necessity Map. The two must not be confused: a process that can be computed cheaply is not thereby shown to be unnecessary, and vice versa.