← The living map spec / 03-evaluation

3. Evaluation: what "solved" means

The model is evaluated against a fixed, versioned benchmark. Each task is scored continuously with a normalized score S, whose ceiling is set by variability between animals and between trials. A level counts as passed when the lower bound of S's interval reaches a pre-registered threshold on held-out data. Continuous scores show progress and allow comparison with baselines even before a level is passed. Every task has a frozen specification (below), so a score means the same thing for every model and every release.

Level Test Example data (C. elegans → Drosophila)
E1 Cell Reproduce held-out single-neuron electrophysiology per cell type (I–V curves, plateaus, graded or spiking responses) Patch-clamp literature; fly whole-cell recordings
E2 Causal pairs Predict sign, amplitude and kernel shape of responses to optogenetic stimulation of neuron i in neuron j Signal propagation atlas (Randi et al. 2023, ~23k pairs); fly optogenetic + imaging datasets
E3 Spontaneous state Predict whole-brain spontaneous activity statistics, manifold geometry and state transitions Kato et al. 2015; Atanas et al. 2023; fly whole-brain spontaneous imaging
E4 Behavior Closed-loop behavior distributions match (postures, eigenworms, ethograms, gait, chemotaxis/thermotaxis indices), including learned behavior: the change in behavior after a training protocol Worm tracking; salt and temperature associative learning in worms; fly walking/flight datasets; fly olfactory conditioning
E5 Perturbation generalization Predict unseen mutants, ablations and drugs, including slow compensation unc-7/unc-9 (gap junctions), neuropeptide mutants, ablation literature
E6 Individual Given one animal's connectome plus sparse recordings, predict that animal's specific activity Same-animal functional + EM datasets (larval zebrafish CLEM, 2025)

"Solved" for an organism = E1–E5 passed, with E6 attempted.

Scoring

For each task t, with loss ℓ_t:

S_t = (ℓ_null − ℓ_model) / (ℓ_null − ℓ_ceiling)

Task specifications

Correlation is never the only loss: amplitude and timing errors must count.

Level Unit Primary loss Diagnostics (reported, not scored) Null predictor Ceiling
E1 (cell type, protocol) Squared error of electrophysiological features, each standardized by its within-type across-cell SD; plus trace MSE where voltage traces exist Per-feature errors Mean over all types Leave-one-cell-out within type
E2 (stimulated j, responder i) MSE between predicted and measured response traces in raw fluorescence, through the observation model, normalized by trial variance. This scores sign, amplitude, kernel shape and latency jointly Sign (Brier score of predicted sign probability); amplitude (|log ratio| for detected responses); shape (1 − correlation after amplitude normalization); latency error Mean training response for the pair's class (or no response) Across-animal replicate
E3 Recording (animal) Standardized distance between predicted and measured statistics: correlation matrix, log power spectra, state occupancy and transition rates, participation-ratio dimension. Each is standardized by its between-animal SD, plus an energy distance on activity windows Each statistic separately Phase-randomized surrogate preserving each neuron's spectrum Between-animal
E4 Animal track or trial Energy distance on posture distributions (eigenworm coefficients, fly joint angles); KL divergence of ethogram transition matrices; absolute error of task indices (chemotaxis, thermotaxis). Learned behavior: difference-in-differences (after − before training, trained − control) Each component Measured distribution pooled across conditions (no condition specificity) Between independent cohorts
E5 (perturbation, readout) Squared error of predicted vs. measured effect (perturbed − control) on E2–E4 statistics, divided by the effect's squared standard error. Compensation trajectories: the same over time points, so amplitude and timing count, plus |log| errors of κ and the half-time t½ Trajectory correlation No effect (perturbed = control) Replicate cohorts of the same perturbation
E6 Individual animal E2/E3 loss on that animal. Score is the improvement of the individual model (its own connectome plus sparse recordings) over the population model — Population model (reference connectome) Within-animal split-half

Thresholds for passing and the margins δ_t are set per task at M0, with a written justification.

Generalization axes and split units

Each task is scored separately along each axis it supports. Scores are never pooled across axes, because they measure different abilities.

Axis Held out May overlap with training Leakage rule
G-pair (j, i) pairs Neuron identities, animals The held-out pair, and its reciprocal for tests involving gap junctions, never appear in training
G-stim Stimulated neurons Responders, animals No pair with a held-out source appears in training
G-identity Neuron classes Nothing involving those classes No data of any kind involving the class, as source or responder; no prior fitted on its functional data
G-animal Animals Neuron identities, pairs No data from held-out animals, including normalization statistics
G-intervention Mutants, drugs, ablations Everything else No data under the intervention; no prior derived from it
G-environment Environments and rearing conditions Everything else No data from the held-out condition
G-time Later time points of perturbation trajectories Earlier time points Forecasting only; no smoothing across the split

E2 prediction of unseen pairs among observed neurons (G-pair) is a legitimate task, distinct from generalizing to unseen identities (G-identity). Both are reported.

Frozen benchmark artifacts (M0)

Forecast contract. Each forecasting task also declares its observable, forecast horizons, observation cutoff, permitted conditioning data and modalities, assimilation/update policy, intervention schedule and observation model. Report skill by horizon and generalization axis. Future measured responses, state estimates smoothed using post-cutoff data and realized post-cutoff body feedback cannot enter a forecast issued at an earlier cutoff. Rolling assimilation forecasts and forecasts adapted using measurements under an intervention are distinct tasks from zero-shot intervention prediction. Comparisons use the same permitted information.

Predictive uncertainty. For probabilistic tasks, freeze a suitable proper score and report calibration with predictive interval width by horizon and condition. These assess animal-level predictive distributions; the hierarchical bootstrap interval on S assesses uncertainty in a benchmark score. Neither replaces the other. The rationale, A0 evidence and proposed history-state tests are in 11-prediction-and-history.md. These are prospective requirements, not claims that calibrated filtering or forecasting has already been implemented.

For each task and axis, M0 freezes and hashes a specification: data version, loss, null predictor, ceiling estimator (with its finite-replicate correction), interval method, δ_t, pass threshold, split unit and split seed, plus the applicable forecast contract, probabilistic score and calibration protocol above, and the task's information budget τ_T with its split between bias and noise (12-information-budget.md). The budget is computed from the dataset design and observation model only, never from held-out outcomes. The benchmark version is the hash of all specifications. A change to any of them is a new benchmark version, and results on different versions are not compared.

A key deliverable along the way is the Necessity Map. For each physical process (04-physics.md), remove it and measure how far the model drops on the ladder. This answers directly the open question "which biological details are needed for faithful emulation" raised by the State of Brain Emulation Report 2025.

Baselines

Passing a level means little unless CAKE beats simpler explanations. Every level is also scored for these baselines, trained on the same data and splits:

Baseline What it represents What CAKE must show against it
Connectome-free data-driven models (latent dynamical models and foundation models of neural activity trained on the same recordings) How far prediction goes with no anatomy at all CAKE may tie on E2/E3 within the training distribution; it must win on E5 (unseen perturbations) and transfer to new conditions, or the anatomy and physics aren't earning their cost
The functional atlas as a model (measured pair responses used directly as a linear network, as Randi et al. 2023 did) The best purely empirical connectivity CAKE must predict pairs that were never measured, and nonlinear and slow effects
Simplified connectome models (leaky integrate-and-fire with uniform parameters; task-optimized point-neuron networks) The current state of the art CAKE must beat them on every level it claims, or the extra physics isn't justified
Direct-fit biophysical models (F1 baselines a and b) CAKE without Regulatory Closure As specified in F1

Controls

Ablation vs. reduction

Ablation vs. reduction. The Necessity Map removes a process (e.g. no peptides, no gap junctions) to test whether the biology is needed. A reduction keeps the process but computes it more cheaply (e.g. fewer compartments). Reductions are numerical approximations, validated by measured error against higher fidelity (06-compute.md §3, §9), and don't wait for the Necessity Map. The two must not be confused: a process that can be computed cheaply is not thereby shown to be unnecessary, and vice versa.