11. Prediction, developmental history and limits
Status: Design rationale and proposed evaluation requirements, 2026-10-07. The A0 results below are implemented synthetic evidence. The filtering, calibration and history-compression experiments described here are requirements for later work, not completed capabilities or organism validation.
CAKE aims to predict intervention outcomes from physical rules, the animal's accumulated state, and uncertainty about both. A useful forecast specifies its observable, time horizon, conditioning observations and intervention. It can be a short response trajectory, a distribution of behaviors, or a slow compensation curve. Statistical forecasting does not relax the requirement to predict a measured causal response when that response is reproducible.
The spec already includes posterior ensembles, statistical E3/E4 tasks and explicit equilibrium selection. This document connects those choices to a further research question: what compact, measurable state summarizes the parts of developmental history that matter for future responses?
Lessons from other fields
These are methodological parallels, not evidence that H* is true or that nervous systems obey a particular universality class.
| Parallel | Useful lesson for CAKE | Boundary of the analogy |
|---|---|---|
| Weather and climate | Lorenz's 1963 convection model illustrates sensitivity to initial conditions. Forecast ensembles represent uncertainty in initial conditions and the model; assimilation uses new observations to update a state estimate. | Forecasting also improves through better observations, resolution and physics. Ensemble methods complement those improvements. Resemblance to predictive processing does not establish that a brain implements a weather-assimilation algorithm. Lorenz, ECMWF, forecast uncertainty. |
| Statistical mechanics and renormalization | Search for observables and effective descriptions that remain predictive after microscopic detail is reduced. Universality explains shared large-scale behavior in specified regimes. | It does not make every microscopic detail irrelevant to every perturbation. Neural criticality must be tested if invoked; it is not an assumption of CAKE. The Necessity Map and numerical reduction checks determine which details matter for each task. Wilson, 1983. |
| Evolution and population genetics | Model processes that generate a distribution of outcomes, with explicit inheritance, variation and environmental conditions. | The Price equation is an identity that requires additional assumptions for a dynamically sufficient forecast. Calling an organism a model of its environment is an analogy, not a replacement for those assumptions. Likewise, a regulatory constraint alone does not specify a developmental outcome. Frank's analysis of the Price equation. |
| Bayesian nonparametrics and latent causes | Infer whether observations belong to an existing context or support another latent context. A flexible latent-state model is a candidate for representing experience. | An unbounded number of categories still has specified priors and observation assumptions; it is not an unrestricted generator of scientific models. Gershman and Niv discussed structure learning in 2010. The remapping model by Sanders, Wilson and Gershman appeared in 2020 and is a computational account, not proof of a neural implementation. Gershman and Niv, Sanders et al.. |
| Computability | State the model class and prediction problem before claiming a general shortcut or a general impossibility. | Turing/Rice-style undecidability concerns universal decision procedures for arbitrary programs. It does not establish that a particular finite-horizon neural forecast requires step-by-step simulation. Computational cost, undecidability, chaotic sensitivity and statistical predictability are distinct questions. No irreducibility claim is assumed in CAKE's compute budget. Formal treatment of Rice-type results. |
| Turbulence | Statistical scaling under stated assumptions can be useful while microscopic dynamics and mathematical foundations remain difficult. | Kolmogorov's 1941 theory has a specified high-Reynolds-number regime and local-isotropy assumptions. The Navier–Stokes existence and smoothness problem is a separate mathematical question, not a synonym for every unresolved question about turbulence. Kolmogorov, 1941; translation 1991, Clay problem statement. |
The design inference from these parallels is to choose predictive quantities carefully and validate their uncertainty. None supplies a theorem that the brain is chaotic, critical, incompressible or unidentifiable.
Prediction requires rules and state
Let G denote the anatomical structure, Φ the regulatory programs, x_t the fast neural/body state, z_t the full slow regulatory/plastic state, D_≤t the permitted observations up to a cutoff t, and A a specified future intervention. A population or individual forecast has the form
p(Y_future | D_≤t, G, do(A))
= ∫ p(Y_future | x_t, z_t, Φ, G, do(A))
p(x_t, z_t, Φ | D_≤t, G) dx_t dz_t dΦ
Y includes the experiment's readout through the observation model. G is shown as known for clarity; structural, body and observation-model parameters are also marginalized when uncertain. do(A) means changing the model according to a defined intervention; this notation does not itself establish that the intervention model is correct.
The second factor represents uncertainty about both the generator and the animal it has generated so far. Similar present activity can leave different slow states plausible. Averaging their conductances into a single animal can erase differences in future response; forecasts must retain the relevant joint uncertainty and correlations.
For an unobserved animal, the state distribution comes from the developmental population model. For an observed animal, permitted recordings update it. Information about a compensation rate must come from measurements that constrain that rate; more endpoint replicas cannot repair a structural lack of information in the tested invariant rule.
What A0 establishes
The A0 demonstrator has three regulated neurons and a reduced body feedback loop. For each neuron, define
c_i = log(h_i) − rho_i log(g_i)
Here g_i and h_i are the two channel conductances, rho_i is their regulatory gain ratio, k is the shared regulation rate, and s_i is the activity sensor with target r_i. Under the invariant rule, c_i is conserved. Under the constructed local history rule,
dc_i/dt = k eta (r_i − s_i)^2
c_i(T) = c_i(0) + ∫_0^T k eta (r_i − s_i(t))^2 dt
At equilibrium, the terminal c_i plus the specified targets, rule and environment select the conductances. Three scalars summarize the history information needed for this particular endpoint calculation. They are not a sufficient state for arbitrary transient forecasts: away from equilibrium, fast states and the remaining slow coordinates are still needed.
The recorded experiment found:
- With the invariant branch specified, the endpoint solver and direct development agree to a maximum absolute state error of about 2×10⁻¹⁰.
- Different rearing inputs under the same history-dependent rule produce adult conductances differing by up to 13.2%, while adult activity targets match.
- Voltage endpoints leave the regulation rate unidentified. Recovery trajectories identify all four fitted parameters locally; rate errors were below 0.7% across three synthetic noise seeds, with the wiring, body, gain ratios, adult branch and intervention strengths known.
For the history-dependent rule, the endpoint check receives terminal c from direct simulation. This is an oracle check of the representation, not an inference of c from recordings or a shortcut for computing it. The squared-error rule is a constructed deterministic example; persistent sensor fluctuations can keep driving it. A0 establishes neither biological sufficiency of these coordinates nor calibrated posterior coverage, stochastic stationarity or general global identifiability.
Testing a compact history state
Call a proposed compressed history state ξ_t. It may replace part of z_t only over a declared family of rearing conditions, interventions, observables and horizons. It is not an extra free vector to fit independently for every animal without a generative prior; doing that would discard the parameter-sharing claim of H*.
The empirical claim is approximate predictive sufficiency: after conditioning on the retained state and ξ_t, additional details of the past should not materially improve forecasts on that family of tasks. Test it as follows:
- Generate or collect different rearing histories, including animals with matched present activity. Freeze which measurements are available to infer their latent states.
- Fit a candidate history summary on training histories. Compare models with no history state, the candidate state, and a richer history representation under matched data and compute conditions.
- Freeze the models and forecast new interventions on held-out histories and animals. Compare conditional response distributions, including latency, amplitude and slow recovery, using predeclared margins.
- Check whether discarded history still predicts residuals. If it does, enlarge the retained state or narrow the claimed domain. Agreement on spontaneous activity alone is insufficient.
Quantitatively, the residual is the conditional mutual information I(Y_future ; past | ξ_t, retained state, D_≤t). It is estimated as the held-out log-score gain of the richer history model over the candidate, which is a model-dependent estimate, not an exact value. ξ_t passes when this gain, scaled to the frozen dataset size, is within the task's information budget (12-information-budget.md §1, §5). This puts the sufficiency test in the same units as the log score and the compute tolerances.
In synthetic work, compare with a full-state oracle to distinguish a poor representation from a representation that the available observations cannot identify. In biological work that oracle is unavailable, so conclusions remain relative to the measured interventions and observations. The next A0 extension would infer c from partial noisy observations and test held-out rearing histories; it has not been implemented here.
Forecast horizons and assimilation
Forecast skill is indexed by horizon, task, intervention and generalization axis. There is no assumed universal neural prediction horizon.
| Forecast | What must be predicted | Relevant evaluation |
|---|---|---|
| Short response | Sign, amplitude, latency and waveform under a specified stimulus | E1/E2; transient portions of E5 |
| Activity and behavior over longer windows | Distributions, correlations, occupancy, transitions and task-dependent behavior | E3/E4; distributional portions of E5 |
| Adaptation | Compensation magnitude, time course and dependence on prior experience | E5, F3 and F7 |
These are observable-dependent regimes; numerical horizon boundaries are fixed in each benchmark task. The data-rate gate and the model-based predictability horizon T_pred ≈ ln(ε_tol/ε₀)/λ_max (12-information-budget.md §6) inform where trajectory scoring gives way to distributional scoring; they do not set a universal horizon. Statistical scores complement the existing trajectory tests and do not replace them.
Assimilation estimates present state; forecasting tests what follows. Each task declares a cutoff, permitted modalities, observation schedule and update policy. After that cutoff, the scored forecast receives no future measured response, latent-state estimate derived from it, or realized body/sensory feedback taken from the animal. Controlled experimental inputs specified in advance remain allowed; otherwise the model closes the body–environment loop itself.
Rolling forecasts may assimilate newly available data before issuing the next forecast, but every issued forecast is archived and scored at its original horizon. Smoothing with future observations is retrospective reconstruction and is reported separately. A forecast updated with measurements under a previously unseen intervention is an adaptation task, not zero-shot G-intervention generalization. All models in a comparison receive the same permitted information.
Uncertainty, calibration and provenance
Ensembles must represent the uncertainties relevant to a task: developmental state, programs, latent connectivity, observation noise and declared model discrepancy. Merely varying random seeds does not establish uncertainty coverage. A confidence interval on a benchmark score is also different from an animal-level predictive interval.
For each probabilistic task, freeze a suitable proper scoring rule at M0 (for example, a log score for a specified likelihood, CRPS for scalar outcomes, or an energy score for multivariate samples). Report calibration and predictive interval width together, by horizon and condition, so arbitrarily broad intervals cannot count as successful prediction. Evaluate observation-space predictions against observation-space outcomes; report latent-state uncertainty separately. Thresholds and pass criteria require the real task's replicate structure and are not supplied by A0.
The scoring-rule basis is Gneiting and Raftery (2007). Metric choice, multivariate scaling and estimators for finite ensembles must be fixed with the task; naming a proper score alone does not validate an uncertainty model.
Record the observation cutoff, conditioning-data hash, history-state definition/version, posterior or ensemble artifact, intervention, forecast horizon, observation model, assimilation policy and uncertainty sources with every forecast. Cached conditioned states additionally record their posterior provenance and particle/initial-state identities. Reusing a state from a different history or assimilation cutoff requires a new validity check, even when present activity and per-type programs match.
These requirements extend the task specifications in evaluation and the posterior model in inference. They are prospective: changing an already frozen task requires a new benchmark version. The regulation solver retains its endpoint-validity conditions; statistical prediction does not justify an arbitrary equilibrium choice or bypass gradient validation.