← The living map spec / 09-roadmap

9. Roadmap and milestones

Organism phases

Phase Organism Why Key data
P0 C. elegans hermaphrodite Best ground truth: complete connectomes across development (Witvliet et al. 2021), whole-animal (Cook et al. 2019), transcriptome per class (CeNGEN), neuropeptide connectome, ~23k-pair causal atlas, whole-brain imaging with identified neurons (NeuroPAL), existing embodied baselines (BAAIWorm) WormWiring, CeNGEN, Randi 2023, Atanas 2023
P0.5 Drosophila larva Stepping stone from 302 to ~10⁵ neurons: complete larval brain connectome (~3,000 neurons, ~548k synapses), insect channels and transmitters, transparent body for whole-CNS imaging, rich behavior. Lets the fly toolchain (cell types, insect channel library, H2) be validated at a scale where reference fidelity is still cheap Larval brain connectome (Winding et al. 2023), larval nerve cord data, larval imaging and behavior datasets
P1 Drosophila ventral nerve cord + body Motor control with a body; nerve cord connectomes (MANC, FANC, BANC); CPG modeling (Pugliese et al. 2025) MANC, BANC, flybody
P2 Drosophila full CNS FlyWire (female brain), BANC (Nov 2025), male CNS (Sept 2026: 166.7k neurons, ~125M synapses); ~8k cell types make per-type programs feasible FlyWire, neuPrint, Fly Cell Atlas
P3 Larval zebrafish Vertebrate; same-animal imaging + EM (CLEM) allows test E6 Correlative datasets (2025)
P4 Mouse cortical volume (MICrONS) Mammalian test of H2/H3 on a partial, open system with functional co-registration MICrONS

P0 exit criteria:

P1 (full fly emulation) does not start until the decision rule has been applied. If H* is undecided, P1 proceeds with the model chosen by predictive selection, as the rule's consequence table specifies. Track B analyses (below) use fly data earlier, because they test H2 and H3, not H*.

Tracks and early results

A project this large needs publishable results long before it can emulate a whole fly. The work runs as three parallel tracks, each producing results that stand on their own whatever happens to H*:

Track Question First result Needs
A. Regulatory Closure (worm) Does H* hold? A0 demonstrator (synthetic): equilibrium selection, gradient validity and identifiability of regulation, before any organism model. Then the M3 identifiability gate (year 1): which properties of a nervous system its existing data can pin down at all A0: none; M3: worm data, M0–M2b
B. Electrotonic anchoring (fly) Does morphology predict effective connectivity better than synapse counts (H2, F5)? B1 (year 1): passive-cable transfer impedances for FlyWire neurons. B1 measures headroom (how much location-weighting changes effective weights); it selects F5's test and control types. F5 itself then compares the two weightings in predicting measured transfer, on a named dataset meeting F5's resolution requirements FlyWire morphologies + fly functional data with sufficient temporal resolution; L1 only
C. Invisible connectome (worm) Do model residuals reveal gap-junction and peptide links (H3, F6)? M5: a ranked list of predicted extrasynaptic links for wet-lab testing Track A model at M2 or later

Track B needs only L1 physics and existing data, so it is the cheapest route to a first paper and to early validation of the reduction method (06-compute.md §3).

Status (2026-10-07): B1 prototype in prototypes/b1-electrotonic, run on 3,397 neurons of 16 cell types and a stratified sample of 600 neurons across FlyWire's 10 super classes, with the full synapse table. For slow signals, synapse counts are an adequate proxy almost everywhere (no super class median above 0.074 of input weight moved). With spike-initiation-zone targets (validated in projection neurons against the known site where the axon leaves the antennal lobe), spiking neurons stay near counts even at 100 Hz (projection neurons, descending neurons, MBON01: ≤ 0.08). Headroom is concentrated in fast signaling by central-brain and visual centrifugal neurons (about 60% move more than 10% of input weight at 100 Hz) and in local routing by graded neurons with large arbors (APL, all 44,678 output partners: non-separability 0.40 slow, 0.58 at 100 Hz). Results use default cable parameters; class rankings are more reliable than absolute values. Next: sensitivity at scale, checked transmission-mode labels, and B1b on pre-registered high- and low-headroom types.

A0 mathematical precursor (2026-10-07): Regulatory closure demonstrator, with reproducible numerical evidence. A synthetic three-neuron circuit with body feedback tests direct development against branch-constrained equilibrium solving, implicit and transient gradients, rearing-history dependence, and recovery of regulation rates from noisy perturbation trajectories. Sensor equality leaves three neutral directions; an explicit branch constraint makes the endpoint solve well defined. A local history-dependent rule can reach different adult conductances at identical activity targets. Equilibrium observations leave the regulation rate unidentified; transient recovery identifies it in the tested four-parameter problem. This is a precursor to M3, not an organism-level H* result or completion of M3.

Data to commission

H* can't be decided from public data alone (see "What H* claims, and its rivals"). Partner labs are asked for these datasets, in priority order:

  1. Timed channel depletion in adult worms (auxin-inducible degron alleles of K⁺ and Ca²⁺ channels, fluorescently tagged so the depletion time course is measured), with whole-brain imaging and pair stimulation at several time points over hours. Tests F3 and identifies regulation rates.
  2. Chronic activity manipulation during development (e.g. HisCl1 silencing of identified neurons, rearing at different temperatures), with pair-response mapping both during the manipulation and after it ends and a declared recovery period. Tests F7a and F7b.
  3. Same-animal whole-brain imaging and EM connectome in the worm, for E6 and for separating individual wiring variability from model error.
  4. Electrophysiology of under-characterized worm neuron classes, chosen by the M3 Fisher analysis as the most informative.
  5. Within-type single-cell expression with technical-noise controls (several cells per class, from more than one animal), for the F4 covariance statistics where existing atlas data are underpowered.

Each dataset is specified with the active-learning tools (05-inference.md), so experiments target the parameters that matter. Each request quotes the sample sizes that the power analysis (02-hypotheses.md) requires for the tests it serves.

First milestones (P0, C. elegans)

  1. M0, Data and benchmark freeze:

    • Ingest Cook 2019 + Witvliet 2021 connectomes, morphologies, CeNGEN, peptide–GPCR maps, Randi 2023 atlas, Atanas 2023 whole-brain recordings.
    • Freeze held-out splits along every generalization axis, and the task specifications (loss, null, ceiling, interval, margin, threshold) as hashed artifacts (03-evaluation.md).
    • Implement noise ceilings and continuous scoring.
    • Pre-register the falsification margins, test and control cell types, recovery periods and minimal effects of interest (02-hypotheses.md).
    • Freeze each task's information budget τ_T (12-information-budget.md), which sets the tolerances of every reduction, solve and sampling choice scored on it.
    • Run the leakage audit.
    • Train the connectome-free and functional-atlas baselines, so every later result has a comparison.
  2. A0, Regulation demonstrator (Track A; needs no worm data, runs alongside M0–M2): a few regulated neurons in a reduced closed loop with known initial states, developmental histories and perturbations. It must show:

    • endpoint mode reproduces directly simulated developmental endpoints, and rejects roots that are unstable or not reached from the initial state;
    • implicit gradients agree with finite differences of simulated endpoints, and their error tracks the conditioning of the loop gain;
    • the predicted within-type covariance matches the spread of endpoints over initial states;
    • lasting history effects appear only through the mechanisms L7 lists (e.g. absent for constant-gain rules, present for non-involutive ones);
    • which program parameters equilibrium data identify, and which need transients.

    A synthetic instance covers all five items (prototypes/a0-regulatory, including the integral-control checks in evidence/slice). Still open: sensors as time averages of stochastic activity, and a measured single-cell example.

  3. M1, Reference single cell: L1 + L2 + L6 for ~10 well-characterized neuron classes (e.g. AWA, AVA, RMD, AIY). Pass E1. Also build the cost profiler: roofline measurements for the tree solve, gating and synaptic scatter kernels, plus the temporal-blocking and replica-batching speedups (06-compute.md §4), measured on reduced fly neurons taken from FlyWire morphologies.

  4. M2, Wired network: L3 + L4 (worm gap junctions from EM) with both F1 direct-fit baselines (per-neuron hierarchical, per-type). Score E2/E3. Measure the Lyapunov spectrum and the imaging channel capacity of each dataset design, and apply the data-rate gate (12-information-budget.md §6) to decide which E3 tasks are scored as trajectories and which only as distributions.

  5. M2b, Minimal regulation and reduced closed loop: the components M3 and M4 depend on, built minimally:

    • a regulation simulator implementing the canonical form of L7 with endpoint mode, transient mode and certification (06-compute.md §5), validated against direct simulation on M2 circuits;
    • a reduced worm body (B1 level: resistive force theory) with proprioceptive feedback, closing the loop for burn-in. B0 replay remains the body inside the gradient loop.
  6. M3, Identifiability gate (go/no-go for H*): before building the full regulation pipeline, check whether Φ can be recovered at all, using the M2b simulator to generate and re-infer programs.

    • Synthetic recovery: draw ground-truth Φ from the prior, generate data matching what real experiments provide (same neurons imaged, same perturbations, same durations and noise), then re-infer Φ. Pass if the parameters that matter for predictions are recovered (held-out prediction error within the noise ceiling), even if sloppy directions are not.
    • Local analysis: Fisher information spectrum of Φ at the prior mode. Report how many directions are constrained by the real dataset and which ones each experiment type constrains, separately for equilibrium-identifiable parameters (set points, gain subspace, initial-state distribution) and transient-identifiable ones (gain rates, leak) (05-inference.md).
    • Model recovery and power: the model-recovery matrix and the power of every falsification test on the real design (02-hypotheses.md). These set the sizes of the commissioned datasets.
    • Outcome: go → M4. No-go → reparameterize Φ or add the specific experiments the analysis says are missing (fed into the active-learning queue), then re-run the gate.
    • Cheap at worm scale, and it is the fastest way to keep or kill H*.
  7. M4, Regulatory Closure: first pass the AFD positive control (03-evaluation.md, Controls). Then L7 + burn-in in the reduced closed loop of M2b along sampled developmental histories (both rule families), with the rearing-sensitivity analysis. Record which types needed transient mode. Run F1 (fair-baseline protocol), F2, F4. Score all baselines.

  8. M5, Volume transmission + residual cartography: L5. Run the F6 specificity check, then F6. Publish the predicted wireless edges. (F5 runs in Track B.)

  9. M6, Full embodiment: L8 worm continuum body + environments, replacing the reduced body for evaluation. Re-score E2–E4 in the full closed loop and report the change from M2b's reduced body. Score E4.

  10. M7, Perturbations: score E5, F3, F7a and F7b (needs commissioned datasets 1–2). Publish the Necessity Map. Apply the decision rule and decide how to proceed to P1.