14. Enabling programs
Status: Proposals, 2026-10-07. None of this is current design. Each entry is a standalone computational program that CAKE, its three sister projects and other whole-nervous-system efforts could share. The sister projects, abbreviated below, are:
- MC, the molecular compiler (JAX): compiles connectome and molecular data into a simulation through compositional per-molecule rules.
- WS, WormSim (Rust + JAX): a differentiable C. elegans simulator fitted to the Randi et al. 2023 atlas and to whole-brain recordings.
- FB, fly-brain (JS/WASM/WebGPU + Bend): a spiking male-CNS fly brain (165,122 neurons) in a MuJoCo body, with a 17-assay behavioural benchmark.
Each program supplies an input the spec currently assumes, or removes a bottleneck named in 10-risks.md or 13-proposals.md. Entries give the gap, a literature novelty check, an MVP scope and a pre-registerable test with a kill condition. Adoption of any of them goes through the decision log (decisions.md).
Citations marked (verify) were found only through search snippets or were not fetched, and must be confirmed before they enter references.md.
Novelty check: method and summary
On 2026-10-07, four searches (one per group of three programs) covered papers, preprints and package indexes, with 4–10 queries per program. Most items were read at the abstract or snippet level. GitHub and PyPI were not searched directly; repositories were found only where general search returned them. "Not found" therefore means not found by these queries. A repository-level search is still owed for C1, C3 and C9.
| # | Program | Verdict | Closest prior work | Main consumer | In sister projects |
|---|---|---|---|---|---|
| C1 | Literature compiler for worm and fly physiology | Novel | NeuroElectro (mammalian, manual) | E1, E5; L2 priors; MC M5 | FB hand-coded targets; WS phenotype registry planned |
| C2 | Gating kinetics from sequence | Partial | OpenWorm DD025 (proposed, no results found) | L2; MC M5-R1 | None (no channel models in WS or FB) |
| C3 | Probabilistic signed, receptor-typed connectome | Partial (worm), novel (fly) | Fenyves et al. 2020 (worm, deterministic) | L3; MC M6-R3, K3 | WS Fenyves + CeNGEN rule and sign probe; FB graded transmitter sign |
| C4 | Gap-junction prior from contactome × innexins | Partial (worm), novel (fly) | Barabási & Barabási 2020 (worm, binary adjacency) | L4, H3; MC M4 gap head | FB: one hand-added gap (GF→TTMn); WS: gap rectification groups |
| C5 | Spatial volume-transmission kernels | Partial | Ripoll-Sánchez et al. 2023 (expression only) | L5, H3; MC §12.5 | WS well-mixed modulator pools; FB octopamine operators (negative) |
| C6 | Error-certified reducer for EM neurons | Partial | NEAT (Wybo et al. 2021) | N1 rung; B1; MC M6/M7 | FB size-scaled point neurons |
| C7 | Virtual microscope | Partial | NAOMi (mouse 2P); SINETRA (tracking only) | D3 likelihood; I2; MC M9 gauge | WS preprocessing audit; MC M9 |
| C8 | Synthetic organisms (perfect-model test beds) | Partial, close to novel | In silico zebrafish (Lueckmann et al. 2026) | M3; MC Phase 1 diagnosis | FB synthetic individuals and identification test; WS synthetic engineering checks |
| C9 | Differentiable sensory front-ends | Partial | flygym, FlyBrainLab, drosolf, mod-SenseWorm | L8, E4; MC M13 | FB flyvis eye, analytic odour/touch encoders |
| C10 | Connectome-aware experiment planner | Partial | Beiran & Litwin-Kumar 2024; Wagenmaker et al. 2024 | Active learning; commissioned data; MC M11 | FB rankExperiments; MC M11 |
| C11 | Generative model of connectome variability | Partial | Richter & Schneidman 2024; Schlegel et al. 2024 | Developmental wiring model; E6; MC Track I | WS and FB left–right reliability models |
| C12 | Within-type channel covariance atlas | Novel | Schulz, Goaillard & Marder 2006 (crab, qPCR) | F4; spec/13 smaller proposals | None |
What the sister projects have already measured
The three sister projects have run into the problems these programs address. Their results change several scopes and the priority order.
MC Phase 0 (MC spec §10.2, public worm data):
- Sign is set by chloride, which is unmeasured (K3 failed). Almost every glutamate and GABA edge is sign-ambiguous under a 5–40 mM prior on [Cl⁻]ᵢ, because targets nearly always express an anion-channel receptor. A transmitter-plus-receptor sign is not enough, so C3 must carry E_Cl as a variable.
- Protein-language-model neighbours do not recover receptor and transporter families (MC §12.1). C2 is scoped to voltage-gated channels first.
- Peptidergic heads added nothing on the causal atlas (K2), in agreement with Creamer et al. 2024, under a linear-response approximation.
- The indicator-gain proxy is correlated with expression identity (gauge audit). C7 must model per-neuron indicator expression.
- The connectome-only baseline B0 matched the linear-response compiler, and Stage I0 found no compact individual latent in single-stimulation responses.
WS (MOLECULAR-PRIORS, MOLECULAR-SIGN-PROBE, MIRROR-RELIABILITY, LEVEL0-CAPACITY-PROTOCOL, PREPROCESSING-AUDIT):
- The expression rule leaves most worm signs undecided. Fenyves transmitter tables combined with CeNGEN receptor expression, at expression threshold 2, label 3,638 chemical edges as follows: 2,743 conflicting, 257 inhibitory, 8 excitatory, and the rest incomplete or without evidence. Moving to threshold 4 changes the category of 971 edges. This reaches MC's K3 conclusion by a different route.
- Fits do not reveal whether molecular signs are right. The sign probe could detect a planted signal strong enough to move half the labelled groups toward their labels (z = +4.4), but not one that moved a quarter. The real fit gave z = −1.88 and −0.40. Seed restarts were uninformative, because Adam shrinkage moved every group alike.
- In the worm connectome, edge existence is uncertain and edge size is not. Left–right mirror presence is 32.1% at one synapse (rewired chance 9.1%) and 98.0% at 20 or more. When both sides exist, latent reliability is 0.98. Gap junctions are less reliable: both-present correlation is 0.46 and latent reliability 0.94.
- The capacity gate cannot separate undertraining from misspecification. Every optimizer and numerics variant moved training capture monotonically from 74.8% to 79.31%, never reaching the 90% gate and never stationary. Coexisting resting states were found. No planted-truth control exists, and the protocol states that the failure alone does not identify capacity as the cause.
- No raw fluorescence is available. The public Randi atlas is processed ΔF/F. The WormWideWeb traces are whole-recording z-scores, which leak future samples into early ones. No fresh confirmatory cohort has been secured, and DANDI 001075 turned out to duplicate the OSF recordings.
FB (31-ablation-ladder, 32-synapse-uncertainty, 34-individual-validation, 30-hypothesis-lab, 40-oa-operators, 10-senses):
- Sign, weights and size all matter.
- Removing sign costs 0.272 on the benchmark, and eight of nineteen ablations cost more than that.
- Binarizing weights costs 0.459.
- Removing the size scaling of point neurons costs 0.327.
- Glutamate is treated as inhibitory throughout, and 3,602 neurons without a consensus transmitter get a graded sign.
- In the fly connectome, the weights are the uncertain quantity. A left–right model of 120,000 mirror pairs puts the measurement SD at 0.64 log units at three synapses, about a factor of two. Weak edges are still real: mirror presence is 32.0% at three synapses, 45 times the null. Empirical-Bayes shrinkage moves 99.4% of weights.
- Electrical synapses are missing. The only one in the model is the hand-added GF→TTMn gap junction, and escape fails without it.
- Behaviour cannot identify individuals in simulation, but motor pools can.
- Of twelve synthetic individuals, motor-pool rates identify 12 and six behavioural observables identify 1.
- Five of the six behavioural observables have ρ ≤ 0.03.
- Readout sensitivity κ explains the failure: at κ = 0 the statistic reports nothing whatever the fit.
- Ensemble disagreement ranks noise unless it has a floor. In the hypothesis lab, the spread between ensemble members exceeded the Delta7-silencing effect being ranked.
- No octopamine operator reproduces starved hyperactivity. That is four synapse-graph and field operators, none spatial. Olfaction is the weakest sensory front end: projection neurons idle near 77 Hz against a few to 20 Hz in recordings, because ORN presynaptic gain control is replaced by a normalization.
Cross-project pattern. All four projects are stalled at the same kind of question: is a failure caused by the model, the data or the fit?
- MC's B0 tie;
- WS's capacity gate;
- FB's behavioural identification null;
- CAKE's M3 gate.
None has a known-answer control at its own scale. That makes C8 the top priority. Two other programs are cheap because the precursors already exist: WS and FB each built a left–right reliability model independently (C11), and both import transmitter-based signs (C3).
Data the field lacks
C1. Literature compiler for worm and fly physiology
Gap. E1 needs per-type electrophysiology, E5 needs quantitative perturbation effects, and L2 needs kinetic priors. For worms and flies all three are scattered across about 10⁴ papers, often only as figures. WormBase phenotypes are qualitative. MC's Phase 1 curated its kinetic family records by hand.
Program. An extraction pipeline, with an LLM doing the first pass and curators confirming it. It writes a provenance-linked database of three record types:
- Cell physiology: cell type, measured property (input resistance, τ_m, capacitance, rheobase, I–V points, plateau and spike features, channel kinetics from heterologous expression), value with dispersion and n, and conditions (temperature, saline, genotype, preparation, recording mode).
- Perturbation effects: intervention (mutant, ablation, drug, optogenetic silencing), readout (neural activity, behavior index, posture statistic), effect size with error, control and conditions.
- Assay definitions: the protocol, the quantity extracted and the scoring rule for a benchmark target, so that a model's score cannot drift away from what the source measured.
Every record points to a figure, table or sentence. Figure values go through a digitizer with recorded pixel uncertainty.
Sister projects.
- FB hand-codes its calibration targets. An audit found three benchmark terms fixed at zero by arithmetic, and one term (
tarsalPER) scoring the wrong condition: it rewarded olfactory idle activity. The repair moved the benchmark score from 0.794 to 0.697. That is why the third record type exists. - WS lists a "source-backed phenotype registry" as an unmet requirement. Its perturbation tasks are not demonstrated for lack of phenotype data.
- MC's kinetic family records were curated by hand.
C1 replaces all three hand-built tables with one shared table.
Novelty. Not found for any invertebrate.
- NeuroElectro (Tripathy et al. 2014) is mammalian and manually curated, and its authors name curation as the bottleneck.
- Yemini et al. 2013 is a single-lab behavioral phenotype database for 305 strains. It is a source to link, not a compilation.
- LLM extraction of numeric values is established in materials and protein science (verify: arXiv 2604.07584, 2605.11221), but not in neurophysiology.
MVP.
- Scope: C. elegans only, with two record sets. The first is cell physiology for the ~40 classes with any patch-clamp data. The second is perturbation effects for gap-junction (unc-7/unc-9), peptide-processing (egl-3, egl-21, unc-31) and chloride-transporter (kcc-2) mutants. Together these feed E1, the E5 axes CAKE scores first and MC's K2 and K3 queues.
- Output: a versioned Parquet table with a schema shared with MC
kinetics.json, plus an extraction audit. - Size: a few hundred papers.
Test.
- Before extraction starts, curators who do not see the pipeline's output label a random 10% of papers. The pass criterion is value-level precision ≥ 0.95 and recall ≥ 0.85, with every value's conditions correct. A value counts as correct within the digitizer's stated uncertainty.
- Usefulness test: E1 null and ceiling losses (03-evaluation.md) are computed from the database. They must agree with hand-assembled values for the M1 classes.
- Kill: precision below 0.9 after one round of prompt and schema revision. The fallback is manual curation assisted by the pipeline, with the same schema.
C2. Gating kinetics from sequence
Gap. L2 lists genes without measured kinetics as a risk and falls back on homology priors. MC M5-R1 assigns priors from curated families, because protein-language-model neighbours failed its family-recovery check.
Program. A regression from channel protein sequence (embeddings plus structure features of the voltage sensor and pore) to Hodgkin–Huxley parameters: V½ and slope of activation and inactivation, and the voltage dependence of τ. Its uncertainty must be calibrated, and it must widen with distance from the training data. Training data are characterized channels and variants: published variant electrophysiology for KCNQ1, KCNH2, SCN5A and Kv families, deep mutational scans where they report gating, and C1 records.
Novelty. Partial.
- OpenWorm DD025 proposes the same aim for C. elegans (AlphaFold 3, BioEmu, ESM; target under 30% relative error). It has status "Proposed" and no code or results were found (verify). Contact it as a possible collaborator.
- Single-gene predictors exist. Q1VarPred covers KCNQ1. Kroncke et al. regressed V½ for SCN5A and KCNQ1 with R² ≈ 0.29 (verify). A k-NN predicted Kv V½ with a mean absolute error of 7 mV.
- MissION (2025) classifies gain and loss of function and does not output parameters (verify).
No calibrated cross-family regression was found.
Sister projects. Neither WS (Level 0 graded neurons) nor FB (LIF point neurons) models ion channels yet, so MC is the only near-term consumer. FB's ablation ladder also suggests a lower priority for flies: at the behavioural level, graded weights cost more to remove than any cellular mechanism.
MVP.
- Scope: voltage-gated K⁺ and Ca²⁺ channels only. Receptors and transporters stay on curated labels, per MC §12.1.
- Outputs: per-gene parameter distributions for the worm and fly voltage-gated channelomes, and a ranked list of genes whose predictive interval is widest. That list is sent to C10 and to data request 4 (09-roadmap.md).
Test.
- Hold out whole subfamilies (leave-subfamily-out, not leave-variant-out).
- Pass: (i) the predictor beats MC's nearest-curated-family prior on the log score for held-out subfamilies; (ii) coverage of the 90% intervals is between 0.85 and 0.95.
- Downstream test: M1 E1 scores using C2 priors are non-inferior to scores using curated family priors.
- Kill: if (i) fails, C2 is reduced to a variant-effect corrector within curated families and is not used as a cross-family prior.
C3. Probabilistic signed, receptor-typed connectome
Gap.
- Whole-brain fly models assign sign from the presynaptic transmitter. Shiu et al. 2024 treat glutamate as inhibitory.
- Eckstein et al. 2024 predict transmitter from EM, but not the receptor or the sign.
- In the worm, MC's K3 shows that sign is decided by intracellular chloride, which is unmeasured.
Program. For every edge, a posterior over the receptor classes present and the sign. It combines:
- the EM transmitter posterior;
- postsynaptic receptor expression (CeNGEN; Fly Cell Atlas and the optic-lobe atlases), with a dropout model;
- E_Cl as a latent per type, with a prior from cation–chloride cotransporter expression (KCC-2, NKCC homologues).
The sign is reported as a function of E_Cl, not as one number. The main derived output is a chloride leverage map: the edges and types whose sign flips across the E_Cl prior, ranked by how much they change E2 predictions.
Novelty.
- Worm: partial. Fenyves et al. 2020 assign polarity to 73% of worm chemical synapses from transmitter and ionotropic receptor rules (EleganSign). The method is deterministic, with no dropout model and no chloride term. A KU Leuven study added ACh-gated anion channels (verify).
- Fly: novel. Davis et al. 2020 (67 optic-lobe types) and Sanfilippo et al. 2024 (tagged receptor subunits in a few types) cover parts of the problem. No whole-brain fly sign posterior and no chloride-dependent sign model were found.
Sister projects. WS and FB both use transmitter-based signs, and their results set the output format.
- WS's Fenyves + CeNGEN rule leaves 2,743 of 3,638 worm edges "conflicting", and moving the expression threshold from 2 to 4 changes 971 categories. So:
- "conflicting", "no evidence" and "uncertain" must stay distinct outputs, not collapse to p = 0.5;
- the expression threshold is a source of uncertainty to marginalize over, not a setting.
- WS's sign probe shows that agreement between fitted and molecular signs is too weakly powered to validate the prior. Validation must use the frozen known-sign set and held-out responses, not the drift of fits.
- WS mirror reliability shows many labelled inhibitory edges sit at low synapse counts (median 2) and reproduce across sides less often (43% against 57% for all edges). Edge reliability (C11) is therefore an input to the sign posterior.
- FB treats all glutamate as inhibitory. Glutamate carries 16% of input there, and removing sign costs 0.272 on its benchmark. This makes fly C3 directly testable: replacing FB's transmitter signs with posterior samples must not lower the 17-assay score. Assays that depend on GluClα pathways, such as Delta7→EPG inhibition, are reported separately.
MVP.
- Worm: all edges, with the chloride leverage map. This is the direct input to MC M11 priority 1 and to CAKE L3 priors.
- Fly: the antennal lobe and one optic-lobe motion pathway, where validation cases exist, then the whole male CNS as a drop-in sign table for FB.
- Host: WS already ingests Fenyves and CeNGEN with hashes and has sign groups and shuffled-label controls. The worm product is built there and exported.
Test.
- Known-sign validation set, frozen before fitting. It includes GluClα-mediated inhibition in the antennal lobe, motion-opponent inhibition in the lobula plate, AVR-14 and LGC-47/ACC-1 synapses in the worm, and EXP-1 excitatory GABA.
- Pass: posterior sign is better calibrated (Brier score) than transmitter-only and Fenyves rules on this set.
- Worm functional check: for monosynaptic pairs in the Randi et al. 2023 atlas, the sign of the predicted response under the best-supported E_Cl beats transmitter-only sign on held-out pairs. Results are reported separately for edges that are sign-stable and sign-ambiguous.
- Kill (fly): if the calibration gain over transmitter-only is within the bootstrap interval, the fly product is limited to the receptor-class posterior without sign.
C4. Gap-junction prior from contactome × innexin expression
Gap. Fly EM connectomes show no gap junctions, and L4 has to infer them from model residuals. MC M4 has a gap-junction head that takes contact areas a_ij as input, but no fly contact areas exist.
Program.
- A whole-brain fly contactome: membrane contact area between every pair of touching neurons, computed from FlyWire, BANC and male-CNS meshes.
- Innexin expression by cell type from the Fly Cell Atlas, combined with hemichannel pairing rules, including heterotypic and rectifying ShakB isoform pairs.
- A model P(gap junction | contact area, innexin pair compatibility, type), fitted on the worm, where gap junctions are annotated in EM. Contact areas there come from Brittin et al. 2021 and innexin expression from CeNGEN and the Bhattacharya map.
- The fitted form, transferred to the fly with the fly's own innexin biology.
Novelty.
- Worm: partial. Barabási & Barabási 2020 (verify authors) inferred 19 innexin interaction rules from binary adjacency and expression. Brittin et al. 2021 give worm contact areas but no gap-junction model.
- Fly: novel. No fly contactome and no fly gap-junction detector in EM were found. Ammer, Vieira & Fendl 2022 map innexins at light level and note that electrical synapses are invisible in connectomes.
Sister projects.
- FB has one electrical synapse, the hand-added GF→TTMn junction, and looming escape fails without it. A posterior over fly electrical synapses lets FB replace a hand-written component with a measured prior.
- WS's gap rectification groups (GAP-RECTIFICATION) can take C4's innexin pairs as their grouping.
- WS mirror data show worm gap annotations are themselves noisy (both-present correlation 0.46), so the worm "ground truth" must be weighted by C11 reliability, not taken as binary.
MVP.
- Worm model with held-out validation, labels weighted by left–right reliability.
- Fly contactome for the optic lobe or one neuropil, with gap-junction posteriors.
- Fly validation pairs: giant fiber with TTMn and PSI (ShakB), VS–VS and VS–HS coupling, and Inx7 in the mushroom body.
Test.
- Worm: leave-one-animal-out (Cook 2019 vs. Witvliet adult). Area under the precision–recall curve for annotated gap junctions must beat (a) contact area alone, (b) expression alone and (c) the Barabási rules.
- Fly: the known coupled pairs rank in the top decile of the posterior for their neuropil.
- CAKE use: the C4 posterior becomes the L4 prior, and the H3 specificity check (02-hypotheses.md, F6) is rerun with it.
- Kill: if the worm model does not beat contact area alone, there is no evidence that expression adds information, and the fly product is reduced to the contactome as a data release.
C5. Spatial volume-transmission kernels
Gap. L5 and MC §12.5 couple peptides either all-to-all or with one decay length. The worm neuropeptide connectome (Ripoll-Sánchez et al. 2023) is expression-based, with short-, mid- and long-range variants defined by anatomy, not by release geometry.
Program. Dense-core vesicle (DCV) detection in EM, release-site maps per neuron, and diffusion with uptake through the extracellular space. Where fixation does not preserve extracellular space, a tortuosity parameter stands in for it. The result is combined with receptor expression to give per-peptide coupling kernels for every pair of neurons.
Novelty. Partial: each component exists, and the pipeline does not.
- Kinney, Sejnowski et al. 2013 ran diffusion in extracellular space reconstructed from EM, in rat CA1.
- McKim et al. (FlyWire neurosecretory network) and the 2024 fly clock connectome infer paracrine links topologically.
- DCV annotation exists only in small fly EM volumes.
- One study reports that spontaneous DCV release does not co-localize with active zones (verify), so DCV position may not mark the release site.
Why it is last. MC's K2 found no gain from peptidergic heads on the worm causal atlas, in agreement with Creamer et al. 2024. A spatial kernel helps only if the residual signal exists.
Sister projects.
- FB tested four octopamine operators, none spatial, and none reproduced starved hyperactivity.
- WS has well-mixed modulator pools and a multirate solver. It is blocked on source-backed release and receptor maps, not on spatial structure.
Both point to the same first step: a provenance-carrying, non-spatial map (which neurons release which peptide or amine, and which express its receptors), built with C1's machinery. Spatial kernels come after that map exists.
MVP (a kill test first). Use the Randi et al. 2023 atlas and the existing unc-31 comparison. Compare the residuals of a synapse-only model on pairs without a wired connection, with and without the Ripoll-Sánchez mid-range variant as a regressor. Continue to DCV detection only if this gives a positive held-out gain, with the 95% interval excluding zero.
Test (if continued). The predicted kernel must explain held-out unc-31-dependent responses better than the expression-only graph. It must also reproduce at least one measured spatially restricted peptide effect, frozen as a target before fitting.
C6. Error-certified reducer for EM neurons
Gap. The N1 rung of the fidelity lattice (06-compute.md §3) needs reduced neurons with certified error over a declared domain. B1 found that cable parameters and diameters dominate the uncertainty. MC's full morphology reduction (M6 and M7) is a later delivery.
Program.
- Diameter and membrane-area estimates with uncertainty from EM meshes, including a shrinkage correction.
- Reduction to the smallest compartment model whose transfer impedances, to the spike-initiation zone and between output sites, stay within a stated frequency-dependent bound. The bound is the H2 or information-budget tolerance of 12-information-budget.md.
- Uncertainty from the diameter estimates is propagated into the bound.
- Batch runs over whole connectomes. The output is CKL-ready and MC-ready compartment graphs.
Novelty. Partial.
- NEAT (Wybo et al. 2021;
neatdend) places compartments from impedance in the frequency domain. It is the method to beat, and whether its criterion is a stated bound must be checked. - neuron_reduce (Amsalem et al. 2020) has no error bound.
- A Janelia preprint compared full and reduced models of medulla neurons (verify).
No certified, uncertainty-aware batch reducer for EM data was found.
Sister projects. FB's point neurons scale PSPs by neuron volume to the power −0.57, and removing even that crude morphology costs 0.327. FB is therefore the fly-scale test of whether N1 reductions beat a size-scaled point neuron at behaviour level. Its packed graph does not keep synapse positions, so C6 has to export per-synapse compartment assignments with the reduced models.
MVP.
- The B1 neuron set (3,397 neurons in 16 types, plus the 600-neuron stratified sample).
- Bounds at 1, 10 and 100 Hz.
- Report compartments needed per type, and the fraction of neurons where diameter uncertainty, not reduction error, dominates the bound.
Test.
- Against full-morphology passive simulations on held-out neurons, the bound holds (true error ≤ bound) in ≥ 95% of (neuron, frequency, site) cases.
- At equal error, the compartment count is no larger than NEAT's.
- With diameter uncertainty propagated, the B1 class rankings of headroom are unchanged.
- Kill: if diameter uncertainty alone exceeds the tolerance for most types, the product becomes a diameter-measurement request (which neurons, which branches), not a reducer.
Simulation infrastructure
C7. Virtual microscope
Gap.
- D3 scores against raw fluorescence, but there is no forward model of the worm or fly rigs, with motion, z-scan timing, neuron identification errors and indicator expression.
- I2 (developmental GCaMP buffering) and MC's gauge result (indicator gain correlated with identity) both need such a model.
Program. A renderer from simulated Ca²⁺ (and, optionally, voltage) to raw volumes for named rigs. It models:
- the indicator: kinetics, buffering and per-neuron expression drawn from an expression prior;
- the optics: point-spread function and scan geometry;
- time: z-scan timing and bleaching;
- the animal: posture deformation in freely moving worms;
- the labels: NeuroPAL colour channels.
Existing extraction and identification pipelines are then run on the rendered volumes, so their errors become part of the observation model and can be measured.
Novelty. Partial.
- NAOMi (Song et al. 2021) is mouse two-photon only.
- SINETRA (2024) synthesizes worm-like tracking data without indicator, activity or identity layers.
- fDNC (Yu et al. 2021) and targeted augmentation synthesize point clouds or deformed images.
- WormID-Bench (2025) provides real-data metrics but no generative model.
No activity-to-raw-volume renderer for worm or fly was found.
Sister projects. WS's preprocessing audit shows that the public worm data are not raw:
- the Randi atlas is processed ΔF/F, with spike removal, smoothing and bleach correction;
- the WormWideWeb traces are whole-recording z-scores, so later samples change the earliest values by 0.92–2.11 z in a synthetic test;
- per-event pulse duration and optical power are missing.
D3 asks for likelihoods against raw fluorescence, and that cannot be met on these datasets. C7 must therefore render through the published preprocessing as well as to raw volumes, so that preprocessing bias can be measured and models scored like-for-like. WS supplies the audits and the confidence-weighted masking convention.
MVP.
- Scope: immobilized NeuroPAL worm whole-brain imaging (the Randi et al. 2023 rig class).
- Reuse: SINETRA's deformation and noise components; WormID-Bench metrics for scoring.
- Outputs: (i) a rendered synthetic atlas from an MC, WS or CAKE model, both raw and passed through the published Randi preprocessing; (ii) the error of the standard extraction pipeline in ΔF/F amplitude, sign and identity, as a function of indicator expression.
Test.
- Realism: on held-out real recordings, rendered and real volumes match within the between-animal spread on summary statistics frozen in advance (noise spectra, photobleaching curves, cross-talk between neighbouring neurons, identification error rates on WormID-Bench).
- Use: the extraction-induced error in response amplitude must be quantified per neuron. If it correlates with expression identity as strongly as MC's gauge proxy does, MC's amplitude flag is confirmed as an observation effect and CAKE adds the C7 error model to D3.
- Kill: if real and rendered volumes cannot be matched on these statistics, the product is reduced to the indicator and expression layer without optics.
C8. Synthetic organisms
Gap. Every inference method in the field is tested on its authors' own toy problem. Neither CAKE's M3 nor MC's Phase 1 can tell whether a failure means the model is wrong, the data are insufficient or training is inadequate. MC's B0 tie is exactly this ambiguity.
Program. A library of fully specified worm-scale synthetic organisms. Each has:
- a known θ, generated by several mechanisms: a CAKE regulation program under H*, R1 hardwiring and MC compositional molecular rules;
- known developmental histories and known latent edges.
Each organism generates datasets that match real designs: the neurons imaged in Randi et al. and Atanas et al., the same stimulations, durations, observation noise, and C7 rendering where available.
Two uses:
- Method benchmark: any pipeline is scored on recovering held-out predictions and the parameters that matter.
- Power and cause analysis: given the real design, could a correct model beat B0 at all? If it could, by what margin?
Novelty. Partial, close to novel at this scale.
- Lueckmann, Jain & Januszewski 2026 use an in silico zebrafish for system identification (verify; abstract only). It is the closest competitor and has no fluorescence rendering or design matching.
- Guo et al. 2023 test model comparison metrics on known ground truth.
- Cleo (2023) simulates devices on spiking networks.
- The Brain Emulation Challenge (Carboncopies) poses small synthetic system-identification tasks (verify).
No connectome-constrained, design-matched worm test bed was found.
Sister projects.
- FB has already built the smallest version.
- 34-individual-validation generates twelve synthetic individuals and runs a pre-registered identification test with a power grid over observable reliability ρ, readout sensitivity κ and fit error β.
- An adjoint over all 165,122 per-neuron gains exists.
- The finding that behaviour identifies 1 of 12 individuals while motor pools identify 12 shows what a known-answer test is for: it caught an uninformative observable before any fly was recorded.
- Lesson: report κ, the sensitivity of each observable to the planted parameters, alongside every score.
- WS shows what happens without one. The capacity gate has failed at 74.8–79.31% across every optimizer variant, and the protocol cannot attribute the failure. A planted-truth organism drawn from WS's own Level 0 class, fitted with the same harness, would answer whether 90% is reachable at all.
MVP.
- Three organisms: H*, R1, and MC rules with a known molecular rule. Each is on the Cook 2019 wiring and generates a synthetic Randi-design atlas.
- Two WS arms on the same wiring: one with a planted Level 0 truth inside the model class, and one with planted truth outside it (for example, a graded-release or adaptation mechanism Level 0 lacks). Both are scored on the ADAL/ADAR capacity target.
- A0's code is the starting point for the regulation organism; FB's
motor_identifyharness is the template for fly-scale organisms. - Release as frozen, hashed datasets, with the true parameters held privately for blind scoring.
Test.
- Run the MC Phase 1 pipeline and the CAKE M2 baselines on all three organisms.
- The test bed is useful if it separates causes. In at least one organism, the correct-mechanism model must beat B0 by more than the real-data interval width; otherwise the real B0 tie says nothing about mechanism.
- It must also reproduce the training-adequacy distinction of MC §10.3 with a known cause.
- WS arms: if the in-class arm reaches the 90% gate with the same optimizer budget, WS's failure is misspecification. If it also stalls near 79%, the gate is an optimization or identifiability limit, and the gate itself needs revising.
- Generalize FB's harness: every task in the test bed reports κ per observable, and a score on an observable with κ near zero is not reported as evidence.
- Kill: none for the test bed itself. If a correct model cannot beat B0 on the synthetic Randi design, that is the result, and it moves the commissioned datasets up in priority.
C9. Differentiable sensory front-ends
Gap. L8 and MC M13 close the loop through a body, but sensing is hand-coded per project. E4 tasks such as chemotaxis, thermotaxis and odour-guided walking depend on stimulus-to-receptor mappings that no shared library validates.
Program. A library of transduction modules with one interface. Each module maps a stimulus field at the sensor to a receptor-neuron current, with adaptation, and is fitted to and validated against recordings. The library is differentiable throughout and runs with JAX and the MC or CAKE engines. Initial modules:
- worm: AFD thermosensation (also the AFD positive control in 03-evaluation.md), ASE salt, AWC/AWA odour, touch;
- fly: ORN dose–response and adaptation from DoOR and Hallem–Carlson, photoreceptors.
Novelty. Partial: the components exist separately.
- flygym and NeuroMechFly v2 (Wang-Chen et al. 2024) render vision and odour concentration at the sensor but do not model transduction.
- FlyBrainLab's odorant transduction process (Lazar et al.) is fly-only and framework-bound.
- drosolf is a lookup table.
- mod-SenseWorm (2025) encodes stimuli configurably but is not validated per neuron.
- BAAIWorm's sensory mapping is bespoke to that model.
No worm-side library and no unified, validated, differentiable one were found.
Sister projects.
- FB has the fly visual front end working: flyvis in WASM, matching PyTorch to 2×10⁻⁶. It also has a C¹ soft coupling map for visual gradients.
- FB's odour, taste, touch and proprioception encoders are analytic Poisson rates, not fitted or differentiable. Its weakest point is olfaction: projection neurons idle at about 77 Hz against a few to 20 Hz measured, because ORN presynaptic GABA_B gain control is replaced by a normalization.
- WS specifies
sensory_inbut has not implemented a body.
So the fly MVP is olfaction with presynaptic gain control, not vision.
MVP. AFD and ASE modules plus fly ORNs for 20 DoOR odorants with presynaptic gain control, each with a validation report, plus adapters to flygym and the MC M13 body ladder.
Test.
- Each module predicts held-out recordings (stimulus protocols not used in fitting) with a normalized score above a threshold frozen per module, using the E1 scoring rules.
- FB check: with the C9 ORN layer, antennal-lobe projection-neuron idle rates fall into the measured range without the
AL_NORMnormalization, and FB's odour assays do not lose score. - Closed-loop check: replacing a hand-coded sensor in an existing worm closed loop with the C9 module does not lower E4 chemotaxis and thermotaxis indices.
Inference and experiment design
C10. Connectome-aware experiment planner
Gap. CAKE's active learning (05-inference.md) and MC M11 rank experiments inside their own pipelines. Labs need a planner they can run before they build a strain or book a rig. It should say which neurons to image and stimulate, at what frame rate and for how long, and what each choice buys.
Program. The planner takes a connectome, a reduced or linearized model with a parameter prior, a candidate design and an observation model (C7 when available). It returns:
- the expected information gain on stiff directions, in nats, using the units of 12-information-budget.md;
- the minimal recording set that resolves the dynamics (Beiran & Litwin-Kumar 2024);
- which hypotheses (F-tests, MC K-criteria) each design can decide at the stated power.
It runs locally from a config file and reports its assumptions.
Novelty. Partial.
- Beiran & Litwin-Kumar 2024 give the theory for prioritizing recorded neurons, but not a tool, and they do not treat frame rate, duration or the indicator.
- Wagenmaker et al. (NeurIPS 2024) choose photostimulation targets actively, without a connectome prior. They are the baseline for the stimulation part.
- Lewi, Butera & Paninski 2006 and Bayesian microcircuit-mapping work are single-neuron or address connectivity.
No pre-experiment planner for connectome-constrained models was found.
Sister projects.
- FB's hypothesis lab already ranks experiments by ensemble disagreement (
rankExperiments), and MC M11 does the same. FB found that the spread between ensemble members exceeded the Delta7-silencing effect being ranked, so the ranking sorted noise. C10 must therefore score information gain against a noise floor measured from the ensemble itself, and report designs below the floor as uninformative. - WS shows the other constraint: confirmatory data are scarce. There is no fresh worm cohort, and one candidate turned out to duplicate the training data. The planner should report whether a design can be tested at all on data that exist, separately from what it would gain if run.
MVP.
- Worm only, with the MC linear-response model as the default model. FB's lab code is the fly port.
- Candidate designs: stimulation targets, imaged subsets, frame rate and duration.
- First real use: size commissioned dataset 1 (09-roadmap.md) and the MC K3 chloride and transporter mutant panel.
Test.
- On C8 organisms, designs ranked higher by the planner give lower held-out prediction error after fitting.
- Pass: Spearman correlation ≥ 0.6 between predicted information gain and realized error reduction across about 20 designs, and the top-ranked design beats random and Wagenmaker selection.
- Kill: if the ranking is no better than a heuristic (stimulate the highest-degree neurons, image everything), the planner is reduced to the power calculator.
C11. Generative model of connectome variability
Gap. Each connectome is used as if it were the animal. CAKE's developmental wiring model and E6 need an explicit split between age trend, individual deviation and reconstruction error. MC Track I found no individual latent in responses; whether wiring differences predict individuality is untested.
Program. One hierarchical model fitted jointly across the worm series and the fly connectomes. Worm datasets: Witvliet 2021 (8 stages), Cook 2019 and adult datasets from 2024–2026. Fly datasets: FlyWire, hemibrain, MANC, FANC, BANC and male CNS. Left–right homologues serve as within-animal replicates. For each edge it estimates the age trend, individual variance and detection error, with a prior from published synapse precision and recall. It samples plausible individuals, and it can condition on a partial connectome.
Novelty. Partial.
- Richter & Schneidman 2024 give a stochastic developmental generative model for the worm, without a split between error and individual variation.
- Schlegel et al. 2024 compare FlyWire and hemibrain descriptively and find that edges above about 10 synapses are conserved.
- Witvliet et al. 2021 describe a stereotyped core and variable background.
No joint, cross-dataset, per-edge decomposition with sampling was found.
Sister projects. WS and FB have each fitted a censored hierarchical left–right model. Their results differ in a way the joint model must keep:
- Worm (WS): existence is the uncertain quantity. Dropout probability is 0.64 at expected count 0.5, and size reliability is 0.98 when both sides exist. The model overpredicts presence by 4–6 points at counts 1–3.
- Fly (FB): weight is uncertain, with SD 0.64 log units at three synapses. Reliability is 0.87 at three synapses and 0.98 at 100 or more. Empirical-Bayes shrinkage moves 99.4% of weights, taking the total from 104.2 M to 65.6 M synapses.
Both are lower bounds: within one animal, left–right disagreement mixes reconstruction error with true asymmetry. Only a second animal separates them (the Witvliet series in the worm; FlyWire against the male CNS in the fly, planned in FB but not done). So C11 is mostly integration: one model family with species-specific existence and weight components, fitted across animals.
MVP.
- Worm: Witvliet series plus Cook 2019, with left–right pairs, starting from WS's model.
- Outputs: per-edge variance components, a sampler, and the predicted distribution of E2 responses induced by wiring variability alone. That last output is the wiring part of the E2 ceiling.
Test.
- Held-out connectome likelihood (leave one animal out) beats Richter–Schneidman, a per-edge binomial baseline and the left–right-only WS model.
- Fly: fitted on the male CNS left–right pairs, the model predicts FlyWire edge presence and weight for matched types better than FB's current empirical-Bayes weights.
- Variance attributed to reconstruction error matches independent proofreading audits where they exist.
- CAKE use: the wiring-induced response spread is compared with the between-animal spread in Randi et al. 2023. The fraction it explains bounds what E6 can gain from an individual connectome. If the fraction is small, E6 needs physiology, not wiring, which is consistent with MC I0.
C12. Within-type channel covariance atlas
Gap. H* predicts conserved low-variance combinations of channel expression within each type (F4; "Expression covariance as channel-kinetics data" in 13-proposals.md). They have been measured only in a few crustacean neurons with single-cell qPCR.
Program. For every type in CeNGEN, the Fly Cell Atlas and the optic-lobe atlases:
- Deconvolve technical noise (sampling, dropout, depth) using an existing correlation-correction method such as BigSur or Sanity, not a new one.
- Estimate the within-type covariance of channel, receptor and transporter genes.
- Report the low-variance directions, their stability across animals and batches, and whether they are shared by types with similar channel sets.
- Report per-type power: which types have enough cells to detect a given correlation.
Novelty. Novel as specified.
- Schulz, Goaillard & Marder 2006 and Tobin, Cho & Marder 2009 measured correlations in crab neurons with qPCR.
- O'Leary et al. 2013 give the theory (correlations emerge from homeostatic tuning rules).
- Allen Patch-seq work relates genes to properties, not channel–channel covariance.
- BigSur (2024) supplies the noise correction.
No atlas-wide noise-corrected channel covariance was found.
Sister projects. None measures this yet. WS imports CeNGEN at class level only (128 clusters). FB individualizes neurons with independent lognormal gains (σ = 0.25), which a covariance structure from C12 could replace.
MVP. CeNGEN, all 128 classes. Types below the power limit are reported, not dropped. Fly optic lobe second.
Test. Preregister which quantities are reported and which hypothesis each decides.
- (i) Correlations among channel genes exceed those of matched control gene sets with similar expression level and dropout.
- (ii) The low-variance directions replicate across independent animals and batches.
- (iii) H* predicts the low-variance directions span row(J), so they should align with what each type's channel kinetics imply. They are compared against R1, which predicts covariance driven by shared transcriptional regulation and therefore aligned with regulon membership.
The F4 statistics in 02-hypotheses.md are computed on these covariances. If neither alignment exceeds a permutation null, the result is "undetectable with atlas data", and data request 5 (09-roadmap.md) is sized from the measured power.
Priority and hosting
| Order | Program | Host | Reason |
|---|---|---|---|
| 1 | C8 | CAKE (A0) with WS and FB arms | The question that stalls all four projects (MC's B0 tie, WS's capacity gate, FB's identification null, CAKE's M3) is only answerable with a known-answer control |
| 1 | C11 | WS (worm), FB (fly) | Both left–right precursors exist; the remaining work is cross-animal fitting and a joint model family. Feeds C3, C4 and E6 |
| 1 | C3 | WS (worm), FB (fly) | WS's 2,743 conflicting edges and MC's K3 failure point at the same gap. FB is an immediate fly test bed |
| 2 | C12 | CAKE | Public data only; early evidence on H* versus R1 |
| 2 | C1 | Shared | Replaces three hand-built tables (MC kinetics, WS phenotype registry, FB calibration targets) |
| 2 | C7 | WS | WS's audit shows the public worm data are preprocessed; needed before amplitude or D3 claims |
| 3 | C4 | FB (fly), WS (worm) | Removes FB's hand-added electrical synapse; needs C11 reliability for worm labels |
| 3 | C10 | FB lab code, MC M11 | Needs C8 to be tested, and a noise floor (FB) |
| 3 | C6 | CAKE (B1) | Extends B1; FB tests whether it beats size-scaled point neurons |
| 4 | C9 | FB (fly olfaction), WS (worm) | Fly olfaction first; worm when WS has a body |
| 4 | C2 | MC | No channel models in WS or FB yet; must pass leave-subfamily-out |
| 5 | C5 | WS, FB | Spatial kernels only after a source-backed non-spatial map exists and the kill test passes; MC K2 and FB's operator results are both negative |
Shared conventions
- Every program's outputs are versioned, content-hashed artifacts with provenance and licence tier, following MC
data-register.json, WS input hashes and run manifests, and the CAKE Data Commons (07-platform.md §7.6), so any of the four projects can consume them. - Programs use each project's canonical neuron identifiers and ship mapping tables, not re-keyed copies of the connectome.
- Each program's test, threshold and kill condition is frozen before its first real-data analysis, as for F1–F7 and the MC K-criteria.
- A program that fails its test still publishes its negative result and its data product, labelled with the failure.