← The living map spec / 02-hypotheses

2. Hypotheses and falsification

H*: Per-neuron biophysical parameters are not free parameters to be fitted. They are the steady state of a regulatory process that each cell runs, and that process can be simulated.

Specifically: θ for each neuron is the outcome of activity-dependent homeostatic regulation: an equilibrium of that regulation, selected by the neuron's initial state and developmental history (04-physics.md, L7). That regulation is encoded per cell type, acts on the set of channel and receptor genes the cell type expresses, and runs while the neuron is embedded in the real connectome, which is embedded in a real body interacting with an environment.

If H* holds, the unknowns shrink from per-neuron conductances (~10⁹ in the fly) to per-cell-type regulatory programs (~10⁴ cell types × ~10–50 parameters). The behaviorally relevant (stiff) directions of the system then become identifiable from data that can actually be collected: whole-brain imaging, a few hundred targeted perturbations, and behavior.

What H* claims, and its rivals

The scientific content of H* is locality and activity dependence: each neuron's parameters are set by a rule that uses only that cell's own signals (its voltage, Ca²⁺ and second messengers) and its type's genetic program. The specific functional form (e.g. calcium-sensor rules after Liu et al. 1998) is a modeling choice, not the hypothesis, so CAKE also fits a learned rule family restricted to local signals (04-physics.md, L7). This keeps H* from being rejected merely because one rule form was wrong.

H* is tested against two named rivals:

Hypothesis Parameters are set by Unknowns
H* Regulatory Closure A cell-autonomous, type-specific, activity-dependent rule running in the closed loop Per-type regulatory programs
R1 Genetic hardwiring Cell type alone, with no dependence on activity Per-type parameters (F1 baseline b)
R2 Idiosyncrasy Each neuron individually, with no shared rule Per-neuron parameters (F1 baseline a)

Acute data cannot separate H* from R1. In an animal reared under standard conditions, a regulation fixed point and a per-type parameter set can produce the same activity. Optogenetic pair responses (F1) therefore mainly test H* against R2. Separating H* from R1 needs experiments that change a neuron's activity history and look for the parameter changes H* predicts and R1 forbids (F3, F7 below). These experiments are essential; 09-roadmap.md lists the data to commission.

One family, not three separate models. H*, R1 and R2 are limits of one parameterized family, so the comparison estimates interpretable quantities instead of only picking a winner. For neuron i of type c, in the canonical form of 04-physics.md, L7:

dz_i/dt = λ_c G_c e_i − D_c (z_i − m_i),     m_i ~ N(m_c, τ_c²),     s*_i ~ N(s*_c, Σ_η,c)
Limit Model
λ_c = 0, D_c > 0, τ_c = 0 R1: per-type parameters, no activity dependence (F1 baseline b)
λ_c = 0, D_c > 0, τ_c free R2: per-neuron parameters pulled toward the type (F1 baseline a)
λ_c > 0, D_c small, τ_c and Σ_η,c small H*: shared, activity-dependent regulation

Three quantities summarize where the data put a cell type in this family. Each is estimated with an interval:

Why this is plausible

What is new here

Novelty check, 2026-10-07: a search of recent literature found homeostatic regulation studied in single neurons and small circuits, and connectome-constrained models fitted directly to imaging data, but no work using simulated regulation as the inference mechanism for whole-connectome parameters. Repeat this check before each major release.

Homeostatic conductance regulation is well studied in single neurons and small circuits (STG models). Connectome-constrained whole-brain models are well studied with simplified neurons. As far as we know, nobody has used simulated regulation as the inference mechanism for whole-connectome biophysical parameters. In this approach:

  1. Fitting targets change. We fit set points and regulatory gains per cell type, not conductances per neuron.
  2. Degeneracy becomes a prediction. Within-type variability of conductances emerges from different initial states that regulation carries to different points of the same equilibrium set. To first order its covariance is P Σ₀ Pᵀ + G (J G)⁻¹ Σ_r (J G)⁻ᵀ Gᵀ (04-physics.md, L7). The combinations of conductances that the sensors control vary only as much as set points and inputs do. This structure can be tested against Patch-seq and single-cell RNA data (F4).
  3. The model predicts slow responses to perturbation. A fitted static model cannot predict how a circuit compensates hours or days after a channel knockout, cell ablation or drug. A regulatory model must. This is the strongest falsification test available.

Two companion hypotheses

H2: Electrotonic anchoring. Effective synaptic weight should be computed, not fitted. The physically measurable part of a synapse's influence is its dendritic location, which sets the cable transfer impedance from the synapse to the cell's integration and output sites. EM morphology gives this geometry. Replacing "weight ∝ synapse count" with EM-derived transfer functions removes a large block of free parameters and should improve zero-shot prediction of functional connectivity.

Headroom versus validation. The B1 analysis (09-roadmap.md, Track B) measures headroom: how much replacing counts by EM-derived transfer changes the effective weights. Headroom is necessary for H2 to matter, but it is not evidence that H2 predicts biological responses better. Only F5, on functional data, tests that.

H3: Residual cartography of the invisible connectome. Once the chemically wired, electrotonically anchored, regulation-closed model is fitted, its structured residuals against functional perturbation data are not noise. They mark missing physics: gap junctions, neuropeptide/GPCR links and monoamine volume transmission. These latent edges get priors from gene expression (innexin/connexin co-expression plus membrane apposition in EM; ligand–receptor co-expression for peptides) and are inferred by Bayesian model selection. The output is a predicted wireless and electrical connectome that wet-lab experiments can test directly.

Residual structure is not evidence of missing edges by itself. Errors in channel models, stimulation strength, the observation model and morphology also produce structured residuals. H3 therefore attributes residuals with a structured decomposition in which these sources compete with latent edges. Each source enters with the algebraic signature it must have in the linearized pair-response residual R(ω): rows are responders, columns are stimulated neurons, measured pairs form the set Ω. Observation and stimulation errors act directly on the model prediction. The other sources perturb the linearized network operator by ΔW(ω), which appears in the residual to first order as L(ω) ΔW(ω) R_p(ω), where L and R_p are the model's linear propagators to responders and from stimulated sources.

Source Signature
Observation error in responder i (indicator gain, kinetics) Row scaling of the model prediction, with frequency profile from the indicator model
Stimulation-strength error in source j (opsin expression, light) Column scaling of the model prediction, frequency-flat
Intrinsic (channel) error in neuron i Diagonal ΔW
Chemical-synapse weight error (incl. morphology, H2) ΔW supported on EM synapse edges
Gap junction i–j ΔW symmetric, Laplacian-structured (g_ij enters off-diagonal and with opposite sign on both diagonals), frequency-flat, off the EM chemical support
Peptide / monoamine ΔW = Σ_p K_p(ω) · diag(receptor_p) D_p diag(ligand_p): one coefficient per ligand–receptor pair, spatial diffusion kernel D_p, low-pass K_p

The components are fitted jointly by penalized regression: ridge on row, column, diagonal and synaptic terms, group sparsity on gap-junction edges, sparsity on peptide-pair coefficients. Before fitting, an identifiability diagnostic computes, for each pair of components, the largest principal-angle cosine between their design subspaces restricted to Ω (mutual coherence). Component pairs above a pre-registered coherence threshold are reported as confounded in that dataset, not attributed. The active-learning queue proposes experiments that reduce the coherence. All of this is in the linear-response regime, so the residuals used come from small-signal stimulation or are locally linearized.

Falsification criteria (decided before building)

Every test below is scored with the task specifications of 03-evaluation.md: a defined loss, a normalized score S with a hierarchical-bootstrap interval, and a pre-registered equivalence margin δ. Each test has three outcomes:

ID Prediction Pass Fail
F1 Regulation-closed model (per-type programs) predicts held-out optogenetic responses at least as well as each direct-fit baseline trained on the same data (fair-baseline protocol below). Scored separately against baseline (a) and baseline (b) Non-inferior: lower bound of S_H* − S_baseline > −δ Upper bound of S_H* − S_baseline < −δ
F2 Shuffling cell-type set points, or using a degree-preserving shuffled connectome, destroys prediction Lower bound of S_real − S_shuffled > δ S_real − S_shuffled within [−δ, δ] (equivalence)
F3 After a timed perturbation in adults (e.g. auxin-inducible degradation of a K⁺ channel, depletion measured by a fluorescent tag, tracked over hours), the model predicts the compensation trajectory in amplitude and timing, not only its shape Trajectory score of the regulation model beats the frozen-regulation model by > δ, and lower bound of κ_slow > κ_min κ_slow within [−δ_κ, δ_κ] (no slow compensation)
F4 Within-type conductance/expression covariance has the structure regulation predicts (below) Both F4 statistics beat their null distributions at α = 0.05 Neither statistic distinguishable from the null, with the test powered
F5 (H2) EM-derived transfer predicts measured transfer better than synapse counts for pathways in pre-registered high-headroom types, increasingly with frequency (below) Criteria (i)–(iii) below (i) and (ii) fail, with the test powered
F6 (H3) Residual-inferred latent edges, called at controlled false-discovery rate, are enriched for experimentally verified peptide and gap-junction links beyond expression-only priors. Runs only after the specificity check (below) passes Lower bound of the enrichment odds ratio (residual ranking vs. prior-only ranking) > 1 Upper bound ≤ 1
F7a (H* vs. R1) During a sustained manipulation of activity in development (e.g. HisCl1-mediated silencing of identified neurons, or rearing at a different temperature), the parameters of the manipulated neurons and their partners shift as predicted by re-running burn-in with the altered history. Measured through pair responses, electrophysiology or expression Measured shift in the pre-registered readouts lies outside [−δ₇, δ₇], and the score of the re-run burn-in prediction beats the no-shift null by > δ Shift within [−δ₇, δ₇] (equivalent to no shift)
F7b (rule structure) After the manipulation ends and a declared recovery period passes, the remaining fraction ρ of the shift matches the rule family's prediction. Constant-gain integral rules predict ρ ≈ 0 unless branch selection or saturation applies (L7) Predicted and measured ρ agree within δ_ρ Disagree beyond δ_ρ

The margins δ, κ_min, δ_κ, δ₇ and δ_ρ, the recovery period, the cell types tested and the power requirements are fixed at M0 (F1–F4, F6, F7) or in the Track B pre-registration (F5). They are not changed after data are seen.

F4 statistics. In log-expression (or log-conductance) coordinates, with J the model's sensor sensitivity for the type at its operating point:

  1. Controlled-combination variance: v = tr(J Σ̂ Jᵀ) / E_null[tr(J Σ̂_null Jᵀ)]. The null shuffles cells independently per gene within the type, which keeps every marginal variance and removes covariance. Regulation predicts v ≪ 1.
  2. Subspace alignment: the principal angles between the row space of J and the empirical lowest-variance k-dimensional eigenspace of Σ̂, compared with their distribution under the same null and under random k-dimensional subspaces.

Technical noise (estimated from the count model or spike-ins) is modeled, not subtracted after the fact. Statistic 1 is invariant to per-gene scale factors in log coordinates. Statistic 2 depends on the channel models through J, so F4 also tests the channel library.

F5 criteria and data requirements. For pathways j → i in pre-registered test types, let W^H2_ij(ω) = Σ over synapses of q · Z_syn→target(ω; θ_i), with the target set by transmission mode (spike-initiation zone for spiking neurons, output sites for graded neurons), and W^count_ij = q · n_ij. Let Δ(ω) = S_H2(ω) − S_count(ω) be the score difference in predicting measured transfer, in frequency bands.

Test and control types are chosen from the B1 headroom ranking before any functional data are examined. The dataset is named in the pre-registration and must provide:

Until such a dataset is named, B1 results remain headroom estimates, and F5 is reported as not yet run. F5 is tested mainly in the fly; most worm neurons are compact.

F6 specificity check (precondition). On the real measurement design (the same pairs, noise and stimulation), plant known errors of every non-edge source in the H3 table at realistic magnitudes, with and without planted gap-junction and peptide edges. The decomposition must:

Component pairs that the coherence diagnostic flags as confounded are excluded from F6. Verified links used to score F6 must be independent of the priors (leakage audit, 03-evaluation.md).

F1 fair-baseline protocol. F1 compares our model with a rival, so the rival must be as strong as we can make it:

Power requirements

Each test pre-registers:

The required numbers of animals, neurons, pairs and time points are computed by simulation in the M3 pipeline: data are simulated from each generating model (H* with both rule families, R1, R2) at the real experimental design, and the full fitting and testing pipeline is run on them. The same simulations give a model-recovery matrix: how often the pipeline selects each model when each one generated the data. Data requests to partner labs (09-roadmap.md) quote these numbers. A test whose data fall short of its requirement can only be inconclusive.

Decision rule (pre-registered)

The verdict separates two questions that a single table previously merged:

A model losing a prediction contest shows that the tested formulation is insufficient. It does not by itself show that the biological mechanism is absent.

Step 0: Validity gates

No verdict is drawn until all gates pass. A failed gate makes the overall verdict inconclusive (machinery), with the failure named.

Gate Requirement
G1 AFD positive control passes (03-evaluation.md)
G2 M3 synthetic recovery: every generating model is selected correctly at a rate ≥ 0.8 in the model-recovery matrix, and each test meets its power requirement on the real design
G3 F2 passes. If F2 fails, stop and diagnose before any scale-up: the model is not using the connectome meaningfully
G4 Optimization adequacy: restart stability (fair-baseline protocol), and self-consistency (refitting synthetic data generated from each fitted model recovers it)

Step 1: Predictive model selection

Candidates: the regulation-closed model (each rule family), baseline (b) and baseline (a).

This is F1 read as a selection rule.

Step 2: Mechanistic evidence

Each claim is assessed from its own tests:

Claim Tests Supported Not detected (at stated power) Inconclusive
Activity-dependent regulation (H* vs. R1) F3, F7a F3 or F7a passes Both fail (powered) Otherwise
A shared rule rather than idiosyncrasy (H* vs. R2) F1 vs. (a), F4, ι F1 vs. (a) passes and F4 passes F1 vs. (a) fails and F4 fails Otherwise
Rule structure (which mechanisms carry history) F7b, F4 subspace Reported as estimates of ρ and subspace alignment, with intervals — —

Conflicting tests. F3 can pass while F7a fails, or the reverse. This is not a contradiction: it is evidence for activity-dependent regulation in adults but not during development, or the reverse. The claim is then stated only for the setting where it passed. A failed powered test and a passed test on the same setting (e.g. F3 in two types) are reported per type, not pooled.

Before any negative mechanistic statement, four alternative explanations must each be checked and reported:

  1. Missing mechanisms: Necessity Map ablations and H3 residual structure show no unmodeled process that could carry the effect.
  2. Weak intervention: the measured efficacy of the manipulation (depletion fraction, silencing strength) exceeds its pre-registered minimum.
  3. Insufficient measurement: the test met its power requirement.
  4. Optimization failure: G4 passed for the models compared.

Only when all four are excluded does CAKE state "evidence against activity-dependent regulation of these parameters, in these types, under these conditions". Otherwise the statement is "the tested formulation does not account for the data", with the open alternatives listed.

Step 3: Consequences

Predictive selection Activity-dependent regulation Verdict Consequence
Regulation-closed Supported H* supported Proceed to P1 with regulation-closed models
Baseline (b) Supported Regulation present; tested formulation not predictive (this covers F1 losing only to the per-type baseline) Proceed to P1 with per-type direct fitting; revise rule family and sensors as a research track
Baseline (b) or regulation-closed Not detected R1 preferred Proceed with per-type direct fitting; regulation layer kept for research only
Baseline (a) Any Per-type structure insufficient for prediction Proceed with hierarchical per-neuron fitting; prioritize same-animal data (E6). A claim of biological idiosyncrasy additionally needs F4 to fail and the four alternatives to be excluded
Any Inconclusive H* undecided Proceed to P1 with the predictive selection; extend the commissioned F3/F7 datasets to the size the power analysis requires; the H* verdict stays open and is reported as such

F5 (H2) and F6 (H3) are reported independently and don't change the H* verdict.