13. Proposed theories for the main bottlenecks
Status: Proposals, 2026-10-07. None of this is current design. Each proposal names the bottleneck it targets, the claim, the test that would decide it and what it would change in the spec if adopted. Adoption goes through the decision log (decisions.md). One numerical check (T4) was run during drafting; it is described below but is not stored evidence. Citations marked (verify) have not yet been checked and must be confirmed before they enter references.md.
Two items are not proposals but problems in the current spec. They are listed first.
Issues found in the current spec
I1. Equilibrium fluctuations carry information about regulation rates. L7 (04-physics.md, "Identifiability") states that steady-state data, "including within-type covariance, carry zero Fisher information about C_c". This holds for the deterministic averaged dynamics. It does not hold once sensors fluctuate around their averages with correlation time τ_mix: the stationary spread of the regulated state then scales with the rate, and its autocorrelation time equals τ_reg (T4 below). The statement should be restricted to static, time-averaged equilibrium data. The A0 roadmap line already lists "sensors as time averages of stochastic activity" as open (09-roadmap.md); T4 says what that extension should measure.
I2. The calcium indicator is present during development, not only at readout. L6 puts GCaMP in the model as a buffer and as the readout. Pan-neuronal indicator strains express it throughout development, so under H* every imaged animal was reared with an added Ca²⁺ buffer, and regulation has had the chance to compensate for it. Burn-in for predictions about imaged animals must therefore include the indicator from its developmental onset; predictions about non-imaged wild-type animals (most E4 behavior data) must burn in without it. T8 turns this into a test.
Bottlenecks targeted
| # | Bottleneck | Where in the spec | Proposals |
|---|---|---|---|
| B1 | H* and R1 are separable only with commissioned chronic or timed-depletion data | 02-hypotheses.md, "What H* claims"; 09-roadmap.md, data to commission | T1, T3, T7, T8 |
| B2 | Regulation rates C_c are identified only by transients | L7 identifiability; M3 | T4 |
| B3 | The history state ξ is defined but not observable; A0 obtains it from an oracle | 11-prediction-and-history.md | T3 |
| B4 | Burn-in is nested in the optimizer; branch and basin failures | 06-compute.md §1, §5 | T2, T5 |
| B5 | Locality is assumed at the level of the cell | 02-hypotheses.md, L7 learned family | T6 |
| B6 | Φ may be unidentifiable, and nothing transfers it to unseen types (G-identity) | 10-risks.md; 03-evaluation.md | T9 |
T1. The connectome as a dose–response curve (B1)
Claim. Under H*, neurons of one type that receive more input compensate by changing their intrinsic parameters, so that the sensor-controlled activity does not depend on input. Under R1 intrinsic parameters do not depend on input, and activity rises with it.
Derivation. The linearized adult deviation in L7 is
δz* = P δz₀ + G (J G)⁻¹ δr
Part of δr is each neuron's difference in input drive from its type mean, δr_i = −J_in,i δu_i, and δu_i (input synapse counts per presynaptic type, or the H2 transfer-weighted drive) is measured by EM. H* therefore predicts:
- E[δz* | δu] = −G (J G)⁻¹ J_in δu: a regression of intrinsic parameters on EM input with rank ≤ k, lying in range(G);
- J δz* = −J_in δu: the sensor-controlled combinations exactly cancel the input difference, so mean sensor activity (e.g. mean Ca²⁺ in the sensor band) is uncorrelated with input within the type.
R1 predicts zero slope for intrinsic parameters and a positive slope of activity on input.
Statistic. A cross-sectional compensation fraction κₓ, defined like κ in 02-hypotheses.md but across neurons rather than over time: the slope of activity on input, relative to the slope the frozen-regulation model predicts. H* predicts κₓ → 1, R1 predicts κₓ = 0.
Data. Same-animal measurement of input and of intrinsic properties or activity. The best candidates are fly types with many members and measurable input variation, for example Kenyon cells, whose claw number varies several-fold and can be counted in the recorded cell by dye fill. Compensatory variability between Kenyon-cell input and excitability has been reported (Abdelrahman, Vasilaki & Lin 2021 (verify)). In the worm, left–right pairs with asymmetric input in a same-animal EM dataset (data request 3) give small-n paired tests.
Confound. R1 can produce the same correlation only if input and intrinsic parameters are both specified by a genetic subtype. Two checks separate this: the slope must persist within transcriptionally homogeneous clusters, and its direction must lie in range(G) as estimated from acute perturbation (F3).
Changes if adopted. Adds an F-test (κₓ) to the H* vs. R1 row of the decision rule. It can run before commissioned datasets 1 and 2 arrive.
T2. Critical periods are near-bifurcations of the slice equation (B4, F7)
Claim. Development acts as a homotopy in the coupling strength. Early wiring is sparse, the network term J_o of the loop gain is small, and the slice equation is close to block-diagonal and well posed. The biological branch is the one continuously connected to this weakly coupled solution along the developmental schedule. Developmental critical periods are the stages at which σ_min(J G) dips along this path: there the endpoint is sensitive to activity, and transient manipulations can switch branches (mechanism 1 of "When history matters", L7).
Computation. Replace the pseudo-transient start at a = 0 (06-compute.md §5.2) with numerical continuation in a coupling parameter λ, scaling W(t) from its earliest measured stage to the adult. Track σ_min(J G) and fold or branch points along the path. This is both a branch-tracking solver and a direct model of the selection rule the schedule implies.
Prediction. For each type, the stage at which σ_min(J G) is smallest. Timed manipulations (e.g. HisCl1 silencing) that span that stage should give the largest persistence ρ; manipulations of equal strength and duration at other stages should give ρ near the constant-gain prediction. F7b then tests a stage-specific prediction rather than the presence of persistence.
Link to T4. Near a dip in σ_min(J G), regulation relaxes slowly. Fluctuations of the regulated state should then grow in variance and autocorrelation time (critical slowing down), which gives an independent marker of the predicted window.
Changes if adopted. Continuation becomes a solver mode in §5; F7 pre-registration names stages from the model.
T3. History is written in the conserved coordinates (B3, F7b)
Claim. Under Case 1 the n − k quantities N_cᵀ z_i are conserved, so adult values of N_cᵀ z record the initial draw. The mechanisms that let history persist (non-involutive gain, higher-order error terms, saturation) are exactly the processes that move N_cᵀ z. The compressed history state of 11-prediction-and-history.md is therefore
ξ_i(t) = N_cᵀ (z_i(t) − z_i(t₀))
In log coordinates these are log-ratios of channel densities. A0's c_i = log h_i − ρ_i log g_i is one of them.
Test (snapshot classifier). Measure within-type expression (or conductances) in animals reared under two conditions. Decompose the mean shift against range(G), which is estimated from the direction of the response to acute perturbation, since dz ∝ G e:
- a shift inside range(G) is the adult environment acting through the slice (Case 1);
- a component outside range(G) is a lasting history effect.
This needs one time point per condition instead of a manipulation, a recovery period and a second measurement.
Changes if adopted. Gives the A0 extension proposed in §11 (infer c from recordings) a concrete observable: expression ratios, not activity. Adds a snapshot form of F7b to the commissioned datasets.
T4. Regulation rates from spontaneous fluctuations (B2)
Claim. If regulation is one linear integral controller, the autocorrelation of a regulated quantity in an unperturbed neuron and the recovery kernel after a small perturbation are the same propagator (Onsager's regression hypothesis applied to homeostasis):
C_z(τ) = exp(−G C J τ) C_z(0)
So regulation rates can be measured without a perturbation, and agreement between the two measurements is a test of the linear L7 model. Disagreement indicates nonlinearity or several controllers.
Scalar case. For dz/dt = c g (s* − j z − ξ), with ξ an Ornstein–Uhlenbeck sensor fluctuation of variance σ_ξ² and correlation time τ_mix, and λ = c g j:
Var(z) = (g / j) · c σ_ξ² τ_mix / (1 + λ τ_mix)
autocorrelation time ≈ 1/λ = τ_reg (λ τ_mix ≪ 1)
The spread scales with the rate c, so G → G C is not a symmetry of the stationary distribution. In sensor units the extra variance is of order ε σ_ξ², with ε = τ_mix/τ_reg as in L7: small in amplitude, which is why the averaged dynamics ignore it, but measurable through its timescale.
Check during drafting. Scalar simulation with g = j = σ_ξ = τ_mix = 1, 2×10⁷ steps of dt = 0.01:
| c | Var(z)/c | Theory 1/(1 + c) | Autocorrelation time | 1/c |
|---|---|---|---|---|
| 0.005 | 0.976 | 0.995 | 196 | 200 |
| 0.01 | 0.981 | 0.990 | 96 | 100 |
| 0.04 | 0.966 | 0.962 | 27 | 25 |
Data. Longitudinal imaging of endogenously tagged channels (knock-in fluorescent fusions) in identified neurons over days, in unperturbed animals. Duration must cover several τ_reg; sampling must resolve it.
Prototype (2026-10-07). A0 cake-a0-fluctuation, on the two-sensor neuron with an Ornstein–Uhlenbeck input current: the stationary mean is rate-invariant while the spread grows with the rate (3.7× for 4× the rate, within 1% of the linear theory); loop-gain eigenvalues estimated from unperturbed fluctuations agree with the true values within 6% and with depletion-recovery estimates within 10%; a hidden expression cascade makes the estimate lag-dependent (77% against 6%), so the consistency test detects missing state. Synthetic; the fluctuations come from an imposed input, not network activity.
Changes if adopted. Corrects I1. Adds a fluctuation–recovery consistency test to F3. Extends M3 so that C_c can be partly identified from unperturbed longitudinal data. Gives A0's stochastic-sensor extension its target.
T5. Regulation as a constrained prior: all-at-once inference (B4)
Claim. Under Case 1 with a unique slice equilibrium, H* is equivalent to a prior on the adult regulatory state z* directly, without simulating regulation inside the optimizer. With F(z) = [s* − E[σ(z)] ; Nᵀ z] as in 06-compute.md §5.2 and set-point jitter Σ_η:
p(z*) ∝ N(Nᵀ z*; Nᵀ μ_c, Nᵀ Σ₀ N) · N(e(z*); 0, Σ_η) · |det ∂F/∂z (z*)| · 1[J G stable at z*]
The first factor is the projection of p₀,c onto the conserved coordinates; the second is the sensor condition; the determinant is the change of variables from (initial state, jitter) to z*.
Computation. Fit z* jointly with Φ and the data. Each optimizer iteration needs one closed-loop averaging run to evaluate e(z*), instead of a converged decomposed solve. This is the "one-shot" or all-at-once approach of PDE-constrained optimization. The log-determinant is estimated from the per-neuron k×k blocks plus a stochastic estimate of the coupling term. Stability (V2) and basin (V3) certification run on the final iterate, as in §5.2.
Rival family. The nested family of 02-hypotheses.md becomes a set of relaxations of one prior: R2 drops the sensor factor; R1 sets Σ₀ → 0. Model comparison is then between priors on the same unknowns.
Diagnostic. The standardized per-neuron sensor residual e_i(z*) under the posterior is a regulation-violation map: it shows where the data pull neurons away from their regulated equilibrium. It plays the role for H* that structured residuals play for H3.
Expected gain. A factor equal to the number of outer decomposition iterations per burn-in (§5.3), plus fewer branch failures during optimization. Not applicable where V3 or V6 fail; those types stay in transient mode.
Prototype (2026-10-07). A0 cake-a0-oneshot, on 200 mean-field-coupled two-sensor neurons: nested and one-shot fits reach the same estimate (to 6×10⁻¹²) with 122 and 7 network evaluations respectively; the one-shot optimum passes stability and basin certification; the per-neuron misfit flags idiosyncratic neurons (AUC 0.99 at the largest planted spread, 0.65 when the spread is comparable to the expression noise); the density with |det[J; Nᵀ]| is normalized and without it is not. Synthetic, with G and the noise model known; evaluation counts are not an organism-scale speedup.
Changes if adopted. A third solver mode in §5; A0 is the place to compare it with endpoint mode.
T6. Locality at the synapse, not the cell (B5)
Claim. The learned rule family restricts inputs to the cell's own signals. The best-characterized homeostatic mechanism in Drosophila, presynaptic homeostatic potentiation at the neuromuscular junction, is retrograde: the postsynaptic cell senses and the presynaptic terminal changes (reviewed by Davis and colleagues (verify)). If central regulation works this way, the current family would reject H* when the actual finding is regulation that is local to the synapse.
Rival. Add R3 to the nested family: rules whose inputs include the sensors of synaptic partners, weighted by β. β = 0 is the current H*.
Signature. Deplete a channel in neuron A and measure, through pair stimulation, the transfer from a presynaptic partner j to A and from j to an unperturbed target B:
- cell-local or network-mediated changes in j scale j→A and j→B together;
- retrograde regulation changes j→A only.
Changes if adopted. One new rival and one readout in F3. This should be settled before the M0 pre-registration freezes the hypothesis family.
T7. Within-animal mosaics for F3 (B1, power)
Claim. Genetic mosaics, in which only some neurons of a type carry the perturbation (MARCM or FLP-out in the fly, mosaic transgene loss in the worm), give each perturbed cell an unperturbed sibling of the same type in the same animal. This removes between-animal variance from κ and raises power for a given number of animals.
| Hypothesis | Perturbed cells | Unperturbed siblings |
|---|---|---|
| Cell-local H* | Compensate | Change only through network activity |
| R1 | Do not compensate | Do not change |
| Network-level homeostasis | Shift with siblings | Shift with perturbed cells |
| Synapse-local (T6) | — | Inputs onto perturbed cells change; inputs onto siblings do not |
Changes if adopted. Mosaic designs added to commissioned dataset 1; power computed by the M3 pipeline as for the current design.
T8. The indicator as a chronic manipulation (B1, I2)
Claim. Developmental GCaMP expression adds a Ca²⁺ buffer. A non-saturating buffer leaves steady free Ca²⁺ at constant influx nearly unchanged but slows and shrinks transients. If regulation responds to indicator level, the sensors include fast Ca²⁺ bands; if it does not, they are slow.
Test. Compare animals with and without the indicator, or with indicator variants and expression levels of different buffering capacity, using readouts that do not buffer Ca²⁺ (electrophysiology, voltage imaging, expression). Cell-to-cell variation in indicator expression within a type is also a natural dose for a T1-style regression. Existing strains suffice.
Changes if adopted. Fixes I2 in the burn-in specification. Adds a cheap F7a-like dataset and a sensor-identification test.
T9. Regulatory programs as a function of the terminal-selector code (B6)
Claim. Φ_c is a smooth function of the type's transcription-factor code: Φ_c = f(TF_c). Every C. elegans neuron class has a unique homeobox code (Reilly et al. 2020 (verify)), and the Fly Cell Atlas gives TF profiles per type.
Effect. About 10⁴ independent programs become one function of a few hundred features. This is a principled form of the existing fallback "share regulatory gains across families of types" (10-risks.md), and it is the only proposal here that predicts programs for types with no functional data.
Test. G-identity (03-evaluation.md): predict Φ for held-out classes from their codes and score their E1/E2 tasks. Compare with per-type programs and with a type-family pooling baseline.
Changes if adopted. A parameterization option for Φ in 05-inference.md; evaluated at M3 alongside per-type programs.
Smaller proposals
- Expression covariance as channel-kinetics data. The low-variance directions of within-type expression covariance estimate row(J), which depends on the kinetics of the channels involved. This constrains genes that have only homology priors (L2). Types used this way must be excluded from F4 scoring, under the leakage audit.
- Temperature as an H3 identifiability axis. Gap-junction conductance has Q₁₀ ≈ 1.3, while chemical release and GPCR cascades have Q₁₀ of 2–3 or more (verify per mechanism). Pair responses at two temperatures add a separating signature for components that the coherence diagnostic reports as confounded.
- Type-centred perturbative burn-in for the fly. Solve one representative per type, or per orbit of a repeated structure such as visual columns, and give each member the linear correction G (J G)⁻¹ δr. Certify with shadow neurons (06-compute.md §9). Valid where within-type input variation is small; T1 identifies the types where it is not.
Priority
| Order | Proposals | Reason |
|---|---|---|
| 1 | T1, T3, T8 | Evidence on H* vs. R1 from snapshot data and existing strains, before commissioned datasets arrive |
| 2 | T4, T5 | Theory and solver work that can be prototyped by extending A0 now |
| 3 | T6 | Changes the hypothesis family, so it must be decided before the M0 pre-registration |
| 4 | T2, T7, T9 | Depend on M2b, commissioned datasets or M3 respectively |