7. Platform architecture
CAKE is planned as a long-lived, community-built open-source scientific code, on the scale of the large simulation codes in physics, chemistry and climate science, not as one lab's research repository. The closest analogue is an Earth system model (E3SM, CESM): several physical components (atmosphere, ocean, land, ice), each a large code in its own right, joined by a coupler, run in standard configurations, compared with other models through community intercomparison projects, and governed by a consortium. CAKE has the same shape: neural tissue, chemistry, body and environment as components, organisms as configurations, and the Emulation Ladder as the intercomparison protocol.
7.1 Projects we learn from
| Project | Domain | What CAKE borrows |
|---|---|---|
| NEURON + CoreNEURON / NMODL | Multicompartment neurons | Hines solver; a mechanism language compiled to fast kernels; decades of validated channel models |
| Arbor | Multicompartment neurons on HPC/GPU | Modern C++ GPU design; clean separation of model description from execution |
| NEST | Large point-neuron networks | Min-delay distributed communication; exascale scaling discipline |
| Brian2 / GeNN | Equation-defined spiking models | Equations-as-code with units checking; code generation for GPUs |
| Jaxley | Differentiable biophysics | Gradients through detailed neurons; parameter sharing |
| OpenWorm | Whole C. elegans | Open-science organization around one organism; worm body (Sibernetic) |
| Blue Brain Project tools (SONATA, BluePyOpt) | Detailed cortical models | Network file formats; large-scale parameter optimization workflows |
| NeuroML / LEMS, NWB, DANDI | Neuro standards and archives | Model exchange; data import standard; where benchmark data lives |
| GROMACS, LAMMPS, OpenMM | Molecular dynamics | Extreme kernel optimization; plugin force fields with required validation; reference vs. fast implementations |
| PETSc, SUNDIALS, Trilinos | Solver libraries | Use mature solvers rather than writing our own; Krylov and nonlinear solver interfaces |
| FEniCS / Firedrake | Finite elements | A DSL (UFL) that compiles math to kernels and derivatives |
| AMReX | Adaptive mesh refinement | Performance portability across GPU vendors; adaptive resolution |
| E3SM / CESM + CMIP | Coupled climate models | Component + coupler architecture; standard configurations; model intercomparison protocol |
| Einstein Toolkit (Cactus) | Numerical relativity | Plugin ("thorn") architecture so many groups contribute physics independently |
| Geant4 / ROOT | Particle physics | Swappable "physics lists" with validation suites; multi-decade collaboration governance |
| MuJoCo / MJX | Rigid-body physics | Body simulation (used directly) |
| CASP / Rosetta Commons | Protein structure | Blind prediction challenges that drove AlphaFold; multi-lab consortium governance |
| Folding@home / BOINC | Volunteer computing | Distributing embarrassingly parallel simulation work to the public |
| Astropy | Astronomy software | Core package + affiliated packages; community governance under NumFOCUS |
7.2 Layered architecture
┌───────────────────────────────────────────────────────────────────────────────┐
│ Applications cake CLI · Python API · notebooks · web viewer · bench client │
├───────────────────────────────────────────────────────────────────────────────┤
│ Configurations cake-worm · cake-fly-vnc · cake-fly-cns · cake-fish · … │
│ (organism release + fidelity choices + parameter posterior) │
├───────────────────────────────────────────────────────────────────────────────┤
│ Workflow campaigns · job scheduling · content-addressed cache · │
│ provenance capture │
├───────────────────────────────────────────────────────────────────────────────┤
│ Inference fitting · regulation endpoints · shadowing gradients · │
│ posteriors · identifiability · experiment design │
├───────────────────────────────────────────────────────────────────────────────┤
│ Coupler multirate scheduler · waveform relaxation · checkpoint/restart │
│ · conservation monitors · fidelity probes │
├───────────────────────────────────────────────────────────────────────────────┤
│ Components neural tissue (L1–L4, L7) · extracellular chemistry (L5, L6) · │
│ body · environment · observation (indicators, optics, ephys) │
├───────────────────────────────────────────────────────────────────────────────┤
│ CKL compiler CAKE Kinetics Language → kernels + derivatives + unit checks │
├───────────────────────────────────────────────────────────────────────────────┤
│ Runtime JAX/XLA · Pallas/Triton · C++/CUDA/HIP kernels · MPI · I/O │
└───────────────────────────────────────────────────────────────────────────────┘
▲ Data Commons: organism releases, priors, recordings, provenance graph
Core data model. Everything is built on four objects: Organism (immutable organism release), Model (compiled arrays for one fidelity configuration), Params (what inference changes) and State (per replica). Separating them, as MuJoCo separates mjModel from mjData, is what makes replica batching, caching and differentiation straightforward. Layout rules are in 06-compute.md §2.
Fidelity as configuration. A configuration chooses a fidelity level per component per cell type from the fidelity lattice (06-compute.md §3). The coupler can mix levels in one run (e.g. most fly neurons reduced, one circuit at full morphology), which is also how fidelity probes and multilevel Monte Carlo work.
Components and coupler. Each component (neural tissue, extracellular chemistry, body, environment, observation) owns its state and time step, and exchanges only declared fields with the coupler: membrane currents, concentrations, muscle activations, sensor signals, fluorescence. A component can be swapped (worm continuum body vs. resistive-force body; reference vs. reduced neurons) without touching the others, as climate models swap ocean or land components. The coupler implements the multirate scheduling and waveform relaxation of 06-compute.md §4.
CKL, the CAKE Kinetics Language. Every mechanism (channel, receptor, pump, release machinery, GPCR cascade, regulation rule) is written in a small declarative language, in the tradition of NMODL, LEMS, Brian equations and UFL:
mechanism Kv1_Shaker {
genes = [shk-1 @ C.elegans, Sh @ D.melanogaster]
ion = K
states = m, h
gbar ~ LogNormal(prior_from = "expression") [S/cm^2]
dm/dt = (m_inf(v) - m) / tau_m(v) * q10(T)
dh/dt = (h_inf(v) - h) / tau_h(v) * q10(T)
i = gbar * m^4 * h * (v - E_K)
provenance = { kinetics: "doi:…", temperature: 22 [degC] }
}
The compiler does dimensional analysis, generates forward kernels for each backend, generates derivatives (JVP/VJP) for inference, chooses integrators (Rush–Larsen for gating, implicit for stiff kinetics), and refuses to compile a mechanism without provenance metadata. Contributors add biology by writing CKL files with tests, not by writing GPU code. This is how the channel library (L2) grows to cover every channel and receptor gene in each organism.
Adopt first, build later. CKL, the coupler and custom kernels are not built up front. The first worm and Track B work runs on existing tools (Jaxley for differentiable multicompartment simulation, MJX for bodies, NMODL channel models translated as needed). Platform components are extracted from working science code once the worm milestones show what is actually needed, with interfaces fixed by CEPs at that point. This avoids spending the first years building infrastructure for requirements that may change when H* is decided.
7.3 Repository layout
A monorepo with separately versioned packages:
cake/
core/ # runtime, array/state layout, kernels, MPI, I/O, RNG (Philox)
ckl/ # language spec, parser, unit system, codegen backends, derivative generation
components/
neural/ # L1 cable · L2 channels · L3 synapses · L4 gap junctions · L7 regulation
chemistry/ # L5 volume transmission · L6 ions, pumps, glia, energy
body/ # worm continuum body · fly MJX body · muscles · proprioception
environment/ # fluids, surfaces, odor/thermal fields, visual scenes
observation/ # calcium/voltage indicators, microscope optics, ephys, behavior trackers
coupler/ # multirate scheduling, waveform relaxation, checkpoint/restart, monitors
reduce/ # neuron reduction, synapse lumping, diffusion homogenization
probes/ # shadow-neuron fidelity probes, adaptive refinement
surrogate/ # per-type regulation surrogates with verification hooks
infer/ # fitting, shadowing gradients, posteriors, identifiability, experiment design
workflow/ # campaign definitions, scheduler adapters (Slurm, cloud, volunteer), content-addressed cache
mechanisms/ # the CKL library: channels, receptors, pumps, cascades, regulation rules
ingest/ # FlyWire/CAVE, neuPrint, WormWiring, CeNGEN, Fly Cell Atlas, NWB/DANDI
organisms/ # organism release recipes (worm, fly-vnc, fly-cns, fish, …)
bench/ # Emulation Ladder, noise ceilings, frozen splits, submission client
necessity/ # layer ablation harness → Necessity Map
verification/ # manufactured solutions, convergence tests, cross-simulator checks
docs/ # tutorials · how-to guides · reference · explanations
spec/ # this design spec, decision log, CEPs
7.4 Standards and interoperability
| Purpose | Standard |
|---|---|
| Import existing models | NeuroML/LEMS, NMODL (translated to CKL), SONATA networks |
| Experimental data in | NWB, pulled from DANDI and other archives |
| Connectomes and morphology | CAVE/FlyWire, neuPrint, SWC, meshes (precomputed / OBJ) |
| Bodies | MuJoCo MJCF |
| Simulation output | Zarr (chunked, cloud-friendly) with CAKE output conventions (units, provenance, organism release ID), similar to CF conventions in climate science |
| HPC I/O | ADIOS2 for in-situ analysis of very large runs |
| Export | NeuroML export of reduced models, so other simulators can run CAKE results |
7.5 Verification and validation
Borrowed from computational physics practice. Verification means "are we solving the equations correctly"; validation means "are they the right equations".
- Code verification:
- Method of manufactured solutions for the cable, diffusion and mechanics solvers.
- Convergence-order tests under refinement of time step and mesh.
- Gradient checks (finite differences vs. JVP/VJP) for every CKL mechanism.
- Cross-simulator verification: identical models built in NEURON and Arbor must match CAKE's reference components to stated tolerances, run in CI.
- Invariants: conservation of charge, ion mass and mechanical energy are checked continuously, not just in tests.
- Regression: golden runs per organism configuration; any change in outputs beyond tolerance needs an explanation in the pull request.
- Performance regression: benchmark runs on real GPUs and HPC nodes on every release, with roofline comparisons. A slowdown is treated as a bug.
- Validation: the Emulation Ladder (03-evaluation.md), scored with its frozen task specifications; the Necessity Map; and the falsification tests F1–F7 with the synthetic specificity and power checks that precede them (02-hypotheses.md).
- Solver certification: every regulation endpoint records which validity conditions held (06-compute.md §5.5), and gradient checks against finite differences of certified solves run on a sampled subset in every fit.
- Uncertainty quantification: every published prediction carries a posterior interval, a measured local reduction error (fidelity probes), and that error propagated to the reported quantity (06-compute.md §9).
7.6 Organism releases and the Data Commons
An organism is released the way a genome assembly is: as a versioned, citable build.
cake-worm-2027.1= specific dataset versions (connectome, morphologies, transcriptome, peptide maps) + ingest code version + curation decisions, content-hashed.- Model releases are separate from code releases, as in climate science: a model version = code version + configuration + parameter posterior. Results always cite the model version.
- The Data Commons stores organism releases, priors, fitted posteriors and benchmark data with a provenance graph: every parameter traces back to whether it was measured, came from a prior or was inferred, and from which dataset.
- Licensing: code under Apache-2.0; organism releases, posteriors and benchmarks under CC-BY 4.0 (subject to upstream dataset licenses, recorded in provenance).
7.7 Community benchmarks: NeuroMIP and blind challenges
Two community instruments, open to any nervous-system model, not just CAKE:
- NeuroMIP (Nervous-system Model Intercomparison Project), modeled on CMIP:
- a standard set of experiments (stimulation protocols, mutants, environments) with fixed input and output formats;
- every participating model runs the same protocols, and results are compared on the Emulation Ladder;
- the goal is to learn which modeling choices matter, across groups and methods.
- Blind prediction rounds, modeled on CASP:
- partner labs collect new perturbation data (new optogenetic pairs, new mutants, new behaviors) and hold it back;
- modelers submit predictions before release; scoring is against the noise ceiling;
- repeated every 1–2 years. This makes progress measurable and makes overfitting to public data impossible.
The active-learning queue (05-inference.md) feeds both: the experiments CAKE says are most informative are good candidates for blind rounds.
7.8 Workflow, campaigns and caching
Fitting the fly means hundreds of thousands of jobs: replicas, burn-ins, probes, benchmark runs. Like large simulation projects with their own campaign tooling, CAKE needs a workflow layer as a first-class part of the code, not ad-hoc scripts.
- Campaigns are declarative: organism release, fidelity configuration, inference recipe, data, compute budget. A campaign file plus the code version fully determines the result.
- Content-addressed cache (06-compute.md §8): before any job runs, the scheduler checks whether its result already exists under the same input hash, across the whole team and all past campaigns.
- Scheduler adapters for Slurm/HPC batch systems, cloud and (later) volunteer computing, behind one interface.
- Provenance capture is automatic: every output records the campaign, inputs, code version and cache keys that produced it.
- Budget accounting: each campaign reports GPU-hours by stage against its declared budget, so cost estimates in 06-compute.md §10 get replaced with measurements over time.
7.9 Deployment and computing
- Scales: laptop (worm, reduced fly sub-circuits) → single GPU (reduced fly brain) → cluster (ensembles, inference) → leadership-class HPC (reference-fidelity fly, Necessity Map sweeps).
- Vendor portability: NVIDIA, AMD and Intel GPUs, since the largest supercomputers use all three. JAX/XLA covers most of this; hand-written kernels need CUDA and HIP versions behind one interface, with SYCL considered later.
- Packaging: pip/conda for users; Spack for HPC; Apptainer containers for reproducible runs.
- Checkpoint/restart for every component, so long runs survive node failures and can be resumed or branched.
- Compute allocations: apply to national HPC programs (e.g. INCITE and ACCESS in the US, EuroHPC in Europe) and to cloud research credits.
- Volunteer computing (later phase): replicas, fidelity checks and ensemble members are independent, whole-brain-per-device jobs (06-compute.md §4.1), which suits a Folding@home/BOINC-style public network. Results from untrusted machines are checked by redundant computation before use: bitwise within a reproducibility class, statistically across classes (06-compute.md §4.6).
Principles that apply everywhere:
- Provenance on every parameter.
- Held-out benchmark splits frozen and hashed before any model fitting.
- Every reduced model validated against higher fidelity before use, with error monitored in production (fidelity probes).
- Reproducible stochastic runs: identical random draws everywhere (counter-based RNG keyed by replica, neuron, synapse and step); bitwise-identical trajectories within a reproducibility class; statistically equivalent results across classes (06-compute.md §4.6).
- Every result traceable to a campaign file and code version.