flybench

Does the simulated fly still do the things a real fly does?

An open benchmark for whole-brain fruit fly simulations: 11 behaviours from the literature, one score per model. The reference model already fails half of them. Beat it.

behaviours tested
11
5 core · 6 hard
best hard-tier score
0.69
reference model — beatable
working gain window
0.4–0.45
published value: 1.0
brain firing at gain 1.0
34%
on one taste of sugar

Three ways in

git clone https://github.com/brandoncho369/flybench && cd flybench && pip install -e ".[dev]"
flybench toy && flybench run        # synthetic brain, 10 s — then see the README for the real one

What it found so far

The published model uses gain 1.0. FlyWire re-predicted every synapse in July 2025 and the connections got heavier; nobody re-tuned. On today's data that gain makes 34% of the brain fire at a taste of sugar. Sweep the one knob:

core reflexes passed
0.000.501.000.30.350.40.450.50.71gain →

shaded: gain 0.40.45, every core reflex passes

share of brain firing (peak)
0%20%40%0.30.350.40.450.50.71gain →

shaded: gain 0.40.45, every core reflex passes

gaincorehardbrain firingwhat happens
0.30
0.67
0.63
9%taste never reaches the proboscis
0.35
0.67
0.53
10%taste never reaches the proboscis
0.40
1.00
0.69
11%passes every known reflex
0.45
1.00
0.69
11%passes every known reflex
0.50
0.83
0.74
12%too much of the brain fires
0.70
0.87
0.82
27%too much of the brain fires
1.00 · Shiu 2024
0.77
0.71
34%too much of the brain fires

Inside the window the model still fails most of the hard tier: no dose response, no adaptation, no lateral inhibition, no selectivity. Each one is a concrete thing your model could add.

Leaderboard

ranked by core, then hard · snapshot 2026-09-11

A model must reproduce the known reflexes before its hard-tier wins count. Every row so far is the reference LIF model at a different gain on FlyWire v783. Submit yours with a pull request.

#runmodelcorehard core tasks hard tasks
1LIF gain 0.4LIFSimulator1.000.69
2LIF gain 0.45LIFSimulator1.000.69
3LIF gain 0.7LIFSimulator0.870.82
4LIF gain 0.5LIFSimulator0.830.74
5LIF gain 1.0LIFSimulator0.770.71
6LIF gain 0.3LIFSimulator0.670.63
7LIF gain 0.35LIFSimulator0.670.53

Filled dot = task passed; hollow = failed (hover for the task name and partial score). Full per-check values in LEADERBOARD.md and results/*.json.

The tasks

Core

Reflexes the reference model must reproduce. A model that fails these is broken.

Silence in, silence out

stability

With no sensory drive the network must not spontaneously light up. A model whose gain is too high "passes" every reflex because everything fires; this task is the guard against that. (Real flies do have spontaneous activity — the LIF model has none by construction, so quiet is the correct answer here.)

Shiu et al. 2024, Nature — model is silent without input by design

Sugar → proboscis extension

sugar_to_proboscis

The canonical reflex. Activating sugar-sensing gustatory receptor neurons should drive MN9, the motor neuron that extends the proboscis. This is the headline result of Shiu et al. 2024 and a behaviour every fly lab can reproduce with a drop of sucrose.

Dethier 1976; Shiu et al. 2024 (Nature) Fig. 2

Bitter suppresses the sugar response

bitter_suppression

Co-activating bitter GRNs with sugar GRNs should reduce MN9 output relative to sugar alone. Flies reject sucrose laced with quinine; the connectome should show why (bitter-driven inhibition onto the sugar pathway).

Shiu et al. 2024 (Nature) Fig. 3; Jaeger et al. 2018

Looming → giant fiber escape

looming_to_giant_fiber

LPLC2 and LC4 are the visual projection neurons that detect an expanding (looming) object. They converge on the Giant Fiber, the descending neuron that triggers the takeoff escape. Drive the detectors; the GF should fire.

von Reyn et al. 2014; Ache et al. 2019; Shiu et al. 2024

Bitter alone does not extend the proboscis

taste_specificity

A negative control. Bitter GRNs alone should not drive MN9; if they do, the model has lost the sign of its inhibitory synapses or the gain is high enough that any input reaches any output.

Shiu et al. 2024 (Nature)

Hard

Behaviours a wiring diagram plus five constants is not expected to give you. The reference model fails most of these on purpose; they are the research agenda.

Proboscis response scales with sugar intensity

dose_response

Real flies extend the proboscis more reliably and more strongly as sucrose concentration rises; GRN firing rate encodes concentration. A model whose MN9 output is all-or-nothing (saturating at the refractory limit the moment any input arrives) has lost that gradation. Weak drive (20 Hz on the sugar GRNs) should give a clearly smaller MN9 response than strong drive (100 Hz), and both should be above baseline.

Dethier 1976; Dahanukar et al. 2007 (Neuron); Shiu et al. 2024

Repeated sugar gives a smaller second response

adaptation

Sensory responses adapt: a second identical sugar pulse shortly after the first evokes a weaker proboscis response (short-term adaptation in GRNs and downstream circuits; over longer timescales, PER habituation). A memoryless point-neuron model cannot do this by construction — every pulse looks like the first — so the reference LIF is expected to fail here. That is the point: this task measures something the wiring diagram alone does not give you.

Duerr & Quinn 1982 (PNAS, PER habituation); Paranjpe et al. 2012 (J. Neurosci.)

One odour channel stays one channel

olfactory_sparse_coding

Olfactory receptor neurons of a single glomerulus (DA1, the cVA pheromone channel) synapse onto their own projection neurons. Lateral inhibition by local neurons keeps the antennal lobe output sparse: driving DA1 ORNs should fire DA1 PNs strongly while most PNs of the other ~50 glomeruli stay quiet, and the third-order lateral horn should not light up wholesale. A model with the wrong excitation/inhibition balance turns one odour into "all odours".

Olsen & Wilson 2008 (Nature); Wilson 2013 (Annu. Rev. Neurosci.); Kohl et al. 2013

Taste does not trigger escape; looming does not trigger feeding

crosstalk

Negative controls across circuits. Sugar on the labellum should not fire the Giant Fiber, and a looming shadow should not extend the proboscis. Both circuits pass their own reflex tests in the core tier; this task asks whether they stay separate, which fails when the global gain is high enough that activity spreads through shared interneurons.

von Reyn et al. 2014; Shiu et al. 2024

Looming recruits the takeoff ensemble, not the whole descending tract

looming_dn_ensemble

The Giant Fiber is not alone: looming also drives DNp02, DNp11 and other descending neurons that prepare the takeoff (leg extension, wing raise), and the response is selective — a few dozen of the ~1300 descending neurons, not a general alarm. Checks that the named looming DNs fire and that the descending population as a whole stays mostly silent.

Ache et al. 2019 (Curr. Biol.); Namiki et al. 2018 (eLife); Dombrovski et al. 2023 (Nature)

A uniform flash is not a looming object

flash_is_not_loom

LPLC2/LC4 respond to expanding edges, not to the whole eye lighting up. Driving every photoreceptor at once (a full-field flash) should therefore not fire the Giant Fiber the way looming does. In a pure connectome model motion selectivity has to emerge from the lamina/medulla circuitry and the synaptic delays alone — no dendritic nonlinearities, no adaptation — so this is expected to be hard. Failure here means the model's visual front end is a brightness detector, which is worth knowing before wiring it to a game.

von Reyn et al. 2014 (Nat. Neurosci.); Card & Dickinson 2008 (Curr. Biol.); Klapoetke et al. 2017 (Nature)