OPEN RESEARCH PROGRAM 2026

Beyond reflexes.
Toward a mind.

We are evolving a fly brain from simple collision avoidance into robust, adaptive embodied intelligence—one challenge, one learner, and one honest experiment at a time.

SENSORY INPUTADAPTIVE POLICY↗ EVOLUTIONARY SEARCH
09learner families
05embodied challenges
600FlyWire proxy neurons
hypotheses to test
01 / THE MISSION

A tiny nervous system is a
serious research question.

The fruit fly has a compact brain, rich behavior, and a body that has to make decisions in the real world. That makes it a sharp starting point for asking what intelligence needs—and what it can learn.

Learning To Fly is a reproducible research workspace for evolving controllers in progressively harder environments. The primary track stays grounded in the fly’s immediate needs: avoid walls, find resources, recover from damage, and adapt when the world changes.

02 / THE SEARCH SPACE

One world.
Many ways to learn.

Every approach gets the same sensors, actions, lifetime, online updates, compute budget, and held-out shifts. The question is not which name sounds smartest—it is which mechanism survives contact with the world.

CONTROL

Frozen baseline

A non-learning reference that tells us whether adaptation is doing real work.

ONLINE LEARNING

CLRM

A compact continual learner with matched observation, action, and update budgets.

ONLINE LEARNING

IntuiKit-CLRM

A parallel CLRM formulation kept separate so intuitions can be tested fairly.

VECTOR-FIELD PROXY

IntuiCell-inspired

Decentralized recurrent units with local attractor dynamics and prediction-error plasticity—an auditable public-research-inspired proxy, not the proprietary system.

PREDICTIVE LATENTS

JEPA-style proxy

Predicts future latent state and learns useful representations without a full external JEPA.

CAUSAL MEMORY

LLM-style proxy

A bounded causal token-memory controller that tests sequence prediction under tight budgets.

SELF-REVISION

Gödel-style proxy

A constrained self-rewriting controller exploring whether useful policy changes can be discovered online.

CONTINUOUS TIME

Liquid LTC

Sparse liquid time-constant dynamics with adaptive rates and local plasticity.

CONTINUOUS TIME

Liquid CfC

A closed-form continuous-time sibling probing fast, bounded recurrent adaptation.

03 / THE SEARCH ATLAS

The brain is
not one variable.

Every candidate is a point in a much larger space. This projection shows the axes we vary—mechanism, wiring, dynamics, memory, and plasticity.

VIEW
GENERATION
Random hypotheses form a cloud while evolution traces a selective path through the same space. Higher-dimensional search axes radiate from a central controller node.FLYBRAINFAMILYJEPA · INTUICELL · LTC · CfCTOPOLOGYFLYWIRE · RANDOMTIME CONSTANT0.05 — 8.0GATE SCALEINPUT-DEPENDENTRECURRENT GAINSPARSE MEMORYINPUT / OUTPUTSENSOR → MOTORLEARNING RATEONLINE PLASTICITYMEMORY BUDGET4 — 64 UNITS↗ PROJECTION OF A HIGHER-DIMENSIONAL SEARCH SPACE
RANDOM HYPOTHESES EVOLVED PATHGEN 1
03 / THE METHOD

Evolve.
Stress-test.
Learn.

Evolution proposes the learner and its parameters. The environment supplies the pressure. Held-out challenges decide whether an apparent improvement is real.

01

Compose the world

Food, water, procedural mazes, relocation, sensor damage, and multi-challenge carryover create a richer behavioral substrate.

02

Search across mechanisms

Evolution mutates families, memory, rates, gains, and structure under bounded, matched budgets.

03

Separate signal from story

Controls, ablations, disjoint seeds, and honest negative results keep the primary claim gated.

THE FLIGHT PLAN

From avoiding walls
to finding a way home.

01Collision avoidance
02Homeostatic foraging
03Relocation adaptation
04Damage recovery
05Multi-challenge carryover
CURRENT READOUT

We found the blocker, and then we cleared it.

A ceiling audit found 0.518 of score that roughly forty experiments had been contesting a hundredth at a time. Ablating privilege located it in two super-additive levers — knowing where you are, and knowing the maze — and a sweep put a number on the first: route planning is worth +0.18 when position error is under 0.15 cells and −0.18 once it passes 0.23. Our dead reckoning sat at 0.67. Giving the fly a lattice of known landmarks to correct against drops it to 0.048, and score rises from 0.214 to 0.419 — the first arm in the programme to clear its preregistered gate by more than its cohort can resolve.

PRIMARY GATEPASSED & REPLICATED+0.163 and +0.114 on a score rebuilt to measure foraging
PRIMARY TRACK · LATEST

Knowing where you are was the whole problem

0.048 cells position error+0.205 score, CI [+0.168, +0.243]184/256 worlds reach food after it moves

Three anchors clustered in one corner were not enough; spread over a lattice, a small correction on every step they are visible arrests drift outright — late error stops exceeding mean error. Matched resource contacts triple, mobility goes 0.762 → 0.964, and the gain is not the extra sensor: an arm carrying the lattice while ignoring it scores −0.027, its interval crossing zero. The cliff predicted +0.18 at this accuracy. We measured +0.205 — a forecast from privileged sweeps, met by an arm that reads no privileged pose. A second untouched 256-world split passed every criterion again at +0.174, CI95 [+0.130, +0.218], with the carrying control at +0.001: nine landmarks you ignore are indistinguishable from not having them.

LEARNER GATE · CLOSED

We gave learning its fair test, and it closed the question

−0.032 plastic vs frozen, drifting frame−0.001 plastic vs frozen, accurate frame~30 historical nulls, now explained

Every learner-versus-frozen comparison this programme ran — roughly thirty — drove the residual from a pose drifting 0.67 cells, so we preregistered a two-by-two crossing the frame with the plasticity on 128 untouched worlds. The regimes separated reliably, but not the way the exploratory sixteen-world estimate claimed: plasticity harms the drifting navigator (−0.032, CI95 [−0.057, −0.009]) and trended positive on the accurate one — the first plastic-over-frozen comparison whose official interval cleared zero. We registered that nomination’s confirmation before running it, sized at 512 fresh worlds from the measured variance. It came back −0.0006, with an official interval tight enough to exclude the nominated effect itself. Learning ran identically in every cohort — 178 reinforcements per world, readout moved 0.19 — and did nothing. The track is closed by its own registered rule: this residual family reliably damages a navigator that is lost and cannot improve one that knows where it is. Everything that moved this benchmark is structural.

LANDMARK SPARSITY

Placement, not count

Thinning the field shows the fly does not need nine landmarks — it needs mutually visible ones. A fix requires two anchors in range at once, so four at the edge midpoints (spacing 8.5) recover +0.137 of the +0.20, while four at the corners (spacing 12) produce zero fixes and score exactly the uncorrected baseline. Same count, opposite outcome. Five anchors hold position error at 0.133 cells, still inside the budget, on a third of the observations.

CORRECTION

We checked what our own benchmark measures, and rebuilt it

69% of the score is collision rate5.4% is matched consumption

Across the full competence range — wall follower to omniscient oracle — collision rate accounts for 69% of everything the score can discriminate, and it is counted twice: survival is 1 − collisions/steps and the penalty is −collisions/steps. Matched consumption is 5.4% despite carrying a weight of 8, because a weight cannot rescue a term whose entire range from worst to best agent is 0.073. So we built a score that measures foraging — collisions counted once, every term normalised by the range actually achievable — and re-scored our own headline on its full cohorts. It holds, at +0.163 and +0.114 with both intervals clear of zero, against +0.205 and +0.174 officially: the collision term was inflating it and the honest figure is about a third smaller. Resource acquisition alone gives +0.202 [+0.085, +0.312] and +0.110 [−0.003, +0.221], so the fly does forage better — while the integrated deficit does not move at all, because that term is set by travel time and decay rather than by anything the fly knows.

CLAIM BOUNDARY

This is an observability result

The lattice is disclosed added sensing — nine known immutable landmarks, seen egocentrically within six cells under the usual noise. It shows that given a distributed landmark field, continuous correction plus planning converts most of the pose lever, and that every earlier anchor arm failed for want of coverage rather than mechanism. It is not a claim that the original fourteen-channel fly can localize itself. Replication and sparsity are both done; attempts to earn the same reference from junctions or resource cues inherit the fly’s own drift and do not close that gap.

HYPER ANALOGUE

The architecture is buildable; the missing observation is not

We built the smallest reproducible version of HYPER’s public idea: a structured world model that asks whether each action turned normally or inverted and whether it moved or was blocked, then checks those hypotheses against the next wall scan. Across 64 untouched worlds it reached only 50.1%transition accuracy versus 93.4% for the existing clearance heuristic, and pose error grew to 7.40 cells. It receives no navigation authority. Structured prediction remains sensible, but these eight sparse rays do not contain the persistent scene information such a model needs.

JOINT ESTIMATION

The first mechanism to reach under the cliff — and its reach is short

Scan matching against the fly’s own map failed as theory predicts: the map is built from the drifting pose, so it encodes the drift. Referencing the three immutable anchors continuously instead — a small correction every step rather than one late jump — cut position error to 0.201 cells and gained +0.112 exploratorily. On untouched worlds it closed: +0.0224, CI95 [−0.054, +0.097]. Error there fell only to 0.361, still the wrong side of the cliff, because all three anchors sit in one corner and correction quality tracks how long the fly spends near them.

STATISTICAL POWER

Most of our negatives are unresolved, not refuted

Per-world score deviation is about 0.43, so a 128-world cohort resolves nothing below roughly 0.08 — and almost every mechanism this programme has tested sits under that bound. That reframes the record: the honest reading of most closed arms is “too small to see”, not “shown to fail”. Future cohorts get sized from this variance.

SUBSTRATE TEST

The challenge we built everything around costs 0.016

Deleting the actuator inversion entirely drops heading error to 0.000 and position error from 0.665 to 0.225 cells — but the fly gains only +0.016, and the reachable gap barely moves (0.501 to 0.485). The existing recovery machinery already handles the inversion; what remains is the ordinary odometry floor, which sits above the planning cliff in every episode either way.

WHY EVERY FIX STALLS

The two levers only work together

Giving the fly a real occupancy map — tracing every ray instead of reading one bit per direction — moved score +0.0225 against +0.0221 for planning alone. A richer map cannot pay while the pose that places its cells drifts by 0.665 cells, and better pose cannot pay while the map is thin. Six separate single-lever mechanisms have now returned a hundredth or less each.

DRIFT ANATOMY

One moment destroys the fly’s sense of place

Heading error is exactly 0.000 for the first 300 steps and 0.310 rad after. Dead reckoning is not weak — it is destroyed by the ~28 steps between the actuator inverting and the fly confirming it, every turn of which is integrated with the wrong sign and never repaired.

ACTIVE PROBING

Faster detection, and it still was not enough

Letting the fly deliberately test its own steering when it suspects a change cut confirmation from step 322.4 to 313.0. Score moved −0.0247, CI95 [−0.082, +0.032], so it closes by rule. It can now tell that its body changed sooner, but still not when — and without the onset it cannot undo the turns it got wrong.

HEADROOM AUDIT

Knowing where the food is barely matters

A fly with the true maze but only the six-cell resource cue scores 0.748 against 0.762 for one that is also told the exact resource coordinates. The resource memory the programme has been refining is worth about 0.014; the maze and the fly’s place in it are worth the rest.

FRONTIER PLANNING

Better planning on a drifting map does not help

Replacing the local least-visited junction rule with shortest-path routing to the nearest unvisited cell moved score by +0.0087, CI95 [−0.021, +0.038]. Closed by rule — and predicted by the cliff, since the map it plans over is drifting.

LEARNER FAMILIES

IntuiCell proxy leads the matrix

Five fixed-family searches average 0.734 held-out. This is an auditable public-research-inspired proxy, not the proprietary IntuiCell system.

SEARCH POLICY

Diversity helps coverage, not reliability

Five-bin quality diversity improves archive occupancy, yet matched lottery still leads on mean score and every superiority interval crosses zero.

STRATEGY COMBINATION

Routing beats blending, not JEPA

Winner-take-all routing scored 0.743 and beat blending in all eight runs, while frozen JEPA remained stronger at 0.783.

Experiment history, grouped by finding
Learner-frame two-by-two and replicationPlasticity reliably harms on a drifting frame and does nothing on an accurate one; the learning track closed by its own registered rule.
Distributed anchor latticeFirst replicated positive: pose error 0.048 cells, +0.205 then +0.174 across two untouched 256-world splits.
Pose-accuracy cliffPlanning is worth +0.18 below 0.15 cells of position error and −0.18 above 0.23.
Disclosed-privilege headroom audit0.518 of score is reachable; resource knowledge accounts for only 0.014 of it.
Frontier route planningGlobal routing over a drifting map did not clear zero.
Coverage-memory restartA specific but sparse detector could not support fresh exploration.
Immutable resource baselineSpecificity passed; sparse late detection could not support harmful routing.
Frame-aligned landmark refreshCorrect transforms could not repair a drift-contaminated landmark baseline.
Localized landmark refreshReset specificity improved sharply; stale topology could not exploit it.
Powered pose reanchorAccurate correction disrupted the old topology frame and reliably hurt flight.
Distributed three-anchor localizationFirst pose mechanism to clear efficacy, safety, and coverage together.
Localization replicationPrecision and safety replicated; fixed-pair visibility remained below gate.
Motion-compensated beacon offsetsFirst safe, reliably beneficial pose correction; coverage missed by one world.
Temporal beacon confidenceStationary-pose consensus rejected legitimate motion; motion compensation is now the frontier.
Dual immutable beaconsFull pose became observable in mean; one-shot safety still failed.
Immutable home beaconOne anchor exposed heading drift instead of resolving full pose.
Global transform consensusMultiple anchors filtered aliases substantially, but remained unsafe.
Trajectory identityLonger action-conditioned sequences remained globally ambiguous.
Wall-loop localizationLocal ray sequences were heavily aliased; no correction authority was granted.
CMA-ES and JEPASearch efficiency improved reliably; equal reward composition did not.
Landmark resetCoordinate consensus confused dead-reckoning drift with actual relocation.
Relocation transferSafe traversal persisted, but post-relocation contact did not clear its gates.
Pulse topologyThe unchanged 256-world replication passed every robust-flight gate.
Learning mechanismLow-authority task-weighted learning passed frozen and wall controls.
Powered searchEvolution tied lottery and default; resource contact remained zero.
Resource observabilityGlobal cues did not help; a shortest-path reference succeeded in 31/32 worlds.
Cue navigationFusing cues with wall avoidance improved survival, but not resource contact.
Hand-rule explorationFirst non-privileged contacts: 2/32 food, zero water; gate not passed.
Route memoryExact-odometry memory reached 26/32 worlds; noisy integration comes next.
Path integrationObservation-only dead reckoning reached both resources in 22/32 worlds.
Path learningAlways-on plasticity cut contact worlds from 51/64 to 37/64; gate failed.
Risk-gated learningContacts recovered to 37/64, but score and collision gates still failed.
Aligned eligibilityTrace semantics are fixed; plasticity still tied or trailed both controls.
Forward rewardBoth resources remained, but score trailed all controls within uncertainty.
Marginal creditStrongest learned arm at +0.329; all-control confidence gate still failed.
Predictive navigationNext-ray surprise beat its frozen motor twin and reliably increased contacts.
Trace horizonDecay 0.90 stayed best, but none of four horizons passed the joint gate.
Trace replicationA 128-world retest preserved small gains, but none was reliable.
Immediate eligibilityRemoving the latent trace reliably hurt; the trace remains replication-only.
Collision attributionCounterfactual blame helped slightly, but not reliably or against controls.
Residual authorityGain 0.10 improved contacts over frozen, but failed the joint nomination rule.
Reward noise floorThree deadbands left contacts unchanged and failed to match path-only.
Exploration scheduleA 150-step cutoff led numerically, but failed both confidence requirements.
Confidence gateResidual magnitude misidentified harmful saturated turns as confidence.
Reward evidenceSigned EMA cancellation left all three exploration gates completely inert.
Reward-sign consistencyScale-free gates activated briefly but reliably trailed the clock cutoff.
Replenishment eventsTwo observed jumps beat path-only score, but failed the full learning gate.
Need-triggered resumptionRising deficit resumed exploration, but did not improve score or contact.
Residual authorityGains through 0.30 improved dynamics but not plastic-over-frozen performance.
Marginal slope creditAction-scale normalization reliably beat raw credit, but not frozen or path.
Action-signed creditCorrected eligibility led every mean, but all required intervals crossed zero.
Intui mechanism ablationReadout learning reproduced the harm; recurrent plasticity was nearly inert.
Primary Intui proxyPost-remap plasticity changed the readout, then trailed replay and frozen Intui.
Resource landmarksExact reachable replay is safe; using its landmarks to replace pose is not.
Learner familiesIntuiCell-inspired proxies lead locally, without a proprietary-system claim.
Quality diversityArchive coverage improved, but no reliable advantage over lottery emerged.
Strategy combinationRouting beat blending, but neither beat the strongest frozen component.