Frozen baseline
A non-learning reference that tells us whether adaptation is doing real work.
We are evolving a fly brain from simple collision avoidance into robust, adaptive embodied intelligence—one challenge, one learner, and one honest experiment at a time.
The fruit fly has a compact brain, rich behavior, and a body that has to make decisions in the real world. That makes it a sharp starting point for asking what intelligence needs—and what it can learn.
Learning To Fly is a reproducible research workspace for evolving controllers in progressively harder environments. The primary track stays grounded in the fly’s immediate needs: avoid walls, find resources, recover from damage, and adapt when the world changes.
Every approach gets the same sensors, actions, lifetime, online updates, compute budget, and held-out shifts. The question is not which name sounds smartest—it is which mechanism survives contact with the world.
A non-learning reference that tells us whether adaptation is doing real work.
A compact continual learner with matched observation, action, and update budgets.
A parallel CLRM formulation kept separate so intuitions can be tested fairly.
Decentralized recurrent units with local attractor dynamics and prediction-error plasticity—an auditable public-research-inspired proxy, not the proprietary system.
Predicts future latent state and learns useful representations without a full external JEPA.
A bounded causal token-memory controller that tests sequence prediction under tight budgets.
A constrained self-rewriting controller exploring whether useful policy changes can be discovered online.
Sparse liquid time-constant dynamics with adaptive rates and local plasticity.
A closed-form continuous-time sibling probing fast, bounded recurrent adaptation.
Every candidate is a point in a much larger space. This projection shows the axes we vary—mechanism, wiring, dynamics, memory, and plasticity.
Evolution proposes the learner and its parameters. The environment supplies the pressure. Held-out challenges decide whether an apparent improvement is real.
Food, water, procedural mazes, relocation, sensor damage, and multi-challenge carryover create a richer behavioral substrate.
Evolution mutates families, memory, rates, gains, and structure under bounded, matched budgets.
Controls, ablations, disjoint seeds, and honest negative results keep the primary claim gated.
A ceiling audit found 0.518 of score that roughly forty experiments had been contesting a hundredth at a time. Ablating privilege located it in two super-additive levers — knowing where you are, and knowing the maze — and a sweep put a number on the first: route planning is worth +0.18 when position error is under 0.15 cells and −0.18 once it passes 0.23. Our dead reckoning sat at 0.67. Giving the fly a lattice of known landmarks to correct against drops it to 0.048, and score rises from 0.214 to 0.419 — the first arm in the programme to clear its preregistered gate by more than its cohort can resolve.
Three anchors clustered in one corner were not enough; spread over a lattice, a small correction on every step they are visible arrests drift outright — late error stops exceeding mean error. Matched resource contacts triple, mobility goes 0.762 → 0.964, and the gain is not the extra sensor: an arm carrying the lattice while ignoring it scores −0.027, its interval crossing zero. The cliff predicted +0.18 at this accuracy. We measured +0.205 — a forecast from privileged sweeps, met by an arm that reads no privileged pose. A second untouched 256-world split passed every criterion again at +0.174, CI95 [+0.130, +0.218], with the carrying control at +0.001: nine landmarks you ignore are indistinguishable from not having them.
Every learner-versus-frozen comparison this programme ran — roughly thirty — drove the residual from a pose drifting 0.67 cells, so we preregistered a two-by-two crossing the frame with the plasticity on 128 untouched worlds. The regimes separated reliably, but not the way the exploratory sixteen-world estimate claimed: plasticity harms the drifting navigator (−0.032, CI95 [−0.057, −0.009]) and trended positive on the accurate one — the first plastic-over-frozen comparison whose official interval cleared zero. We registered that nomination’s confirmation before running it, sized at 512 fresh worlds from the measured variance. It came back −0.0006, with an official interval tight enough to exclude the nominated effect itself. Learning ran identically in every cohort — 178 reinforcements per world, readout moved 0.19 — and did nothing. The track is closed by its own registered rule: this residual family reliably damages a navigator that is lost and cannot improve one that knows where it is. Everything that moved this benchmark is structural.
Thinning the field shows the fly does not need nine landmarks — it needs mutually visible ones. A fix requires two anchors in range at once, so four at the edge midpoints (spacing 8.5) recover +0.137 of the +0.20, while four at the corners (spacing 12) produce zero fixes and score exactly the uncorrected baseline. Same count, opposite outcome. Five anchors hold position error at 0.133 cells, still inside the budget, on a third of the observations.
Across the full competence range — wall follower to omniscient oracle — collision rate accounts for 69% of everything the score can discriminate, and it is counted twice: survival is 1 − collisions/steps and the penalty is −collisions/steps. Matched consumption is 5.4% despite carrying a weight of 8, because a weight cannot rescue a term whose entire range from worst to best agent is 0.073. So we built a score that measures foraging — collisions counted once, every term normalised by the range actually achievable — and re-scored our own headline on its full cohorts. It holds, at +0.163 and +0.114 with both intervals clear of zero, against +0.205 and +0.174 officially: the collision term was inflating it and the honest figure is about a third smaller. Resource acquisition alone gives +0.202 [+0.085, +0.312] and +0.110 [−0.003, +0.221], so the fly does forage better — while the integrated deficit does not move at all, because that term is set by travel time and decay rather than by anything the fly knows.
The lattice is disclosed added sensing — nine known immutable landmarks, seen egocentrically within six cells under the usual noise. It shows that given a distributed landmark field, continuous correction plus planning converts most of the pose lever, and that every earlier anchor arm failed for want of coverage rather than mechanism. It is not a claim that the original fourteen-channel fly can localize itself. Replication and sparsity are both done; attempts to earn the same reference from junctions or resource cues inherit the fly’s own drift and do not close that gap.
We built the smallest reproducible version of HYPER’s public idea: a structured world model that asks whether each action turned normally or inverted and whether it moved or was blocked, then checks those hypotheses against the next wall scan. Across 64 untouched worlds it reached only 50.1%transition accuracy versus 93.4% for the existing clearance heuristic, and pose error grew to 7.40 cells. It receives no navigation authority. Structured prediction remains sensible, but these eight sparse rays do not contain the persistent scene information such a model needs.
Scan matching against the fly’s own map failed as theory predicts: the map is built from the drifting pose, so it encodes the drift. Referencing the three immutable anchors continuously instead — a small correction every step rather than one late jump — cut position error to 0.201 cells and gained +0.112 exploratorily. On untouched worlds it closed: +0.0224, CI95 [−0.054, +0.097]. Error there fell only to 0.361, still the wrong side of the cliff, because all three anchors sit in one corner and correction quality tracks how long the fly spends near them.
Per-world score deviation is about 0.43, so a 128-world cohort resolves nothing below roughly 0.08 — and almost every mechanism this programme has tested sits under that bound. That reframes the record: the honest reading of most closed arms is “too small to see”, not “shown to fail”. Future cohorts get sized from this variance.
Deleting the actuator inversion entirely drops heading error to 0.000 and position error from 0.665 to 0.225 cells — but the fly gains only +0.016, and the reachable gap barely moves (0.501 to 0.485). The existing recovery machinery already handles the inversion; what remains is the ordinary odometry floor, which sits above the planning cliff in every episode either way.
Giving the fly a real occupancy map — tracing every ray instead of reading one bit per direction — moved score +0.0225 against +0.0221 for planning alone. A richer map cannot pay while the pose that places its cells drifts by 0.665 cells, and better pose cannot pay while the map is thin. Six separate single-lever mechanisms have now returned a hundredth or less each.
Heading error is exactly 0.000 for the first 300 steps and 0.310 rad after. Dead reckoning is not weak — it is destroyed by the ~28 steps between the actuator inverting and the fly confirming it, every turn of which is integrated with the wrong sign and never repaired.
Letting the fly deliberately test its own steering when it suspects a change cut confirmation from step 322.4 to 313.0. Score moved −0.0247, CI95 [−0.082, +0.032], so it closes by rule. It can now tell that its body changed sooner, but still not when — and without the onset it cannot undo the turns it got wrong.
A fly with the true maze but only the six-cell resource cue scores 0.748 against 0.762 for one that is also told the exact resource coordinates. The resource memory the programme has been refining is worth about 0.014; the maze and the fly’s place in it are worth the rest.
Replacing the local least-visited junction rule with shortest-path routing to the nearest unvisited cell moved score by +0.0087, CI95 [−0.021, +0.038]. Closed by rule — and predicted by the cliff, since the map it plans over is drifting.
Five fixed-family searches average 0.734 held-out. This is an auditable public-research-inspired proxy, not the proprietary IntuiCell system.
Five-bin quality diversity improves archive occupancy, yet matched lottery still leads on mean score and every superiority interval crosses zero.
Winner-take-all routing scored 0.743 and beat blending in all eight runs, while frozen JEPA remained stronger at 0.783.
Our new IntuiCell-inspired vector-field proxy currently leads the broad family matrix at 0.602, ahead of the JEPA-style proxy at 0.565. Functional FlyWire motifs improve evolved liquid topologies over degree-only selection, but the frozen motif winners did not generalize uniformly to held-out seeds. Matched random-dynamics controls still win the aggregate topology comparison. Three search replicates suggest family-specific preferences, not a universal topology winner. A 16-member random ensemble also led both liquid families on mean score, with wide variance—an effect worth studying, not a shortcut to a conclusion. Our latest liquid result is more encouraging: globally initialized evolution beat default-centered evolution on 7 of 8 held-out replicates for both CfC and LTC, improving held-out means from −0.445 to −0.067 and −0.455 to −0.187 respectively. But at larger matched budgets, broad global lottery still won 7 of 8 depth cells. Initialization coverage helps; it has not yet established a reliable evolutionary advantage. We also now expose an IntuiCell-inspired decentralized vector-field proxy as a separate mode; it is an auditable guess from public research themes, not the proprietary system. In the first mechanism ablation, removing attractor memory dropped its held-out score to 0.563, while removing recurrent plasticity or prediction-error modulation did not hurt. Five independent fixed- IntuiCell searches now average 0.734 held-out, suggesting its result is not only family-selection luck. Our first local BetaEvolve-inspired arbitration search combined IntuiCell and JEPA, but reached 0.528 held-out versus 0.602 for single IntuiCell; the combination is not promoted. A follow-up over frozen evolved winners reached 0.686 versus 0.633 for equal weighting, but still trailed the best frozen IntuiCell component at 0.711. A three-way mixture adding frozen Liquid CfC improved over equal weighting but fell to 0.540 held-out versus 0.686 for the two-way mixture and 0.703 for IntuiCell alone; evolution suppressed the weak Liquid arm, but diversity alone was not sufficient. Validation-only selection of a stronger JEPA winner raised the next mixture to 0.696, still below JEPA at 0.700 and IntuiCell at 0.716. Across eight independent arbitration seeds, mixtures averaged 0.711 ± 0.009, beating equal weighting every time but IntuiCell in only 3 of 8 runs. That replication also exposed an inert shared gate term that canceled under softmax. We repaired it with genuinely component-specific contextual slopes and repeated all eight seeds; the contextual gate averaged 0.702 ± 0.010, below the constant mixture and IntuiCell, and never beat IntuiCell. We then froze a five-feature gate from wall clearance, resource cues, and internal deficits and moved final evaluation to untouched seeds 400–407. It averaged 0.696 ± 0.014, versus 0.735 for IntuiCell and 0.741 for JEPA, beating neither component in any of eight runs. The defect is fixed, but contextual softmax gating is not promoted. A fresh winner-take-all routing test then avoided action averaging and beat matched blending in all eight runs, averaging 0.743 ± 0.001. But frozen JEPA scored 0.783 on those untouched seeds and won every comparison. Routing is cleaner and more stable, yet still does not create an advantage over its strongest component. We have now reused NExtAI's AlphaEvolve-style quality-diversity architecture for a matched eight-candidate Liquid CfC pilot. Across four fresh-seed cohorts, archive search led mean score at 0.280, versus 0.204 for global lottery and 0.089 for global evolution. It beat lottery in only 2 of 4 cohorts, so this is a promising direction, not a superiority claim. An eight-cohort preregistered replication on another fresh split did not confirm the lead: QD averaged 0.087, evolution 0.046, and lottery 0.113, with both paired confidence intervals crossing zero. The pilot signal is superseded; no search policy has demonstrated superiority. A matched archive-resolution ablation then found a concrete bottleneck: five bins raised mean occupied cells from 1.0–1.75 to 3.0 and fresh score to 0.137, winning all four comparisons against two bins. The higher resolution is promising, but still only a diagnostic. In an independent eight-cohort five-bin replication, QD rose to 0.408 versus 0.331 for evolution, while lottery led at 0.474. Both paired confidence intervals crossed zero, so improved archive diversity still has not produced a reliable search advantage. Doubling the matched budget to 16 candidates increased five-bin archive occupancy to 5.5 cells, but QD averaged 0.301 versus 0.290 for evolution and 0.358 for lottery across four new cohorts. Both paired confidence intervals again crossed zero. More archive depth did not unlock reliable flight or search superiority. Back on the primary foraging track, we preserved the wall-following reflex and restricted CLRM to a bounded residual. The frozen residual improved to 0.062 across 32 untouched worlds, but enabling plasticity fell to −0.142; plastic minus frozen was −0.203 with its entire confidence interval below zero. The useful prior survives, while the current learning update destroys it. A new 32-world timing ablation then moved feedback after each action's consequences, but scored −0.125 versus −0.117 for legacy timing and 0.048 frozen. Correct causal ordering alone did not help; the running-baseline reward or action-level credit signal needed repair. Replacing it with the measured one-step discomfort change produced the first clean primary-track learning win: transition-delta scored 0.135 across 32 new worlds, versus −0.147 for running-baseline learning and 0.063 frozen, with both paired confidence intervals above zero. This is a calibration-world mechanism result, not robust flight; it must replicate against wall-following in production-scale worlds before powered evolution opens. That 32-world production replication now clears the trivial baseline: transition delta scored −0.455 versus −0.882 for wall-following, a reliable +0.427 gain. But frozen residual scored −0.505, and the learner's +0.050 advantage had a confidence interval crossing zero. Nontriviality cleared; reliable within-lifetime learning did not, so the powered gate remains closed. Causal prediction progress later scored 0.369 versus 0.311 for raw surprise, 0.327 for its frozen-motor twin, 0.363 for path-only, and 0.358 for slope credit on 32 fresh worlds. It led every mean, but all four paired score intervals crossed zero and contact trailed slope by one world, so it was not nominated or extended. An unseen steering-sign inversion then created a genuine within-lifetime shift. Observation-only optical calibration reliably improved score by 0.127, late mobility by 0.178, and post-shift contact by 0.094 over its frozen twin, with every confidence interval above zero. Yet it recovered only 25% of oracle late mobility and contacted resources after the shift in 9/64 worlds versus 53/64 oracle, so the absolute-performance gate failed. A fresh follow-up rewound five recent heading updates after the inferred sign switch. Repair improved score from −0.261 to −0.222 and mobility from 0.150 to 0.191, but every optical-only interval crossed zero and oracle mobility remained 0.758. The latent repair was not nominated. Normalizing ray evidence and adding exact visible-cue kinematics then scored −0.259 versus −0.238 raw ray and reliably lost 0.094 post-shift contact fraction. It recovered only 21% of oracle mobility, so sparse cue fusion was rejected. A fixed three-transition change-point latch then scored −0.085 versus −0.198 for raw EMA and raised late mobility from 0.227 to 0.368, with both paired confidence intervals above zero. It reached 50% of oracle mobility, but contact noninferiority crossed zero and the 85% absolute floor remained unmet. Full-pose replay then scored 0.011 versus 0.004 heading-only and raised contact mean, but its score/contact intervals crossed zero and mobility fell from 0.469 to 0.446. It reached 59% of oracle mobility and was rejected; the simpler heading-only learner proceeded to replication. Across 128 new worlds, it beat raw EMA by 0.186 score, 0.220 late mobility, and 0.227 post-shift contact fraction, with every paired lower bound positive. It reached 57% of oracle mobility, validating the mechanism while leaving the 85% absolute gate closed. Replacing the route map at the latch then reduced score from −0.061 to −0.121, mobility from 0.396 to 0.312, and nearly eliminated water contact. Local reset recovered only 42% of oracle mobility and was rejected; the pre-change map must be preserved. Restoring its snapshot and replaying exactly three evidence-generating transitions then scored 0.074 versus −0.056 heading-only, raised mobility from 0.377 to 0.491, and restored both resources. All paired mechanism gates passed, while 67% oracle mobility still missed the 85% floor. On 128 independent worlds, untouched replay again beat heading-only by 0.127 score, 0.140 mobility, and 0.164 post-shift contact fraction, with all lower bounds positive. Recovery again reached 67% of oracle, making the absolute floor the sole blocker. A five-of-seven evidence fallback then found two additional shifts without any premature latches. It raised mobility from 0.563 to 0.573, but score and mobility confidence intervals versus primary replay both touched zero and recovery reached only 71.6% of oracle. Interrupted evidence has diagnostic value, but detector sensitivity is not promoted as the next bottleneck. A fresh route-validation test then excluded retained targets that contradicted visible clearance after the latch. It raised late mobility from 0.522 to 0.812—111% of oracle—and improved score by 0.152, with both confidence intervals above zero. But it averaged 31 invalidations per world, reduced post-shift contact by 0.188 with the entire interval below zero, and shifted coverage strongly toward food while losing water. Route contradiction is a real mobility mechanism, but continuous invalidation fails the joint foraging gate. Restricting validation to only the target retained at the latch then restored contact to 31/64 worlds versus 32/64 for replay, but mobility fell to 0.553 and both score and mobility intervals crossed zero. The matched continuous arm again exceeded oracle mobility while cutting contact worlds to 15/64. One stale target is not the cause; contradiction handling must remain active without repeatedly overriding route coverage. Durable edge quarantine then reached 0.851 mobility—111% of oracle—but contact fell from 34/64 replay worlds to 9/64 and water averaged only 0.172 contacts. It still made 26 invalidations spanning eight logical edges per world. The retained coordinate frame is systemically inconsistent, so editing that map cannot restore coverage. Guiding each replacement opening toward the more deficient visible resource then kept mobility at 0.789—103% of oracle—but contact fell to 8/64 versus 28/64 replay worlds. Water was exactly unchanged from unguided validation, and all cue-versus-unguided score/mobility intervals crossed zero. Sparse local cues arrive too late to rebuild systematic maze coverage. Switching once at the learned latch to the existing ray-only right-hand policy then retained 109% of oracle mobility and reliably beat replay score and mobility. Contact improved to 18/64 versus 11/64 for continuous validation, but remained below replay at 25/64 and water averaged only 0.078 contacts. Map-free wall topology helps, but a fixed hand rule does not traverse both resource regions. Switching once to the mirrored hand rule after the first observed replenishment then raised water from 0.313 to 0.844 contacts while reducing food from 3.000 to 2.500. It fired in 16/64 worlds, but total contact worlds stayed exactly 18/64 and all switching-versus-fixed intervals crossed zero. Post-success routing cannot repair first-resource discovery. Alternating hand rules on ray-defined dead-end entries then reliably beat replay by 0.098 score and 0.258 mobility, reached 94% of oracle mobility, retained both resources, and reached 28/64 contact worlds versus 29/64 replay. But contact CI95 [−0.156, +0.125] crossed the fixed noninferiority gate. The policy switched 36 times per world and reliably lost mobility versus fixed hand, exposing premature rearming during dead-end exits. Requiring confirmed forward corridor clearance reduced switches from 32.4 to 24.3 on fresh worlds and reliably improved score and mobility over predicate rearming. It reached 103% of oracle mobility, but contact stayed 18/64 versus 28/64 for replay, with CI95 [−0.297, −0.016]. The earlier near-contact result did not replicate; permanent map-free switching discards too much route coverage. Restricting the hand rule to temporary dead-end escapes then restored contact exactly to replay's 27/64 worlds and retained both resources. But it scored 0.032 versus 0.075 for replay, mobility was 0.510 versus 0.529, both paired intervals crossed zero, and it reached only 67.5% of oracle mobility. The handoff works; generic maze geometry does not identify when replay actually needs help. Conditioning escape on a visible contradiction of replay's intended target then raised mobility to 0.745—100.4% of oracle—and reliably beat replay mobility by 0.144. But contact fell from 30/64 replay worlds to 21/64, with CI95 [−0.266, −0.016], and water fell to 0.719 contacts. The route context is useful, but the same blocked target reactivated after each handoff, producing 39.7 escapes per world. The next refinement permits one escape per target lifecycle instead of adding a tuned duration or ray threshold. That fixed debounce reduced escapes from 36.3 to 7.4 per world, but mobility fell to 0.470—68.1% of oracle—and contact fell to 17/64 versus replay's 32/64. Score, mobility, and contact were all reliably worse than replay. Repeated escape was carrying the mobility benefit while replay remained committed to the same contradicted target; the next mechanism must replace that target after escape rather than only suppress retries. Deferred retirement then resolved 4.8 still-blocked targets per world and raised water from 0.391 to 0.984 versus debounce, but mobility reached only 59.6% of oracle and contact fell to 13/64 versus replay's 31/64, with CI95 [−0.406, −0.156]. Local route editing remains unable to restore systematic coverage. The next branch bounds repeated high-mobility escape with the first observed replenishment, then freezes back to replay. That bounded arm reliably beat replay score by 0.097 and mobility by 0.184, reached 87.8% of oracle mobility, retained both resources, and improved water over repeated escape. Yet contact remained 23/64 versus replay's 35/64, with CI95 [−0.297, −0.078]. This clears every fixed gate except contact, but the event occurs after contact and therefore cannot repair first-contact coverage. The next branch protects visible resource approaches before contact. Cue-visible replay handoff raised contact from repeated escape's 22/64 to 25/64, but still trailed replay's 31/64 and reached only 71.0% of oracle mobility. Direct CueWallFollower arbitration controlled 125 steps per world, fell to 15/64 contact worlds, and reached 63.2% of oracle. Visibility at six-cell range is too broad; the next refinement requires the cue to be forward-facing without adding a tuned distance. That restriction reduced handoff from 71 to 47 steps per world, raised contact from 23/64 to 27/64, improved water to 1.219, and reached 86.8% of oracle mobility. But replay reached 36/64 and the contact CI95 was [−0.266, −0.016]; score and mobility intervals versus replay also crossed zero. Restricting protection again to the central 90-degree, sensor-aligned cone reached 87.0% of oracle mobility but only 19/64 contact worlds versus replay's 27/64. Its replay-relative score and mobility intervals crossed zero, and repeated escape was reliably better on both measures. Direction-only cue protection is therefore rejected. The next branch compared replay and escape commands on the same observation and selected strictly greater current-ray clearance. It reached 88.0% of oracle mobility, reliably improved replay mobility by 0.142, and reliably beat repeated-escape score by 0.027. But its replay-relative score interval narrowly crossed zero and contact was 28/64 versus replay's 34/64. This is a promising frontier move, not promotion. A visible-resource angular veto then produced reliable replay gains of 0.097 score and 0.218 mobility, reached 86.8% of oracle mobility, and recovered contact to 35/64 versus replay's 36/64. But its contact interval remained [−0.125, +0.094], and it reliably reduced the stronger clearance-only arm. A categorical opposite-side veto reduced intervention to eight steps per world and reached 92.4% of oracle mobility, but contact fell to 26/64 versus replay's 35/64. Cue-side vetoing is rejected. One-transition confirmation of the same contradicted target then reliably improved replay score by 0.120 and mobility by 0.250, raised contact from 34/64 to 37/64, and reached 111.2% of oracle mobility. Its contact interval [−0.063, +0.156] still crossed zero, so the mechanism was frozen for an untouched 128-world replication. The replication preserved reliable gains of 0.102 score and 0.191 mobility, reached 92.7% of oracle mobility, and led contact 60/128 to 58/128. But contact CI95 [−0.047, +0.086] crossed zero and one early-latch world intervened before the shift. The adaptation effect is real; robust flight is not yet established. Requiring the existing three-observation evidence streak removed early intervention, reliably improved replay score by 0.130 and mobility by 0.214, and reached 97.4% of oracle mobility. Contact rose 31/64 to replay's 29/64 but again crossed zero, while mobility reliably trailed the one-transition arm. More waiting is rejected; the next mechanism requires measured non-progress across the causal transition. That causal ray test reliably improved replay score by 0.112 and mobility by 0.217 and reached 98.1% of oracle mobility, but contact fell to 28/64 versus replay's 30/64. Ray progress is not coverage-safe; the next trigger protected first-pass edges using replay's own route memory. It retained contact at 33/64 versus replay's 31/64, but allowed only 6.5 escapes per world, did not reliably improve score or mobility, and reached 68.5% of oracle mobility. Permanent protection is too strict; the next hybrid gives first-pass edges the existing three-observation grace, then allows recovery. That hybrid reliably improved replay score by 0.124 and mobility by 0.239 and reached 103.9% of oracle mobility, but contact fell to 28/64 versus replay's 32/64. Route-memory timing is exhausted as a coverage trigger. Symmetric midpoint authority then reached only 35.4% of oracle mobility and reliably lost replay score, mobility, and contact; partial turns prolonged escape lifecycles to 120 selected steps per world. Winner-take-all authority is retained. Limiting that authority to a single selected action then passed the fixed nomination gate on fresh seeds: score improved by 0.077 with CI95 [+0.023, +0.131], mobility by 0.130 with CI95 [+0.032, +0.229], and contact rose from 34/64 to 42/64 with CI95 [0.000, +0.250]. The pulse remained pre-shift inert, retained food and water, and reached 86.3% of oracle mobility. It was then frozen for 128 untouched worlds. It again reliably improved replay score by 0.048 with CI95 [+0.006, +0.089] and mobility by 0.072 with CI95 [+0.002, +0.140], stayed pre-shift inert, and reached 86.9% of oracle mobility. But contact was only 71/128 versus replay's 70/128, CI95 [−0.063, +0.078]. The adaptation mechanism replicates; the joint robust-flight gate still fails. A spent-split diagnostic then allowed pulses only while the currently deficient resource cue was visible. Contact rose to 73/128, but mobility fell below replay and to 75.6% of oracle. The visibility asymmetry was not viable, so no fresh worlds were consumed. A fresh composition then retained every cue-blind pulse and vetoed only visible pulses that increased angular error to the deficient resource. It reliably improved replay score by 0.065 and mobility by 0.107 and raised contact from 27/64 to 31/64. But contact CI95 [−0.031, +0.156] crossed zero and mobility reached only 83.4% of oracle; the alignment veto is rejected. A coverage-memory composition then preserved pulses until the first observed replenishment and froze to replay afterward. It reliably improved replay score by 0.113, mobility by 0.195, and contact by 0.109 with CI95 [+0.031, +0.203], reaching 36/64 contact worlds versus replay's 29/64. A later audit found that bounded and pulse-earned modes were omitted from immediate cancellation and accidentally inherited resource alignment. Those statistics remain defect evidence, but their one-action mechanistic interpretation is invalid. After repair, fresh seeds 9200–9263 tested corrected immediate bounding and exactly one selected post-replenishment grace pulse. Grace reliably improved replay score by 0.087 and mobility by 0.137, reached 85.7% of oracle, retained both resources, and stayed pre-shift inert. Contact was 34/64 versus 31/64, CI95 [−0.047, +0.141]. The extra pulse did not improve contact over immediate bounding and is not promoted. A fixed two-action pulse then repeated every selected escape command once. It reached 94.8% of oracle mobility and reliably improved replay score by 0.136 and mobility by 0.259, while also reliably beating one-action recovery. But contact fell to 21/64 versus replay's 23/64 and one-action's 24/64. The second forced action is rejected. A causal refinement then repeated only when the same target remained blocked after the first action. It suppressed 4.5 repetitions per world, reached 102.0% of oracle, and reliably improved replay score by 0.110 and mobility by 0.246. Contact remained 32/64 versus replay's 34/64 and was reliably below one-action pulse's 39/64. The surviving-blockage condition is not coverage-safe. Suppressing every further pulse until that blocked route visibly cleared then cut pulse use from 23.1 to 8.0 actions per world, but score fell to 0.093 versus replay's 0.121 and mobility to 0.558 versus 0.600. Contact tied replay at 36/64, both paired score and mobility intervals crossed zero, and recovery reached only 74.3% of oracle mobility. Repeated pulses were doing useful recovery work; clearance rearming is rejected. A causal middle cadence then suppressed exactly every other independently confirmed pulse. It used 16.8 actions per world, between unrestricted pulse's 23.3 and hard rearming's 8.3, but scored 0.097 versus replay's 0.093 and reached only 73.3% of oracle mobility. Contact rose 33/64 to 36/64, yet score, mobility, and contact intervals all crossed their fixed gates. Cadence throttling is rejected as the coverage-safe lever. We then let each pulse teach the next escape hand: repeat the hand when the same target's clearance improved, otherwise switch. The learner made 27.2 updates per world, but increased pulse use from 23.9 to 28.2 and trailed blind alternation on score, mobility, and contact. It scored 0.042 versus replay's 0.000, contacted 30/64 versus 28/64, and reached 76.7% of oracle mobility; every replay-relative interval crossed zero. One-step clearance is not a coverage-aligned teaching signal. Replacing it with completed dead-reckoned route edges then reliably improved replay score by 0.095 and mobility by 0.141, reached 96.1% of oracle mobility, and raised contact from 29/64 to 33/64. Contact uncertainty still crossed zero, and blind alternation retained stronger score and mobility means. Route completion is directionally credible, not promoted. Restricting each successful hand to its dead-reckoned route edge then reliably improved replay score by 0.089 and mobility by 0.162, reached 85.3% of oracle mobility, and raised contact from 27/64 to 29/64. The controller learned 3.5 edge-hand associations but recalled only 0.7 per world; contact uncertainty again crossed zero. Contextual credit is valid but too sparse, so it is not promoted. Generalizing memory across the same eight-bin target-bearing sensor sector raised recall to 18.2 uses per world and contact to 40/64 versus replay's 36/64. Yet score, mobility, and contact intervals all crossed zero, mobility reached only 78.5% of oracle, and global completion learning led at 43/64. Coarse sector memory overgeneralizes and is rejected. We then froze the unchanged global completion learner for 128 untouched worlds. It reliably improved replay score by 0.089 and mobility by 0.184, reached 89.9% of oracle mobility, and contacted 66/128 worlds versus replay's 59/128. The contact interval still crossed zero, and every comparison with blind alternating pulses crossed zero. Recovery replicates; learned contact superiority does not. Qualifying completion credit by first traversal of a dead-reckoned edge then reliably beat replay score and mobility and reached 87.9% of oracle mobility. But it contacted 36/64 worlds, exactly the same worlds as global completion learning, while replay contacted 34/64 and the interval crossed zero. Edge novelty is not the missing discovery signal. Moving novelty into the observed eight-ray open/blocked pattern then reached 92.8% of oracle mobility and reliably beat replay score and mobility, but contacted only 34/64 worlds versus replay's 32/64. The interval crossed zero and every pulse comparator contacted more worlds. Binary visual novelty is also too coarse and sparse. Replacing it with the stable rank ordering of all eight rays then triggered 19.6 novel repeats per world, but contacted only 28/64 worlds versus replay's 27/64 and global completion's 31/64. Score and contact intervals crossed zero, and mobility reached only 84.86% of oracle. Fine-grained ray order over-differentiates geometry. A compact nearest/widest-ray pair then reached 91.4% of oracle mobility and contacted 36/64 worlds versus replay's 33/64, but score and contact intervals crossed zero and global completion led at 39/64. Lifetime memory left only 0.53 novel repeats per world, and one early-latch world violated zero pre-shift intervention. The compact code saturates before learning. A fresh cohort then cleared extrema memory once at the observation-derived latch. Novel repeats rose to 3.66 per world, score and mobility reliably beat replay, and mobility reached 90.4% of oracle. Contact rose from 31/64 to 35/64 but remained uncertain, while unchanged lifetime extrema reached 37/64. Regime scoping repairs saturation, not discovery. A fresh cohort then used the same nearest/widest-ray pair as a hand-memory key and wrote only after a dead-reckoned route completed. It learned 2.59 associations and recalled 1.64 hands per world, but contacted 28/64 worlds versus replay's 32/64. All replay-relative gates failed, and exact-edge memory reliably led score and mobility. Context granularity is not the missing discovery signal. We then replaced last-success overwriting with signed completion/failure votes in each target sector. The voter made 23.19 updates and 19.77 majority recalls per world, yet tied replay at 39/64 contact worlds and failed every inferential gate. Last-success memory reached 41/64, and one early-latch world caused the shared pre-shift diagnostic failure. Dense evidence is not enough when the credit target does not align with discovery. We next credited escape hands only when the same deficient visible resource moved closer. It reliably beat replay and resource-aligned pulse on score and mobility, stayed pre-latch inert, and reached 90.2% of oracle mobility. Contact was 31/64 versus replay's 32/64 and exactly matched blind pulse; only 6.69 of 20.33 selections had valid cue feedback. Task-aligned credit improves recovery quality, not discovery. We then retained signed cue-progress votes separately for food and water. Seven learned recalls per world reliably improved replay score and mobility and beat resource-aligned pulse score and contact, but contacted the same 33/64 worlds as one-shot learning. Contact uncertainty crossed zero, oracle mobility coverage fell to 72.4%, and one early-latch world failed the shared pre-shift diagnostic. Persistent task memory still does not create discovery. We then moved cue-progress learning upstream to visible three-way junctions. Fewer than one behavior-changing choice and valid update occurred per world; the learner made 0.94 cue-aligned choices and only 0.05 coverage choices. It reliably lost replay score and mobility, contacted 30/64 worlds versus replay's 33/64, and reproduced fixed cue alignment. Euclidean cue pursuit at rare branches destroys useful exploration coverage. We repaired the ordering by keeping edge and cell visits ahead of cue angle. Contact rose to 41/64 versus replay's 39/64 without pre-latch intervention, but all inferential gates crossed zero and mobility reached only 61.3% of oracle. Just 0.59 behavior-changing ties occurred per world; learning matched fixed cue tie-breaking on contact and trailed one-shot cue progress. Coverage preservation removes harm, not the missing learned advantage. We then froze the unchanged one-shot cue-progress learner for 256 untouched worlds. It reliably improved replay score by 0.057 and mobility by 0.100, reaching 86.2% of oracle mobility. Contact rose only from 137/256 to 141/256 with an interval crossing zero; every comparison with blind pulse remained uncertain, and one early-latch world failed the shared pre-latch diagnostic. Higher power confirms recovery quality, not reliable resource discovery. We next credited the immediate effect a turn can control: reduced absolute bearing to the deficient visible resource. The learner made 7.06 valid updates per world, reliably improved replay score and mobility, stayed pre-latch inert, and reached 89.4% of oracle mobility. It led contact at 33/64 versus replay's 28, blind pulse's 30, and cue-distance learning's 31, but its contact interval and every pulse-comparator interval crossed zero. Bearing is cleaner credit, not yet robust discovery. We then accumulated separate food and water bearing votes. The voter made 6.55 valid updates and 10.23 majority recalls per world and reliably beat blind pulse contact, 36/64 versus 32/64. But replay reached 37/64, the replay-relative score and contact gates failed, and voting reliably lost one-shot bearing score. Resource-wide votes overgeneralize local turning evidence. We therefore froze the local bearing learner for 256 untouched worlds. It reliably improved replay score by 0.056, mobility by 0.109, and contact from 124/256 to 139/256; all three paired confidence intervals cleared zero, both resources remained represented, and mobility reached 86.7% of oracle. Yet one false-positive observation-only latch fired before the actual remap in every pulse arm. This is the first complete efficacy pass, but robust flight remains closed on latch safety. A matched safety test then raised the optical confirmation count from three to five without using a clock or remap time. It stayed pre-remap inert, but latch coverage collapsed from 60/64 to 43/64, mean latch time slipped from step 324.7 to 359.9, and contact fell to 18/64 versus replay's 35/64. Count-only confirmation suppresses false positives by missing real changes; the latch needs evidence quality, not just more votes. We then retained three confirmations but required their mean normalized ray contrast to reach 0.1, frozen from consumed diagnostics. On 64 fresh worlds it stayed pre-remap inert, preserved all 60 legacy latches at the same mean step, exactly matched legacy bearing's 41 contact worlds, and reached 93.4% of oracle mobility. Its replay-relative intervals crossed zero at this sample size, and legacy produced no fresh false latch, so the unchanged quality gate was nominated for high-power replication. On 256 untouched worlds it produced zero early latches or interventions while matching legacy latch coverage and timing. Score and mobility reliably beat replay, both resources remained represented, and mobility reached 82.9% of oracle. Contact, however, was 126/256 versus replay's 123 and blind pulse's 129; its confidence interval crossed zero. Latch safety is repaired, but discovery reliability did not replicate. We then combined the safe latch with the earlier visible-resource alignment veto. It suppressed 8.97 pulses per world and reduced contact to 37/64 versus quality-blind pulse's 40/64; every replay-relative efficacy interval crossed zero. Alignment removes useful exploration. The simpler safe pulse again showed the stronger contact direction, 40/64 versus replay's 32/64, and was nominated unchanged for high-power replication. That frozen replication reliably improved replay score by 0.073 and mobility by 0.130, reached 91.4% of oracle mobility, retained both resources, and had zero early latches or interventions. Contact rose from 129/256 to 141/256, but CI95 began exactly at zero; the strict contact gate therefore remained closed. A preregistered 512-world confirmation then retained reliable score and mobility gains, both resources, 84.4% of oracle mobility, and zero candidate early latches or interventions. But contact was 257/512 versus replay's 258/512, CI95 [−0.041, +0.037]. Blind-pulse recovery replicates; discovery does not, so this unchanged line is closed. We next required two consecutive replay transitions to reduce both distance and bearing to the same deficient visible resource before granting direct cue authority. This cut overrides from 120.94 to 6.59 per world and reliably recovered 22 contact worlds over broad direct pursuit, but still contacted 32/64 versus replay's 33/64 and reliably harmed score and mobility. Causal gating fixes over-triggering, not the incompatible direct action; direct cue authority is closed. We then moved prediction error into edge-first route selection without changing corridor motor execution. The model learned from 228.06 transitions per world, but its global sector means changed only one of 29.25 target choices. Contact fell to 27/64 versus replay's 29/64, all replay-relative gates failed, and mobility reached 61.3% of oracle. It matched sector-count novelty and reliably lost blind-pulse score and mobility. Dense surprise needs local candidate structure; global sector averaging is rejected. A fresh production sweep then tested frozen, conservative, default, and aggressive transition-delta rates. All four averaged about −0.50 and every plastic-versus-frozen interval crossed zero. Simple learning-rate tuning is not the fix. A fresh reward-horizon sweep then compared one-step delta with warm-started causal smoothing. Both smoothed arms averaged −0.490, below one-step at −0.412 and frozen at −0.426, with every interval crossing zero. Smoothing added no useful credit, so task-aligned reward content is now the mechanistic target. The next content ablation found task-weighted credit at −0.404 versus −0.514 frozen, winning 20 of 32 pairs, but its interval narrowly crossed zero. Homeostasis-only exactly matched frozen because no resource contacts occurred, locating the current bottleneck in navigation. The task-weighted signal is nominated, not promoted. Its independent 64-world replication preserved the effect: −0.411 versus −0.503 frozen and −0.847 wall-following. It reliably cleared wall-following and won 42 of 64 frozen pairs, but task minus frozen was +0.092 with CI95 [−0.019, +0.199]. The cohort was not extended; the strict learning gate remains closed. A full-cohort audit exposed high paired variance, so a new diagnostic compared residual gains 0.15, 0.25, and 0.35. The two lower gains had positive paired intervals; gain 0.15 reduced delta SD from 0.456 to 0.396 and large harms from five to two. Because gains were compared, 0.15 is nominated for independent replication, not counted as a gate pass. That fixed replication has now passed: task-weighted gain 0.15 scored −0.559 versus −0.922 frozen and −0.952 wall-following. Its gains were +0.362 and +0.393, with both CI95 lower bounds above zero and 56 of 64 paired wins each. Learning and nontriviality are cleared for this controller; powered search and robust held-out behavior remain ahead.