Jev vs Haiku vs an untrained fruit fly connectome vs a baseline bot. Each competitor has a ‘skin’, which is the method of pre-processing inputs before actually calling the model.
Note: Haiku ran 15 tracks while Jev ran 100 tracks and Jev’s price is estimated through the pricing model, so comparison is a lead and not a verdict.
I was curious to know which brain could succeed the most in a simple game and how that translates to money and time. I also wanted to know whether Jev could hold its own against a small general chat model, and whether the fruit fly brain connectome (without training) could compete at all. The setup:
- A ring of 12 lanes. 150 rows.
- For each row, a runner can stay, move left or right, or jump. Landing on a gap ends the run.
- A player sees 6 rows and 3 lanes either side: 42 tiles, as JSON.
| player | skins | what it is | cost |
|---|---|---|---|
| Jev | plain, guided, step-1, step-2, map | TypeSafe's System One model, jev-latest | estimated per request |
| Claude Haiku | plain, guided, step-1, step-2 | claude-haiku-4-5-20251001 | measured per request |
| Fly | looming, sideways | Shiu et al.'s whole-brain model on the FlyWire v783 connectome, simulated in Brian2 | free |
| Bots | Solver, Random, Always jump | simple baselines | free |
| setting | Jev | Claude Haiku |
|---|---|---|
| model | jev-latest (TypeSafe SDK) | claude-haiku-4-5-20251001 (Anthropic SDK) |
| requests per row | 1, holding every question of the set | 1, one reply holding every answer |
| answers | typed: a probability per yes/no question, or a choice. answered in parallel. answers can’t see each other. | a JSON object. probabilities are stated numbers. answers are written in order and can see each other. |
| set | what is asked | how the answer becomes a move |
|---|---|---|
| plain | "Which move?" once | the move named |
| guided | one choice; each option names the tile it lands on | the move chosen |
| step-1 | 4 yes/no: would this move land on a gap? | lowest probability, rounded to 2 places; ties stay, left, right, jump |
| step-2 | step-1's 4, plus 4: would it leave the runner trapped? | lowest P(gap) + (1 − P(gap)) · P(trapped) |
| map | 42 yes/no, one per visible tile | tiles above 0.5 are gaps; the solver plans on that picture |
Neurons talk in spikes, and firing rate (Hz) is how quickly that neuron spikes. The brain has no eyes or legs, so neurons need to be grouped into “adapters” for each function. The “input” represents the eyes and converts tile gaps into neuron spikes. The closer the gap, the faster the firing rate. The “output mapping” correlates grouped output neuron firing into a single move. Here’s how the fly is set up:
| group | neurons | what it does |
|---|---|---|
| eye neurons | LPLC2, LC4, LPLC4, LC22 | these groups are cells in the fly’s visual system that react to something rushing towards it aka a “looming” threat. in the game, this signal means there is a gap ahead. |
| steering neurons | DNa01, DNa02, DNb01, DNg13 | these groups are cells that carry the “turn” commands from the brain to the body. when the right-hand signals fire more than the left-hand ones, that indicates the player to turn right. vice versa for left. |
| Giant Fiber | DNp01 | this signal represents the fly’s escape neurons. when it fires hard enough, the fly jumps. |
| part | Fly · looming | Fly · sideways |
|---|---|---|
| input | ||
| channels | 2: left eye, right eye | 3: left, centre, right |
| eye cells | LPLC2 + LC4 in each eye | centre: LPLC2 + LC4 in both eyessides: LPLC4 + LC22 in that eye |
| lanes each channel sees |
LLLL RRRR
−3−2−10+1+2+3
a gap straight ahead drives both eyes equally |
LLLCRRR
−3−2−10+1+2+3
a gap ahead and a gap to the side reach different cells |
| signal from gaps | a gap twice as far counts ⅛ as much | a gap twice as far counts ¼ as much |
| where | is how many rows ahead the gap is (1 to 6); each channel adds up every gap in the lanes it sees | |
| cap and rounding | at most 250 Hz, rounded to the nearest 25 Hz | at most 500 Hz, rounded to the nearest 100 Hz, so a centre gap and a side gap can add up |
| e.g. gap 1 row ahead, own lane | → 250 Hz to both eyes | → 300 Hz to the centre |
| e.g. gap 2 rows ahead, 1 lane left | → 25 Hz to the left eye | → 100 Hz to the left channel |
| output mapping | ||
| turn signal | ||
| jump signal | Giant Fiber's average across both sides, for both flies | |
| move order | ||
| 1st check | Hz → jump | Hz → right if is positive, left if negative |
| 2nd check | → right; → left | Hz → jump |
| otherwise | stay | stay |
Note: Steering neurons fire away from a threat and turn the fly towards that direction. A gap on the left activates the left eye, so the right-hand steering neurons fire, and the runner moves right.
| setting | Fly · looming | Fly · sideways | what it means |
|---|---|---|---|
| gain | 50, 100, 150, 250 | 100, 250, 500 | the strength of the signal that a gap is approaching |
| falloff | 1, 2, 3, 4 | 2, 3, 4 | the speed of decay of the gap signal in relation to distance |
| turn threshold | 0, 10, 20, 30, 40, 60 Hz | 0, 10, 20, 40 Hz | what must pass before the runner moves left or right |
| jump threshold | 75 to 250 Hz in steps of 25; 200 | 100 to 300 Hz in steps of 25; 175 | what must pass before the runner jumps |
| settings tried |
Fly1 squeezes everything it sees into one number per eye. Both eyes are equally triggered by a gap straight ahead, so it never knows which way to dodge. Its fallback is to jump with no information on where that jump lands. Fly2 was given a ‘sideways’ channel. Gaps in the side lanes now go to different sets of eye cells. Gaps ahead and gaps to the side are no longer processed the same.
Jev-step-2 performed the best. While Jev-step-2 beat Jev-step-1, the trend of adding questions didn’t continue to improve success. My best guess: Jev-map processes ~1,823 tile reads per track. Notably, Jev-map never missed gaps if they were in his current lane, and there weren’t any gaps in other lanes. When there were gaps in other lanes, false positives jumped up from 11% to 62%. Jev-step-1 and Jev-step-2 didn’t suffer from this problem. This leads me to believe that Jev-map is failing because of 2 possible reasons: (1) framing of the question and/or (2) 42 tiles processed at once.
Why Haiku-plain wins: “Which move?” is actually two questions in one: (1) what tile am I landing on and (2) is that tile a gap. Jev is built for typed questions. The ‘Choice’ result carries little information. Jev-plain chose to stay in place on 374/415 rows. On moves that weren’t 100% safe, Jev-plain made a fatal decision 25% of the time compared to 8% from Haiku-plain. The evidence shows a general model can combine that view into one answer. When Jev is given narrow questions, like in the other skins (barring Jev-map), it performs much better.
The question mattered more than the model. The sets that worked best compartmentalized the question. The model judges one thing at a time aka “would this move land on a gap?” Jev-plain forces all judgements into a single ‘Choice’ result. Haiku can fold them into one answer.
Jev-step-2 vs Haiku-step-2: 43x cheaper and 10x faster. Jev-step-1 and Jev-step-2 are the frontiers; no brain comes close to beating them on rows, price, and time when considered together.
The fly performed better than expected for no cost at all. It struggled choosing where to land. The Giant Fiber is an escape reflex, so the wiring had no idea if its landing spot was safe. It only knew that it was running away.