← for the love of the problem

teaching machines to play

learned ML by building game-playing AIs instead of doing the usual tutorials.

i wanted to learn ML, but the usual starter projects weren't doing it for me. games were eye-catching enough that i actually wanted to know what was going on inside them, so i built two: a flappy bird agent with deep Q-learning, and a top-down racing car that learned a circuit with tabular Q-learning. then, while putting this portfolio together, i dusted them both off, re-ran them, and found out neither had really learned anything. so i fixed them. the fixed versions are playable, so you can race them.

Loading…

Pipes passed while trainingThe 2023 recipe's rolling mean over training games never rises above 0.55 pipes, ending at 0.15 pipes. The 2026 rework's greedy score at each saved checkpoint, on validation layouts: 0 pipes after 0 games, 0.65 pipes after 97 games, 26.65 pipes after 240 games, 0.55 pipes after 451 games, 109.2 pipes after 577 games. A hand-written rule scores 69.52 pipes on held-out layouts. Capped at 200 pipes.0100200hand-written rule (held-out)2023 recipe2026 reworkgames played →
2026 rework: greedy score at each saved checkpoint, on validation layouts · 2023 recipe: rolling mean of training games · hand-written rule: held-out layouts · capped at 200 pipes · the last dot is the checkpoint chosen on these same validation layouts, so it flatters; on held-out layouts it averages 85.94 pipes
  • 2023 recipe
  • 2026 rework
  • hand-written rule

the setup

a stripped-down flappy bird clone, and an agent that sees seven numbers every frame: how far the top and bottom of the next gap are from the bird, the same for the gap after it, how far away the next pipe is, its own vertical speed, and its height. it gets two moves, flap or don't. a small neural network scores both and the agent picks the better one.

that's deep Q-learning. nobody tells it anything about pipes. it just gets rewarded, gets punished, and crashes. a lot.

the part where it never learned

back in 2023 i trained it, watched it clear the odd lucky pipe, and very confidently wrote in the README that it had started to learn.

cut to this year. i'm adding it to this portfolio, i re-run that exact recipe on the same game, and it averages 0.15 pipes. that's not learning. that's falling with extra steps.

reading the code back, the bugs are almost a checklist. (i didn't test them one at a time, so i can't tell you which one did the real damage. any of them could have.)

raw pixel values in the hundreds went straight into the network with a learning rate of 0.05, which i think saturated it. that's the stuck-on-one-action collapse 2023 me blamed on exploding gradients.

dying cost −1000 while a pipe paid +10, and with a discount of 0.9, a pipe forty frames away was worth about one and a half percent of its +10. next to the −1000, pipes barely registered.

a safety rule took over near the floor and ceiling, but memory recorded the move the agent had chosen, not the one that actually happened. so it was learning from moves it never made.

exploration switched off after 80 games. and the training target wasn't detached, so every update also nudged the thing it was aiming at.

dusting it off

same game, same family of algorithm, just a recipe that works.

inputs are scaled to roughly −1 to 1 and measured from the bird. rewards are +0.1 per frame alive, +1 per pipe, and −1 for a crash, with a discount of 0.99 so the next pipe actually stays in view. double DQN with a target network, Huber loss, and gradient clipping (which, fun fact, the 2023 README proposed as the fix and then never tried). mini-batches from a replay buffer. and no safety override. it keeps itself alive now.

the change that made the biggest difference was the one i least expected: letting it see the gap after next. consecutive gaps can be far apart with very little time between them, so you have to start moving before you've even cleared the current pipe. without that, every version i tried plateaued.

also, DQN on this task swings between brilliant and hopeless from one evaluation to the next, and most of the later checkpoints in the run fell apart again. so i keep the best checkpoint instead of the last one, and the numbers below are on layouts that weren't used to pick it.

flappy bird, scoreboard

on 100 pipe layouts it never saw in training or selection, capped at 200 pipes, the fixed agent averages 85.9. the 2023 version averages 0.15. and a hand-written rule (flap when you drop near the bottom of the gap) averages 69.5.

so it beats a rule that fits in one sentence. i'll take it.

the car, round one

the racing car came a year later. a top-down pygame game on a trace of the Bahrain circuit, learning with a table of Q-values. when i dusted this one off and re-ran its state, reward, and settings on the rebuilt game, from 32 starting points round the lap it never trained from, it covered 0.40% of a lap on average.

the table had room for millions of states. it visited 13 of them. i think that's because each distance sensor was split into five bins 240 pixels wide, so being a car's length from a wall looked exactly the same as being in the middle of the road.

the reward paid for distance driven, not distance round the track, so driving in circles counted as progress. and the game had bugs of its own: hitting a wall only halved your speed, and the car moved twice per frame whenever you weren't steering.

the car, round two

same fixes as flappy bird, plus one i had to build from scratch: a way to measure progress. i flood-fill the track image once, starting from the finish line, so every pixel of road knows how far round the lap it is. the reward is how much of that the car gained this frame. circles don't pay anymore.

the table becomes a small network. the sensors become seven distances that turn with the car, plus its speed. a crash ends the run. it trains from starting points all round the lap and gets judged on others it never saw.

and same instability as before. the score swung hard between checkpoints and several later ones collapsed, so again i keep the best one, not the last.

the car, scoreboard

from 32 held-out starting points, the fixed agent covers 90.6% of a lap on average and finishes 29 of them. from the start line it does a full lap in 18.05 seconds. the 2024 version covers 0.40%.

a short hand-written rule does better, though: steer toward the more open diagonal, slow down when the road ahead closes. it covers 93.8% and finishes 30. turns out sensors make this track easy if you already know how to drive. the agent had to work that out from reward alone, so i'm calling it a moral victory.