package kof97rl
RL framework for KOF'97 (Neo Geo arcade) with a self-hosted
MAME Lua bridge — no accounts, no Docker. It ships a fully-worked PPO
agent that trains Kyo Kusanagi to beat all 10 sampled arcade opponents including
the boss (via a generalist + a Kensou specialist dispatched by an opp_character
router), plus everything around it: the emulator bridge, observation/action contract,
reward shaping, rollout buffer, training loop with LR/entropy annealing + weight EMA,
and evaluation/recording tools.
PPO agent (kof97rl/algos/ppo/)
|
Agent interface (act / update / save / load)
|
OnPolicyTrainer -> RolloutBuffer -> TensorBoard + CSV + checkpoints
|
VecEnv of KOF environments (canonical obs/action contract)
|
mock backend (tests, debugging) | MAME backend (bridge.lua over TCP)
Character scope: only Kyo Kusanagi is configured in this repo — one shipped moveset (
kof97rl/envs/movesets/kof97_kyo.yaml), theconfigs/ppo_kyo_*experiments, and savestates that seat Kyo in the P1 slot. Training another character means adding a new moveset YAML + savestates; the framework itself is character-agnostic.
Kyo clears the full 10-opponent gauntlet on a single health bar — the efficient generalist (≈54% less damage taken, ~35% faster than the first 10/10 net):
~11s highlight loop (silent): pressuring Blue Mary → flooring Billy Kane → a rush on Sie Kensou → a clean K.O. The full 304-second run with audio is demos/kyo_10opponents_efficient.mp4 (see below).
Stills from the run — a closer look across the roster (deterministic play, one health bar):
A clean K.O.; a rush combo on Sie Kensou (the specialist's matchup); pressuring Blue Mary on the temple stage; Billy Kane floored. Per-opponent fights and full roster reels are generated locally into the (gitignored) demos/ folder — see Watch & record.
▶️ Embed the full 304s run with audio as an inline video (optional)
GitHub plays inline <video> only from an uploaded-asset URL — never from a file
committed to the repo, and never from a private repo's raw path (which is why the GIF
above is used instead). To add the full clip:
- Open https://github.com/xeno0/kof97_rl/issues/new — do not submit the issue.
- Drag
demos/kyo_10opponents_efficient.mp4into the comment box; wait for the upload. - Copy the
https://github.com/user-attachments/assets/<uuid>URL it inserts. - Add it here as
<video src="THAT_URL" controls muted width="100%"></video>.
- Demo — watch Kyo clear all 10 opponents
- Setup — env, MAME, ROMs, savestates
- 30-second quickstart — see the shipped agent fight
- Train — configs, flags, auto-resume, Net2Net warm-start, monitoring
- Watch & record — live window, video/audio capture, all flags
- Play the trained agent (the 10/10 deploy) — router, per-opponent win rates
- Observations, actions, reward
- Add your own algorithm
- Tests · Layout
- 📘 Techniques deep-dive — the full how-and-why (env, PPO, Net2Net, router)
conda activate kof97 # Python 3.11 (torch, gymnasium 0.29)
pip install -e ".[dev]"
sudo apt install mame # emulator (0.242 here; anything 0.227+ works)
mkdir -p ~/roms # supply your own legally-owned ROMs:
# ~/roms/kof97.zip - the game
# ~/roms/neogeo.zip - Neo Geo BIOS
mame kof97 -rompath ~/roms -bench 60 # sanity: expect >= 500% speedROM/BIOS files are your responsibility and are gitignored — never commit them.
RAM addresses and boot/select timings live in
kof97rl/envs/mame/specs/kof97.yaml. The training pool loads fight-ready
savestates (each one pins a different arcade opponent at round start) from
~/.kof97rl/kof97/savestates/kof97/. The shipped pool is clean_01 … clean_10
(plus a few hard_o* / fight_ready states used during development). To (re)create one:
python scripts/make_savestate.py --game kof97 --state-name clean_01 # boots, coins in, selects, saves at round start
python scripts/make_savestate.py --game kof97 --state-name clean_01 --windowed # watch it happenmake_savestate.py flags: --game (default kof97), --state-name (default
fight_ready), --rom-path (default ~/roms), --windowed (show the MAME window).
The repo ships trained checkpoints — you can watch them fight without training anything.
conda activate kof97
# 1) Watch the shipped generalist fight a random pool opponent in a live MAME window:
python scripts/watch.py --run runs/efficiency_20260709_072914 \
--checkpoint runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
--episodes 3
# 2) Prove the full deploy beats ALL 10 opponents (generalist + Kensou specialist via router):
python scripts/router_eval.py \
--run runs/net2net2_20260708_061012 \
--deploy runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
--specialist runs/kensou_spec_20260708_182902/checkpoints/kensou_specialist.pt \
--episodes 2 # -> OVERALL: 10.0/10 at 100%Pre-rendered fights (per-opponent + full roster reels, with audio) are written to the
local demos/ folder (gitignored — large media, not on GitHub); regenerate them with
watch.py --record.
Everything is driven by a YAML config. The winning recipe is a single file,
configs/ppo_kyo_net2net.yaml: obs = pixels+features,
single full-health rounds, a 10-opponent arcade pool, and a Net2Net-expanded
network refined at low LR with LR/entropy annealing + weight EMA.
# Full training run (needs MAME + ROMs). Writes to runs/net2net2_<timestamp>/
python scripts/train.py --config configs/ppo_kyo_net2net.yaml
# Smoke test with NO emulator — the mock backend, for verifying the training path:
python scripts/train.py --config configs/ppo_kyo_net2net.yaml --mock
# Robust multi-hour launcher: auto-resumes from the newest checkpoint on any crash.
# Usage: train_auto.sh <config> <run_name_glob> <hours> [seed_ckpt]
bash scripts/train_auto.sh configs/ppo_kyo_net2net.yaml net2net 12| Flag | Meaning |
|---|---|
--config PATH |
Experiment YAML (layered on top of configs/base.yaml). |
--mock |
Force the mock backend — trains with no MAME/ROMs (fast smoke test). |
--init-from PATH |
Warm-start weights from a checkpoint (fine-tune / continue). |
--reset-progress |
With --init-from, load the weights but restart the anneal schedule at step 0 (for seeding a new run from an external checkpoint). |
--total-timesteps N |
Override the config's training horizon (also drives the LR/entropy anneal horizon). |
--num-envs N |
Override parallel emulator count. |
--seed N, --device cuda|cpu |
Override seed / device. |
--algo NAME |
Override algo.name (e.g. your own registered algo). |
--list-algos |
Print the registered algorithms and exit. |
Key knobs in configs/ppo_kyo_net2net.yaml:
env.obs_mode: both— pixels and the 18-dim feature vector (what the shipped model uses).env.action_mode: discrete—Discrete(17).env.episode_mode: round— one full-health fight per episode (usematchto climb the whole arcade ladder in one episode).env.savestate_pool: [clean_01 … clean_10]— the 10 sampled opponents; eachreset()draws one.env.num_envs: 24— parallel emulators. Each rank gets its own MAME process onbase_port + rank.env.reward:— damage-dealt/taken weights, win/special/super bonuses, whiff penalty (see reward).algo.params:— PPO hyperparameters + the Net2Net widths (feat_blocks,trunk_blocks,hidden_dim),learning_rate,ent_coef,anneal_lr,ema_decay.train.total_timesteps/train.checkpoint_every— horizon and checkpoint cadence.
Each run creates runs/<run_name>_<timestamp>/ containing:
config.yaml— the exact resolved config (every eval/watch tool reads this back).checkpoints/latest.pt, periodicstep_*.pt, and the EMA shadow inside each.pt.- TensorBoard event files + a CSV of scalar metrics.
tensorboard --logdir runs # loss, entropy, per-opponent reward, LR/entropy schedulesPPO on a mixed opponent pool suffers catastrophic interference — it trades
which opponents it can beat between checkpoints. The fix that produced deploy.pt,
one net holding 9 of the 10 arcade opponents (the 10th, Kensou, is handled by
the specialist + router — see Play the trained agent):
- Train a Phase-1 policy to stability with LR/entropy annealing + weight EMA.
- Function-preserving expansion (
scripts/net2net_deeper.py): grow the EMA policy into a deeper net by inserting residual blocks whose second layer is zero-initialized — the bigger net reproduces the source exactly (verifiedmax|Δlogits| < 1e-4), so no capability is lost at the seam. - Refine the expanded net at low LR (1e-4) + low entropy (0.005) so the extra capacity learns the last holdouts without eroding the warm-started foundation.
python scripts/net2net_deeper.py --src-run runs/<phase1_run> \
--feat-blocks 10 --trunk-blocks 6 --out runs/net2net_seed/deep_from_phase1.pt
bash scripts/train_auto.sh configs/ppo_kyo_net2net.yaml net2net 12 \
runs/net2net_seed/deep_from_phase1.pt # 4th arg = seed checkpointnet2net_deeper.py flags: --src-run (source run dir), --src-ckpt (default its
latest.pt), --feat-blocks, --trunk-blocks (target depths), --out.
To specialize or refine a trained net (this is how the Kensou specialist and the efficient generalist were made), point a new config at a seed checkpoint:
# Kensou-only specialist, fine-tuned from the generalist:
bash scripts/train_auto.sh configs/ppo_kyo_kensou_specialist.yaml kensou_spec 4 \
runs/net2net2_20260708_061012/checkpoints/deploy.pt
# "Win cleaner" efficiency fine-tune (2x damage-taken penalty + win_health_bonus):
bash scripts/train_auto.sh configs/ppo_kyo_efficiency.yaml efficiency 6 \
runs/net2net2_20260708_061012/checkpoints/deploy.ptscripts/watch.py loads a checkpoint and plays it live (or captures a
video). It reads the run's config.yaml, so opponents/obs/actions match training.
# Live MAME window (WSLg), 3 episodes vs random pool opponents, default latest.pt:
python scripts/watch.py --run runs/<dir>
# Record a normal-speed clip WITH SOUND (MAME renders + captures its own AVI -> mp4):
python scripts/watch.py --run runs/<dir> --checkpoint model.pt --record out.mp4
# Pin a specific opponent and use the deploy weights:
python scripts/watch.py --run runs/<dir> --savestate clean_07 --ema
# Recording that opens on the fight (skip the ~6.3s VS/ROUND/GO banner):
python scripts/watch.py --run runs/<dir> --record out.mp4 --intro-skip 6.3
# Add variety across takes WITHOUT losing (randomize only among near-tied moves):
python scripts/watch.py --run runs/<dir> --record out.mp4 --tie-tol 0.25 --seed 3| Flag | Default | Meaning |
|---|---|---|
--run PATH |
(required) | Run directory (reads its config.yaml). |
--checkpoint PATH |
<run>/checkpoints/latest.pt |
Which .pt to play. |
--ema |
off | Play the smoothed EMA shadow weights (often the best deploy policy). |
--episodes N |
3 |
Number of episodes to play. |
--savestate NAME |
random from pool | Pin the starting opponent (e.g. clean_07, hard_o4). |
--round |
off | Force single-fight mode (episode_mode=round). |
--record PATH.mp4 |
— | Normal-speed capture with audio (MAME AVI → mp4). |
--video PATH.mp4 |
— | Silent frame-dump via rgb_array (faster, no sound). |
--intro-skip SECS |
0 |
When recording, drop the round-intro overlay so the clip opens on the live fight (~6.3). |
--post-frames N |
150 |
After a KO, keep emulating N frames so the KO + victory pose gets recorded. |
--deterministic |
off | Argmax the policy (no sampling). |
--temperature T |
0.0 |
Sample at temperature T for variety (low = mostly the strong move; weaker overall). |
--tie-tol F |
0.0 |
Randomize only among near-tied actions (within fraction F of the top action) — varied play that still wins. Overrides --temperature. |
--seed N |
20000 |
Seed (vary it with --tie-tol/--temperature for different fights). |
--base-port N |
config | Override env base port (avoid clashing with a live training run). |
--mock |
off | Mock backend (no MAME) — for debugging the loop. |
--record vs --video: --record lets MAME render and grab its own AVI at
normal speed with sound (best for shareable clips); --video dumps rgb_array
frames to a silent mp4 (faster, headless-friendly). Use one or the other.
Variety without losing: --tie-tol keeps the argmax whenever the policy is
confident and only randomizes when two moves are near-tied — so each --seed plays
a slightly different but still-winning match. --temperature gives more variety but
weaker play. --deterministic is the most repeatable.
The agent plays against the CPU arcade opponents. A single fine-tuned net can't hold all 10 — any training that cracks the last holdout (Sie Kensou) drags the shared weights off the other matchups (catastrophic interference, confirmed even when Kensou's reward change is isolated). So deployment is a generalist + specialist dispatched by a router.
| Role | Path | Covers |
|---|---|---|
| Efficient generalist (default) | runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt |
9 of 10, cleaner: damage taken 40→18 (−54%), time −35% vs deploy.pt. |
| Generalist (fallback) | runs/net2net2_20260708_061012/checkpoints/deploy.pt |
The original 9-winning 6.6M-param net. |
| Kensou specialist | runs/kensou_spec_20260708_182902/checkpoints/kensou_specialist.pt |
Sie Kensou only (id 13). |
The generalist beats 9 of the 10 pool opponents at 100%: Terry, Andy, Joe, Blue Mary, Yamazaki, Billy, Chin, Athena, and boss Chizuru.
scripts/router_eval.py reads opp_character live from the
observation and dispatches Kensou to the specialist and every other opponent to the
generalist — 10/10, including the boss:
python scripts/router_eval.py \
--run runs/net2net2_20260708_061012 \
--deploy runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
--specialist runs/kensou_spec_20260708_182902/checkpoints/kensou_specialist.pt \
--episodes 2 # -> OVERALL: 10.0/10 at 100%router_eval.py flags: --run (env/arch config), --deploy, --specialist
(all required), --kensou-id (default 13), --opponents (default = the 10 pool
states), --episodes (default 2), --deploy-ema / --spec-ema (use EMA weights),
--base-port.
# Win rate against a chosen set of opponents (one MAME run per opponent):
python scripts/winrate.py --run runs/net2net2_20260708_061012 \
--checkpoint runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
--opponents clean_01 clean_02 clean_03 clean_04 clean_05 \
clean_06 clean_07 clean_08 clean_09 clean_10 \
--episodes 8
# Aggregate reward/score over N episodes vs random pool opponents:
python scripts/evaluate.py --run runs/<dir> --episodes 10 [--ema] [--stochastic]winrate.py flags: --run (req), --checkpoint, --opponents (default a hard_o*
set), --episodes (8), --ema, --base-port. evaluate.py flags: --run (req),
--checkpoint, --episodes (10), --ema, --stochastic, --mock.
Which checkpoint should I use? For a clean single-net demo, the
efficient_generalist.pt (fast, low-damage, 9/10). For a guaranteed 10/10
including Kensou, run the router. deploy.pt is kept as a fallback generalist.
obs_mode: pixels->{"frame": (H, W, stack) uint8}(grayscale, stacked)obs_mode: features->{"features": (18,) float32 in [-1, 1]}(health, position, power, stocks, character, action, timer, wins for both fighters)obs_mode: both-> both keys (what the shipped model uses)- Actions:
Discrete(17)(no-op, 8 directions, A/B/C/D, A+B, C+D, ↓A, →C) orMultiDiscrete([9, 7])(direction x attack, allows simultaneous move+attack) - Reward (
KOFRewardWrapper, configenv.reward): damage dealt − damage taken (normalized by max health), win bonus, special/super bonuses, whiff penalty, optional time penalty andwin_health_bonus(reward scaled by HP left at the KO — the lever behind the efficient generalist). Round-boundary health resets are ignored.
Create kof97rl/algos/my_algo.py:
from kof97rl.algos.base import ActOutput, Agent
from kof97rl.algos.registry import register
@register("my_algo")
class MyAlgo(Agent):
train_mode = "on_policy"
def act(self, obs, deterministic=False):
... # obs is a dict of batched arrays
return ActOutput(action=actions, extras={"log_prob": ..., "value": ...})
def update(self, buffer): # RolloutBuffer
...
return {"loss/policy": loss.item()} # metrics -> TensorBoardThen set algo.name: my_algo in a config (or --algo my_algo). Use
kof97rl/algos/ppo/ as the reference. On-policy agents return log_prob and
value in ActOutput.extras and may expose n_steps, gamma, gae_lambda, and
a value(obs) method for GAE bootstrapping.
pytest # no emulator needed (mock backend)tests/test_ppo_smoke.py trains PPO on the mock env and asserts reward improves —
run it after touching anything in the training path.
configs/ base.yaml + ppo_kyo_{net2net,kensou_specialist,efficiency}.yaml
kof97rl/envs/ spaces contract, actions, mock env, wrappers, factory
kof97rl/envs/mame/ bridge.lua, protocol, process, MameKofEnv, specs/*.yaml
kof97rl/algos/ base + registry + ppo/ <- add yours here
kof97rl/training/ rollout buffer, OnPolicyTrainer, evaluator
scripts/ train, train_auto.sh, evaluate, winrate, watch, router_eval,
make_savestate, capture_roster, net2net_deeper




