Skip to content

Latest commit

 

History

46 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Reinforcement Learning · King of Fighters '97

package kof97rl

RL framework for KOF'97 (Neo Geo arcade) with a self-hosted MAME Lua bridge — no accounts, no Docker. It ships a fully-worked PPO agent that trains Kyo Kusanagi to beat all 10 sampled arcade opponents including the boss (via a generalist + a Kensou specialist dispatched by an opp_character router), plus everything around it: the emulator bridge, observation/action contract, reward shaping, rollout buffer, training loop with LR/entropy annealing + weight EMA, and evaluation/recording tools.

PPO agent (kof97rl/algos/ppo/)
        |
   Agent interface (act / update / save / load)
        |
OnPolicyTrainer -> RolloutBuffer -> TensorBoard + CSV + checkpoints
        |
   VecEnv of KOF environments (canonical obs/action contract)
        |
mock backend (tests, debugging)   |   MAME backend (bridge.lua over TCP)

Character scope: only Kyo Kusanagi is configured in this repo — one shipped moveset (kof97rl/envs/movesets/kof97_kyo.yaml), the configs/ppo_kyo_* experiments, and savestates that seat Kyo in the P1 slot. Training another character means adding a new moveset YAML + savestates; the framework itself is character-agnostic.

Demo

Kyo clears the full 10-opponent gauntlet on a single health bar — the efficient generalist (≈54% less damage taken, ~35% faster than the first 10/10 net):

Kyo beating arcade opponents, ending in a K.O.

~11s highlight loop (silent): pressuring Blue Mary → flooring Billy Kane → a rush on Sie Kensou → a clean K.O. The full 304-second run with audio is demos/kyo_10opponents_efficient.mp4 (see below).

Stills from the run — a closer look across the roster (deterministic play, one health bar):

a clean K.O. rush combo on Sie Kensou
pressuring Blue Mary Billy Kane knocked down

A clean K.O.; a rush combo on Sie Kensou (the specialist's matchup); pressuring Blue Mary on the temple stage; Billy Kane floored. Per-opponent fights and full roster reels are generated locally into the (gitignored) demos/ folder — see Watch & record.

▶️ Embed the full 304s run with audio as an inline video (optional)

GitHub plays inline <video> only from an uploaded-asset URL — never from a file committed to the repo, and never from a private repo's raw path (which is why the GIF above is used instead). To add the full clip:

  1. Open https://github.com/xeno0/kof97_rl/issues/new — do not submit the issue.
  2. Drag demos/kyo_10opponents_efficient.mp4 into the comment box; wait for the upload.
  3. Copy the https://github.com/user-attachments/assets/<uuid> URL it inserts.
  4. Add it here as <video src="THAT_URL" controls muted width="100%"></video>.

Contents

Setup

conda activate kof97            # Python 3.11 (torch, gymnasium 0.29)
pip install -e ".[dev]"

sudo apt install mame           # emulator (0.242 here; anything 0.227+ works)

mkdir -p ~/roms                 # supply your own legally-owned ROMs:
#   ~/roms/kof97.zip                    - the game
#   ~/roms/neogeo.zip                   - Neo Geo BIOS
mame kof97 -rompath ~/roms -bench 60    # sanity: expect >= 500% speed

ROM/BIOS files are your responsibility and are gitignored — never commit them.

RAM addresses and boot/select timings live in kof97rl/envs/mame/specs/kof97.yaml. The training pool loads fight-ready savestates (each one pins a different arcade opponent at round start) from ~/.kof97rl/kof97/savestates/kof97/. The shipped pool is clean_01 … clean_10 (plus a few hard_o* / fight_ready states used during development). To (re)create one:

python scripts/make_savestate.py --game kof97 --state-name clean_01   # boots, coins in, selects, saves at round start
python scripts/make_savestate.py --game kof97 --state-name clean_01 --windowed   # watch it happen

make_savestate.py flags: --game (default kof97), --state-name (default fight_ready), --rom-path (default ~/roms), --windowed (show the MAME window).

30-second quickstart

The repo ships trained checkpoints — you can watch them fight without training anything.

conda activate kof97

# 1) Watch the shipped generalist fight a random pool opponent in a live MAME window:
python scripts/watch.py --run runs/efficiency_20260709_072914 \
    --checkpoint runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
    --episodes 3

# 2) Prove the full deploy beats ALL 10 opponents (generalist + Kensou specialist via router):
python scripts/router_eval.py \
    --run runs/net2net2_20260708_061012 \
    --deploy     runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
    --specialist runs/kensou_spec_20260708_182902/checkpoints/kensou_specialist.pt \
    --episodes 2                                   # -> OVERALL: 10.0/10 at 100%

Pre-rendered fights (per-opponent + full roster reels, with audio) are written to the local demos/ folder (gitignored — large media, not on GitHub); regenerate them with watch.py --record.


Train

Everything is driven by a YAML config. The winning recipe is a single file, configs/ppo_kyo_net2net.yaml: obs = pixels+features, single full-health rounds, a 10-opponent arcade pool, and a Net2Net-expanded network refined at low LR with LR/entropy annealing + weight EMA.

# Full training run (needs MAME + ROMs). Writes to runs/net2net2_<timestamp>/
python scripts/train.py --config configs/ppo_kyo_net2net.yaml

# Smoke test with NO emulator — the mock backend, for verifying the training path:
python scripts/train.py --config configs/ppo_kyo_net2net.yaml --mock

# Robust multi-hour launcher: auto-resumes from the newest checkpoint on any crash.
# Usage: train_auto.sh <config> <run_name_glob> <hours> [seed_ckpt]
bash scripts/train_auto.sh configs/ppo_kyo_net2net.yaml net2net 12

train.py flags

Flag Meaning
--config PATH Experiment YAML (layered on top of configs/base.yaml).
--mock Force the mock backend — trains with no MAME/ROMs (fast smoke test).
--init-from PATH Warm-start weights from a checkpoint (fine-tune / continue).
--reset-progress With --init-from, load the weights but restart the anneal schedule at step 0 (for seeding a new run from an external checkpoint).
--total-timesteps N Override the config's training horizon (also drives the LR/entropy anneal horizon).
--num-envs N Override parallel emulator count.
--seed N, --device cuda|cpu Override seed / device.
--algo NAME Override algo.name (e.g. your own registered algo).
--list-algos Print the registered algorithms and exit.

What the config controls

Key knobs in configs/ppo_kyo_net2net.yaml:

  • env.obs_mode: both — pixels and the 18-dim feature vector (what the shipped model uses).
  • env.action_mode: discrete — Discrete(17).
  • env.episode_mode: round — one full-health fight per episode (use match to climb the whole arcade ladder in one episode).
  • env.savestate_pool: [clean_01 … clean_10] — the 10 sampled opponents; each reset() draws one.
  • env.num_envs: 24 — parallel emulators. Each rank gets its own MAME process on base_port + rank.
  • env.reward: — damage-dealt/taken weights, win/special/super bonuses, whiff penalty (see reward).
  • algo.params: — PPO hyperparameters + the Net2Net widths (feat_blocks, trunk_blocks, hidden_dim), learning_rate, ent_coef, anneal_lr, ema_decay.
  • train.total_timesteps / train.checkpoint_every — horizon and checkpoint cadence.

Outputs & monitoring

Each run creates runs/<run_name>_<timestamp>/ containing:

  • config.yaml — the exact resolved config (every eval/watch tool reads this back).
  • checkpoints/latest.pt, periodic step_*.pt, and the EMA shadow inside each .pt.
  • TensorBoard event files + a CSV of scalar metrics.
tensorboard --logdir runs        # loss, entropy, per-opponent reward, LR/entropy schedules

The Net2Net warm-start (how the generalist was made)

PPO on a mixed opponent pool suffers catastrophic interference — it trades which opponents it can beat between checkpoints. The fix that produced deploy.pt, one net holding 9 of the 10 arcade opponents (the 10th, Kensou, is handled by the specialist + router — see Play the trained agent):

  1. Train a Phase-1 policy to stability with LR/entropy annealing + weight EMA.
  2. Function-preserving expansion (scripts/net2net_deeper.py): grow the EMA policy into a deeper net by inserting residual blocks whose second layer is zero-initialized — the bigger net reproduces the source exactly (verified max|Δlogits| < 1e-4), so no capability is lost at the seam.
  3. Refine the expanded net at low LR (1e-4) + low entropy (0.005) so the extra capacity learns the last holdouts without eroding the warm-started foundation.
python scripts/net2net_deeper.py --src-run runs/<phase1_run> \
    --feat-blocks 10 --trunk-blocks 6 --out runs/net2net_seed/deep_from_phase1.pt
bash scripts/train_auto.sh configs/ppo_kyo_net2net.yaml net2net 12 \
    runs/net2net_seed/deep_from_phase1.pt          # 4th arg = seed checkpoint

net2net_deeper.py flags: --src-run (source run dir), --src-ckpt (default its latest.pt), --feat-blocks, --trunk-blocks (target depths), --out.

Fine-tuning an existing checkpoint

To specialize or refine a trained net (this is how the Kensou specialist and the efficient generalist were made), point a new config at a seed checkpoint:

# Kensou-only specialist, fine-tuned from the generalist:
bash scripts/train_auto.sh configs/ppo_kyo_kensou_specialist.yaml kensou_spec 4 \
    runs/net2net2_20260708_061012/checkpoints/deploy.pt

# "Win cleaner" efficiency fine-tune (2x damage-taken penalty + win_health_bonus):
bash scripts/train_auto.sh configs/ppo_kyo_efficiency.yaml efficiency 6 \
    runs/net2net2_20260708_061012/checkpoints/deploy.pt

Watch & record

scripts/watch.py loads a checkpoint and plays it live (or captures a video). It reads the run's config.yaml, so opponents/obs/actions match training.

# Live MAME window (WSLg), 3 episodes vs random pool opponents, default latest.pt:
python scripts/watch.py --run runs/<dir>

# Record a normal-speed clip WITH SOUND (MAME renders + captures its own AVI -> mp4):
python scripts/watch.py --run runs/<dir> --checkpoint model.pt --record out.mp4

# Pin a specific opponent and use the deploy weights:
python scripts/watch.py --run runs/<dir> --savestate clean_07 --ema

# Recording that opens on the fight (skip the ~6.3s VS/ROUND/GO banner):
python scripts/watch.py --run runs/<dir> --record out.mp4 --intro-skip 6.3

# Add variety across takes WITHOUT losing (randomize only among near-tied moves):
python scripts/watch.py --run runs/<dir> --record out.mp4 --tie-tol 0.25 --seed 3

watch.py flags

Flag Default Meaning
--run PATH (required) Run directory (reads its config.yaml).
--checkpoint PATH <run>/checkpoints/latest.pt Which .pt to play.
--ema off Play the smoothed EMA shadow weights (often the best deploy policy).
--episodes N 3 Number of episodes to play.
--savestate NAME random from pool Pin the starting opponent (e.g. clean_07, hard_o4).
--round off Force single-fight mode (episode_mode=round).
--record PATH.mp4 — Normal-speed capture with audio (MAME AVI → mp4).
--video PATH.mp4 — Silent frame-dump via rgb_array (faster, no sound).
--intro-skip SECS 0 When recording, drop the round-intro overlay so the clip opens on the live fight (~6.3).
--post-frames N 150 After a KO, keep emulating N frames so the KO + victory pose gets recorded.
--deterministic off Argmax the policy (no sampling).
--temperature T 0.0 Sample at temperature T for variety (low = mostly the strong move; weaker overall).
--tie-tol F 0.0 Randomize only among near-tied actions (within fraction F of the top action) — varied play that still wins. Overrides --temperature.
--seed N 20000 Seed (vary it with --tie-tol/--temperature for different fights).
--base-port N config Override env base port (avoid clashing with a live training run).
--mock off Mock backend (no MAME) — for debugging the loop.

--record vs --video: --record lets MAME render and grab its own AVI at normal speed with sound (best for shareable clips); --video dumps rgb_array frames to a silent mp4 (faster, headless-friendly). Use one or the other.

Variety without losing: --tie-tol keeps the argmax whenever the policy is confident and only randomizes when two moves are near-tied — so each --seed plays a slightly different but still-winning match. --temperature gives more variety but weaker play. --deterministic is the most repeatable.


Play the trained agent (the 10/10 deploy)

The agent plays against the CPU arcade opponents. A single fine-tuned net can't hold all 10 — any training that cracks the last holdout (Sie Kensou) drags the shared weights off the other matchups (catastrophic interference, confirmed even when Kensou's reward change is isolated). So deployment is a generalist + specialist dispatched by a router.

Shipped checkpoints

Role Path Covers
Efficient generalist (default) runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt 9 of 10, cleaner: damage taken 40→18 (−54%), time −35% vs deploy.pt.
Generalist (fallback) runs/net2net2_20260708_061012/checkpoints/deploy.pt The original 9-winning 6.6M-param net.
Kensou specialist runs/kensou_spec_20260708_182902/checkpoints/kensou_specialist.pt Sie Kensou only (id 13).

The generalist beats 9 of the 10 pool opponents at 100%: Terry, Andy, Joe, Blue Mary, Yamazaki, Billy, Chin, Athena, and boss Chizuru.

The router → 10/10

scripts/router_eval.py reads opp_character live from the observation and dispatches Kensou to the specialist and every other opponent to the generalist — 10/10, including the boss:

python scripts/router_eval.py \
    --run        runs/net2net2_20260708_061012 \
    --deploy     runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
    --specialist runs/kensou_spec_20260708_182902/checkpoints/kensou_specialist.pt \
    --episodes 2                                   # -> OVERALL: 10.0/10 at 100%

router_eval.py flags: --run (env/arch config), --deploy, --specialist (all required), --kensou-id (default 13), --opponents (default = the 10 pool states), --episodes (default 2), --deploy-ema / --spec-ema (use EMA weights), --base-port.

Per-opponent win rate & aggregate score

# Win rate against a chosen set of opponents (one MAME run per opponent):
python scripts/winrate.py --run runs/net2net2_20260708_061012 \
    --checkpoint runs/efficiency_20260709_072914/checkpoints/efficient_generalist.pt \
    --opponents clean_01 clean_02 clean_03 clean_04 clean_05 \
                clean_06 clean_07 clean_08 clean_09 clean_10 \
    --episodes 8

# Aggregate reward/score over N episodes vs random pool opponents:
python scripts/evaluate.py --run runs/<dir> --episodes 10 [--ema] [--stochastic]

winrate.py flags: --run (req), --checkpoint, --opponents (default a hard_o* set), --episodes (8), --ema, --base-port. evaluate.py flags: --run (req), --checkpoint, --episodes (10), --ema, --stochastic, --mock.

Which checkpoint should I use? For a clean single-net demo, the efficient_generalist.pt (fast, low-damage, 9/10). For a guaranteed 10/10 including Kensou, run the router. deploy.pt is kept as a fallback generalist.


Observations, actions, reward

  • obs_mode: pixels -> {"frame": (H, W, stack) uint8} (grayscale, stacked)
  • obs_mode: features -> {"features": (18,) float32 in [-1, 1]} (health, position, power, stocks, character, action, timer, wins for both fighters)
  • obs_mode: both -> both keys (what the shipped model uses)
  • Actions: Discrete(17) (no-op, 8 directions, A/B/C/D, A+B, C+D, ↓A, →C) or MultiDiscrete([9, 7]) (direction x attack, allows simultaneous move+attack)
  • Reward (KOFRewardWrapper, config env.reward): damage dealt − damage taken (normalized by max health), win bonus, special/super bonuses, whiff penalty, optional time penalty and win_health_bonus (reward scaled by HP left at the KO — the lever behind the efficient generalist). Round-boundary health resets are ignored.

Add your own algorithm

Create kof97rl/algos/my_algo.py:

from kof97rl.algos.base import ActOutput, Agent
from kof97rl.algos.registry import register

@register("my_algo")
class MyAlgo(Agent):
    train_mode = "on_policy"

    def act(self, obs, deterministic=False):
        ...                            # obs is a dict of batched arrays
        return ActOutput(action=actions, extras={"log_prob": ..., "value": ...})

    def update(self, buffer):          # RolloutBuffer
        ...
        return {"loss/policy": loss.item()}  # metrics -> TensorBoard

Then set algo.name: my_algo in a config (or --algo my_algo). Use kof97rl/algos/ppo/ as the reference. On-policy agents return log_prob and value in ActOutput.extras and may expose n_steps, gamma, gae_lambda, and a value(obs) method for GAE bootstrapping.

Tests

pytest          # no emulator needed (mock backend)

tests/test_ppo_smoke.py trains PPO on the mock env and asserts reward improves — run it after touching anything in the training path.

Layout

configs/          base.yaml + ppo_kyo_{net2net,kensou_specialist,efficiency}.yaml
kof97rl/envs/     spaces contract, actions, mock env, wrappers, factory
kof97rl/envs/mame/  bridge.lua, protocol, process, MameKofEnv, specs/*.yaml
kof97rl/algos/    base + registry + ppo/   <- add yours here
kof97rl/training/ rollout buffer, OnPolicyTrainer, evaluator
scripts/          train, train_auto.sh, evaluate, winrate, watch, router_eval,
                  make_savestate, capture_roster, net2net_deeper

About

Reinforcement Learning for King of Fighters '97

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages