AMD AI Hackathon — Track 2: Video Captioning Agent
A Dockerized video-captioning agent that turns short clips into four style-specific commentaries. The pipeline first extracts grounded visual facts from representative frames, then generates concise captions in formal, sarcastic, humorous-tech, and humorous-non-tech voices.
# 1. Build
docker build -t dragon-commentary .
# 2. Prepare input
mkdir -p input output
cat > input/tasks.json <<EOF
[
{
"task_id": "v1",
"video_url": "https://storage.googleapis.com/amd-hackathon-clips/1860079-uhd_2560_1440_25fps.mp4",
"styles": ["formal", "sarcastic", "humorous_tech", "humorous_non_tech"]
}
]
EOF
# 3. Run
docker run --rm \
-v $(pwd)/input:/input \
-v $(pwd)/output:/output \
-e FIREWORKS_API_KEY="your-key-here" \
dragon-commentary
# 4. Check results
cat output/results.jsonRequired env vars:
| Variable | Description |
|---|---|
FIREWORKS_API_KEY |
Fireworks AI API key |
Model architecture (two models):
| Role | Default | Purpose |
|---|---|---|
| Vision | kimi-k2p6 |
Per-frame visual fact extraction (vision-capable) |
| Text | deepseek-v4-pro |
Caption generation (text-only, faster) |
Override any model with these env vars:
Optional env vars:
| Variable | Default | Description |
|---|---|---|
FIREWORKS_MODEL |
accounts/fireworks/models/kimi-k2p6 |
Default (used for both if no override) |
FIREWORKS_VISION_MODEL |
(same as FIREWORKS_MODEL) | Single-frame vision fact extraction |
FIREWORKS_TEXT_MODEL |
accounts/fireworks/models/deepseek-v4-pro |
Caption generation |
FIREWORKS_REVIEW_MODEL |
(same as vision model) | Backward-compatible config; not used by the deterministic vision path |
FIREWORKS_VALIDATION_MODEL |
accounts/fireworks/models/deepseek-v4-pro |
Caption quality validation |
INPUT_PATH |
/input/tasks.json |
Task file path |
OUTPUT_PATH |
/output/results.json |
Output file path |
CACHE_DIR |
/cache |
Cache directory |
WORK_DIR |
/work |
Working directory |
LOG_LEVEL |
INFO |
Logging level |
ENABLE_CACHING |
true |
Enable disk caching |
MAX_TOTAL_DURATION_SEC |
600 |
Batch/task runtime target (10 min); Phase B reserves 30s for shutdown/output |
MAX_VISION_FRAMES |
24 |
Max frames sent to vision model |
VISION_BATCH_SIZE |
16 |
Legacy compatibility setting; all selected frames are now sent in one vision request |
LORA_MODELS |
{} |
JSON dict mapping style keys to LoRA model IDs (e.g. {"formal":"ft-model-id"}) |
USE_LORA |
false |
Enable per-style LoRA model routing (experimental); default uses base text model |
UNREPEATED_TOPIC |
false |
When enabled, non-formal dragons avoid repeating object/OCR topics already covered by another dragon |
The Track 2 score is driven by caption accuracy and style match across hidden clips. The default batch path is tuned for that evaluation shape:
- Accuracy first: Kimi analyzes selected frames into constrained per-frame tagged fact blocks. Python matches each block by frame ID and deterministically builds the canonical JSON used by every caption.
- Style without losing facts: Each dragon receives the same canonical facts but a different persona/style prompt. Humorous tech analogies are allowed only when anchored to real visible details.
- Completeness under timeout: Docker counts the video tasks in
/input/tasks.jsonbefore processing. When Phase B starts, it targetsMAX_TOTAL_DURATION_SEC - 30sfor shutdown/output, divides the remaining batch time by the number of input video tasks, and gives that task-level retry budget to all dragons for that video. If validation fails and the Phase B retry budget is gone, the last usable generated commentary is returned; local fallback is reserved for cases where no usable model commentary exists. - Parallel by default: Docker defaults to
UNREPEATED_TOPIC=false, so style captions run concurrently usingMAX_WORKERS. The Phase B retry budget is divided by video task count, not by dragon/style count. Optional topic de-duplication can be enabled, but it runs sequentially because later dragons depend on earlier approved topics. - Observable model routing: Logs show each frame-analysis result and which backend/model is called for vision, caption generation, validation, and Whisper.
- LoRA is optional: Fine-tuned adapters were explored and remain routable with
USE_LORA=true, but the default batch path uses the base Fireworks text model for reliability on hidden clips.
[
{
"task_id": "v1",
"video_url": "https://example.com/clip.mp4",
"styles": ["formal", "sarcastic", "humorous_tech", "humorous_non_tech"]
}
][
{
"task_id": "v1",
"captions": {
"formal": "A professional, factual description of the video content.",
"sarcastic": "A dry, lightly mocking take on what's happening.",
"humorous_tech": "A funny caption with programming/tech references.",
"humorous_non_tech": "Everyday humor with no technical jargon."
}
}
]This hackathon arrived at the perfect moment. I've been developing Caro5, a gomoku-like board game, and this captioning pipeline maps directly onto a game review system:
| Hackathon Pipeline | Caro5 Review |
|---|---|
| Video frames | Board position snapshots |
| Scene detection | Turn-by-turn move detection |
| Vision analysis (kimi-k2p6) | Position evaluation (engine) |
| Four dragon personalities | Multiple reviewer personas (formal coach, casual friend, rival) |
| Validation critic loop | Analysis quality scoring |
| Web UI with timed overlay | Interactive move-by-move review |
The same two-phase architecture can analyze a Caro5 game instead of a video clip: Phase A would extract board-state facts, and Phase B would turn those facts into styled coaching commentary for new players and experienced reviewers.
- ✅ Reads
/input/tasks.json, writes/output/results.json - ✅ Supports all four required styles (
formal,sarcastic,humorous_tech,humorous_non_tech) - ✅ 10-minute runtime target with Phase B task-count retry budgeting and 30s shutdown/output reserve
- ✅ Exit code 0 on success, non-zero on failure
- ✅ Valid JSON output with all styles per clip
- ✅ Image size 1.37GB (under 10GB limit)
- ✅ Generalizes beyond example clips
tasks.json
↓
Downloader → FFmpeg → Scene Detection → Representative Frames
↓
Phase A: single-frame VLM facts
↓
Deterministic Normalize + Merge ┐
├→ Merger → Canonical JSON
Audio Track → Whisper Transcript ───────────────────────────┘
↓
Phase B: styled caption generation
↓
Caption Validator (retry loop)
↓
results.json
- Phase A (Vision): kimi-k2p6 receives the selected representative frames in one request and returns a constrained facts block for each frame. Python matches each block by frame ID, then parses, normalizes, and merges the observations into a structured schema for subjects, actions, objects, OCR, camera, and environment.
- Whisper: Audio transcription via
faster-whisper - Merger: Combines visual observations and audio transcript into a unified canonical JSON
- Phase B (Text): deepseek-v4-pro turns the canonical JSON into four dragon-voiced captions: Trí Long, Lão Quân, Gemma Trí Tech, and Du Rong. Styles run in parallel by default, and each style can optionally route to a dedicated LoRA model through
LORA_MODELSwith fallback to the base text model. - Validation: Each caption is scored across 5 weighted dimensions (accuracy 35%, style_fit 25%, specificity 20%, hallucination_control 15%, completeness 5%) by an LLM judge; captions below threshold trigger revision with specific feedback while Phase B retry time remains. The judge prompt distinguishes unsupported literal facts from grounded style metaphors. If the retry budget is exhausted after a validation failure, the last usable generated commentary is kept; if no usable commentary exists, the system falls back to a grounded local caption.
- Web UI Quick Chat: The local web UI samples Modal-generated intro/waiting lines per dragon. Each selected dragon speaks one intro at job start, then the chat switches to non-repeating markdown canned responses plus Modal waiting lines with speaker cooldowns to keep one dragon from dominating.
- Timeout and retry budget: The Docker batch path counts
/input/tasks.jsonvideo tasks up front. Phase B reserves 30 seconds for shutdown/output, then splits the remaining time-to-target across video tasks. Dragons for a single video share that task-level budget because they run in parallel.
pip install -r requirements.txt
# Run the CLI agent
python -m app.main
# Or run the web UI
uvicorn app.web:app --reload --host 127.0.0.1 --port 5173Create a .env file in the project root with your FIREWORKS_API_KEY to power the pipeline outside Docker.
Each dragon style can route to a dedicated Qwen3-1.7B LoRA adapter trained through the Fireworks SFT API. These adapters are optional; the batch path can fall back to the base text model whenever USE_LORA=false or a style-specific adapter is unavailable.
| Style | LoRA Model ID |
|---|---|
formal |
accounts/oak42000-.../models/ft-z5npxv9s-l12gp |
sarcastic |
accounts/oak42000-.../models/ft-hrbnfpjn-sts6p |
humorous_tech |
accounts/oak42000-.../models/ft-nt5fc352-ol5kf |
humorous_non_tech |
accounts/oak42000-.../models/ft-du-rong-base-2k-v4-run001 |
Dataset generation pipeline (lora-ft/):
generate_prompts.py— 6,200 prompt variants across 6 categoriesgenerate_data.py— kimi-k2p6 generates in-character responses (11,200 total)validate_data.py— LLM judge rejects low-quality samples (< 7/10)prepare_data.py— Formats into OpenAI chat JSONLfinetune.py— Submits to Fireworks SFT (rank 16, 3 epochs, lr 1e-4)
Set LORA_MODELS in .env to route captions through fine-tuned adapters.
The same Modal-generated training data also feeds the web UI quick-chat pool. scripts/extract_modal_quick_chat.py extracts clean introduction and waiting examples into assets/dragons/quick_chat/modal_quick_chat.json; /api/dragons/quick-chat?limit=200 returns categorized samples per dragon on each page load. The browser shows one intro per selected dragon, then removes waiting/canned messages from its page-local pool as they are used. If a dragon speaks twice in a row, it cools down until the other selected dragons have spoken.
pytest tests/ -v| Style | Dragon | Tone |
|---|---|---|
formal |
Trí Long | Professional, objective, factual data analyst |
sarcastic |
Lão Quân | Dry, ironic, lightly mocking ancient dragoness |
humorous_tech |
Gemma Trí Tech | Funny, tech/programming references, dragoness engineer |
humorous_non_tech |
Du Rong | Funny, everyday humour, young dragon with fake ancient wisdom |
The project includes a Reveal.js slide deck at htdocs/presentation.html covering the architecture, dragon personas, dataset pipeline, and key features.
app/ # Core pipeline components
├── main.py # CLI entrypoint
├── pipeline.py # Orchestrator
├── downloader.py # Video downloader
├── extractor.py # FFmpeg frame/audio extraction
├── scenes.py # Scene detection
├── sampler.py # Adaptive frame sampling
├── vision.py # Phase A — VLM analysis
├── whisper.py # Audio transcription
├── merger.py # Merge vision + audio
├── caption.py # Phase B — style-specific caption gen
├── validator.py # Caption quality validation
├── prompts.py # LLM prompt templates
├── styles.py # Dragon persona loader
├── config.py # Configuration
├── cache.py # Disk caching
├── utils.py # JSON repair, helpers
├── web.py # FastAPI web server
└── resources/ # Style reference examples
assets/dragons/ # Dragon profiles, sprites, animations
└── quick_chat/ # Modal-generated intro/waiting chat corpus
htdocs/ # Web UI frontend
tests/ # Test suite
Dockerfile # Competition build