Skip to content

Latest commit

 

History

143 Commits

Folders and files

Repository files navigation

Dragon Commentary Studio

AMD AI Hackathon — Track 2: Video Captioning Agent

A Dockerized video-captioning agent that turns short clips into four style-specific commentaries. The pipeline first extracts grounded visual facts from representative frames, then generates concise captions in formal, sarcastic, humorous-tech, and humorous-non-tech voices.

Quick Start (Docker)

# 1. Build
docker build -t dragon-commentary .

# 2. Prepare input
mkdir -p input output
cat > input/tasks.json <<EOF
[
  {
    "task_id": "v1",
    "video_url": "https://storage.googleapis.com/amd-hackathon-clips/1860079-uhd_2560_1440_25fps.mp4",
    "styles": ["formal", "sarcastic", "humorous_tech", "humorous_non_tech"]
  }
]
EOF

# 3. Run
docker run --rm \
  -v $(pwd)/input:/input \
  -v $(pwd)/output:/output \
  -e FIREWORKS_API_KEY="your-key-here" \
  dragon-commentary

# 4. Check results
cat output/results.json

Required env vars:

Variable Description
FIREWORKS_API_KEY Fireworks AI API key

Model architecture (two models):

Role Default Purpose
Vision kimi-k2p6 Per-frame visual fact extraction (vision-capable)
Text deepseek-v4-pro Caption generation (text-only, faster)

Override any model with these env vars:

Optional env vars:

Variable Default Description
FIREWORKS_MODEL accounts/fireworks/models/kimi-k2p6 Default (used for both if no override)
FIREWORKS_VISION_MODEL (same as FIREWORKS_MODEL) Single-frame vision fact extraction
FIREWORKS_TEXT_MODEL accounts/fireworks/models/deepseek-v4-pro Caption generation
FIREWORKS_REVIEW_MODEL (same as vision model) Backward-compatible config; not used by the deterministic vision path
FIREWORKS_VALIDATION_MODEL accounts/fireworks/models/deepseek-v4-pro Caption quality validation
INPUT_PATH /input/tasks.json Task file path
OUTPUT_PATH /output/results.json Output file path
CACHE_DIR /cache Cache directory
WORK_DIR /work Working directory
LOG_LEVEL INFO Logging level
ENABLE_CACHING true Enable disk caching
MAX_TOTAL_DURATION_SEC 600 Batch/task runtime target (10 min); Phase B reserves 30s for shutdown/output
MAX_VISION_FRAMES 24 Max frames sent to vision model
VISION_BATCH_SIZE 16 Legacy compatibility setting; all selected frames are now sent in one vision request
LORA_MODELS {} JSON dict mapping style keys to LoRA model IDs (e.g. {"formal":"ft-model-id"})
USE_LORA false Enable per-style LoRA model routing (experimental); default uses base text model
UNREPEATED_TOPIC false When enabled, non-formal dragons avoid repeating object/OCR topics already covered by another dragon

Evaluation Notes

The Track 2 score is driven by caption accuracy and style match across hidden clips. The default batch path is tuned for that evaluation shape:

  • Accuracy first: Kimi analyzes selected frames into constrained per-frame tagged fact blocks. Python matches each block by frame ID and deterministically builds the canonical JSON used by every caption.
  • Style without losing facts: Each dragon receives the same canonical facts but a different persona/style prompt. Humorous tech analogies are allowed only when anchored to real visible details.
  • Completeness under timeout: Docker counts the video tasks in /input/tasks.json before processing. When Phase B starts, it targets MAX_TOTAL_DURATION_SEC - 30s for shutdown/output, divides the remaining batch time by the number of input video tasks, and gives that task-level retry budget to all dragons for that video. If validation fails and the Phase B retry budget is gone, the last usable generated commentary is returned; local fallback is reserved for cases where no usable model commentary exists.
  • Parallel by default: Docker defaults to UNREPEATED_TOPIC=false, so style captions run concurrently using MAX_WORKERS. The Phase B retry budget is divided by video task count, not by dragon/style count. Optional topic de-duplication can be enabled, but it runs sequentially because later dragons depend on earlier approved topics.
  • Observable model routing: Logs show each frame-analysis result and which backend/model is called for vision, caption generation, validation, and Whisper.
  • LoRA is optional: Fine-tuned adapters were explored and remain routable with USE_LORA=true, but the default batch path uses the base Fireworks text model for reliability on hidden clips.

Input Format (/input/tasks.json)

[
  {
    "task_id": "v1",
    "video_url": "https://example.com/clip.mp4",
    "styles": ["formal", "sarcastic", "humorous_tech", "humorous_non_tech"]
  }
]

Output Format (/output/results.json)

[
  {
    "task_id": "v1",
    "captions": {
      "formal": "A professional, factual description of the video content.",
      "sarcastic": "A dry, lightly mocking take on what's happening.",
      "humorous_tech": "A funny caption with programming/tech references.",
      "humorous_non_tech": "Everyday humor with no technical jargon."
    }
  }
]

Why This Matters — Caro5 Game Review

This hackathon arrived at the perfect moment. I've been developing Caro5, a gomoku-like board game, and this captioning pipeline maps directly onto a game review system:

Hackathon Pipeline Caro5 Review
Video frames Board position snapshots
Scene detection Turn-by-turn move detection
Vision analysis (kimi-k2p6) Position evaluation (engine)
Four dragon personalities Multiple reviewer personas (formal coach, casual friend, rival)
Validation critic loop Analysis quality scoring
Web UI with timed overlay Interactive move-by-move review

The same two-phase architecture can analyze a Caro5 game instead of a video clip: Phase A would extract board-state facts, and Phase B would turn those facts into styled coaching commentary for new players and experienced reviewers.

Competition Rules

  • ✅ Reads /input/tasks.json, writes /output/results.json
  • ✅ Supports all four required styles (formal, sarcastic, humorous_tech, humorous_non_tech)
  • ✅ 10-minute runtime target with Phase B task-count retry budgeting and 30s shutdown/output reserve
  • ✅ Exit code 0 on success, non-zero on failure
  • ✅ Valid JSON output with all styles per clip
  • ✅ Image size 1.37GB (under 10GB limit)
  • ✅ Generalizes beyond example clips

Architecture

tasks.json
    ↓
Downloader → FFmpeg → Scene Detection → Representative Frames
                                           ↓
                            Phase A: single-frame VLM facts
                                           ↓
                            Deterministic Normalize + Merge ┐
                                                             ├→ Merger → Canonical JSON
Audio Track → Whisper Transcript ───────────────────────────┘
                                                             ↓
                            Phase B: styled caption generation
                                                             ↓
                            Caption Validator (retry loop)
                                                             ↓
results.json
  • Phase A (Vision): kimi-k2p6 receives the selected representative frames in one request and returns a constrained facts block for each frame. Python matches each block by frame ID, then parses, normalizes, and merges the observations into a structured schema for subjects, actions, objects, OCR, camera, and environment.
  • Whisper: Audio transcription via faster-whisper
  • Merger: Combines visual observations and audio transcript into a unified canonical JSON
  • Phase B (Text): deepseek-v4-pro turns the canonical JSON into four dragon-voiced captions: Trí Long, Lão Quân, Gemma Trí Tech, and Du Rong. Styles run in parallel by default, and each style can optionally route to a dedicated LoRA model through LORA_MODELS with fallback to the base text model.
  • Validation: Each caption is scored across 5 weighted dimensions (accuracy 35%, style_fit 25%, specificity 20%, hallucination_control 15%, completeness 5%) by an LLM judge; captions below threshold trigger revision with specific feedback while Phase B retry time remains. The judge prompt distinguishes unsupported literal facts from grounded style metaphors. If the retry budget is exhausted after a validation failure, the last usable generated commentary is kept; if no usable commentary exists, the system falls back to a grounded local caption.
  • Web UI Quick Chat: The local web UI samples Modal-generated intro/waiting lines per dragon. Each selected dragon speaks one intro at job start, then the chat switches to non-repeating markdown canned responses plus Modal waiting lines with speaker cooldowns to keep one dragon from dominating.
  • Timeout and retry budget: The Docker batch path counts /input/tasks.json video tasks up front. Phase B reserves 30 seconds for shutdown/output, then splits the remaining time-to-target across video tasks. Dragons for a single video share that task-level budget because they run in parallel.

Local Development (no Docker)

pip install -r requirements.txt

# Run the CLI agent
python -m app.main

# Or run the web UI
uvicorn app.web:app --reload --host 127.0.0.1 --port 5173

Create a .env file in the project root with your FIREWORKS_API_KEY to power the pipeline outside Docker.

LoRA Fine-Tuning Pipeline

Each dragon style can route to a dedicated Qwen3-1.7B LoRA adapter trained through the Fireworks SFT API. These adapters are optional; the batch path can fall back to the base text model whenever USE_LORA=false or a style-specific adapter is unavailable.

Style LoRA Model ID
formal accounts/oak42000-.../models/ft-z5npxv9s-l12gp
sarcastic accounts/oak42000-.../models/ft-hrbnfpjn-sts6p
humorous_tech accounts/oak42000-.../models/ft-nt5fc352-ol5kf
humorous_non_tech accounts/oak42000-.../models/ft-du-rong-base-2k-v4-run001

Dataset generation pipeline (lora-ft/):

  1. generate_prompts.py — 6,200 prompt variants across 6 categories
  2. generate_data.py — kimi-k2p6 generates in-character responses (11,200 total)
  3. validate_data.py — LLM judge rejects low-quality samples (< 7/10)
  4. prepare_data.py — Formats into OpenAI chat JSONL
  5. finetune.py — Submits to Fireworks SFT (rank 16, 3 epochs, lr 1e-4)

Set LORA_MODELS in .env to route captions through fine-tuned adapters.

The same Modal-generated training data also feeds the web UI quick-chat pool. scripts/extract_modal_quick_chat.py extracts clean introduction and waiting examples into assets/dragons/quick_chat/modal_quick_chat.json; /api/dragons/quick-chat?limit=200 returns categorized samples per dragon on each page load. The browser shows one intro per selected dragon, then removes waiting/canned messages from its page-local pool as they are used. If a dragon speaks twice in a row, it cools down until the other selected dragons have spoken.

Testing

pytest tests/ -v

Styles & Dragons

Style Dragon Tone
formal Trí Long Professional, objective, factual data analyst
sarcastic Lão Quân Dry, ironic, lightly mocking ancient dragoness
humorous_tech Gemma Trí Tech Funny, tech/programming references, dragoness engineer
humorous_non_tech Du Rong Funny, everyday humour, young dragon with fake ancient wisdom

Presentation & PDF Export

The project includes a Reveal.js slide deck at htdocs/presentation.html covering the architecture, dragon personas, dataset pipeline, and key features.

Project Structure

app/              # Core pipeline components
├── main.py       # CLI entrypoint
├── pipeline.py   # Orchestrator
├── downloader.py # Video downloader
├── extractor.py  # FFmpeg frame/audio extraction
├── scenes.py     # Scene detection
├── sampler.py    # Adaptive frame sampling
├── vision.py     # Phase A — VLM analysis
├── whisper.py    # Audio transcription
├── merger.py     # Merge vision + audio
├── caption.py    # Phase B — style-specific caption gen
├── validator.py  # Caption quality validation
├── prompts.py    # LLM prompt templates
├── styles.py     # Dragon persona loader
├── config.py     # Configuration
├── cache.py      # Disk caching
├── utils.py      # JSON repair, helpers
├── web.py        # FastAPI web server
└── resources/    # Style reference examples
assets/dragons/   # Dragon profiles, sprites, animations
└── quick_chat/    # Modal-generated intro/waiting chat corpus
htdocs/           # Web UI frontend
tests/            # Test suite
Dockerfile        # Competition build

About

Dragon Den entry

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages