AI agents write the code. kstrl makes sure it actually works.
kstrl (pronounced "kestrel") is a software factory for AI coding agents. Like the bird, it hunts by hovering: it holds position over your codebase, watches everything the agent does, and strikes precisely when something is wrong. You hand it a spec and walk away. It gives the agent the context it needs before it starts, measures what the agent produced with checks the agent did not write, feeds the gap back as the next instruction, and stops only when independent checks agree the work is done.
The problem it solves: a coding agent is powerful and it is also the least reliable witness to its own work. Left alone, it works on one prompt at a time, decides for itself when it is finished, and reports success in the same voice whether the code works or not. kstrl is built on one rule: nothing the agent says about its own output is the final word. Every claim is measured by something else, the measurement decides what happens next, and the whole thing runs without you until it reaches a decision only you can make. Because walk-away automation is only trustworthy when you can see what it did, every run streams a typed event log you can watch live in a terminal dashboard, attach to from another terminal, or replay after the fact.
Most agent wrappers are retry loops: run the agent, ask it whether it is done, run it again if not. The agent is both the worker and the judge, so the loop closes on its own opinion of itself.
kstrl is designed as a loop that closes on evidence instead. The parts are deliberate, and each exists to answer one question:
- What should the agent know before it starts? Computed context: a module map of the tree, built without a model call, beside the project's own CLAUDE.md, golden patterns and memory. The agent reads the code itself with its own tools.
- How do we know what it actually did? Independent measurement: mechanical checks, then a reviewer from a different model family judging every acceptance criterion, then a security reviewer. None of these read the agent's description of its work; they read the diff.
- What happens when the measurement disagrees? The gap becomes the next instruction: each failing check's exit status and the lines of its own output around every file location it names in the tree. The agent tries again against a smaller, sharper target.
- What stops it running away? Bounds on everything: iterations, time, tokens, cost, work in flight, and a written envelope of what a merge may touch. When a bound trips, the run stops loudly and tells you why.
- Where do you stand? On the loop, not in it. Boundary conditions route to you: a spec the architect cannot decompose, a merge you asked to approve, a budget that ran out. Everything else flows, and everything is recorded.
kstrl is those loops, nested, innermost fastest. The innermost is the only one with no check of its own; every loop around it acts on what the one inside produced and measures it with something the agent did not write. The loops, their clock rates and what measures each are the table at the top of ARCHITECTURE.md.
The reviewers are the point: independent, adversarial, and expected to distrust the implementing agent. The harness distrusts the reviewers in turn: an empty, partial or oversized review fails closed, and a reviewer's own claim to have searched thoroughly is shown as a hint and never used as a gate. The only thing that proves a reviewer works is calibration: planting known bugs and measuring how often each role catches them. The full phase-by-phase pipeline lives in ARCHITECTURE.md, and the reasoning behind the loop is in docs/loop-design.md.
Documentation: ARCHITECTURE.md is the detailed system tour (pipeline, iteration loop, factory scheduling, state layout), docs/adversarial-design.md covers the full 8-role taxonomy, docs/env-vars.md every environment variable, docs/runbook.md operator failure recovery, docs/baseline.md the check baseline and the pull-request regression report, and docs/linear-integration.md the optional Linear mirror. examples/ has two sample feature specs.
Install from PyPI (or from a clone for development):
uv tool install kstrl # installs `ks` and `kstrl` (requires Python 3.11+, uv)
# or: git clone https://github.com/0xfauzi/kstrl.git && cd kstrl && uv tool install -e .
cd your-project
ks init . # scaffold kstrl.toml, prompt/PRD templates and a .gitignore
$EDITOR scripts/kstrl/prd.json # define what to build (user stories + acceptance criteria)
ks run 25 # let the agent work for up to 25 iterationsThat is the single-component loop, which creates no PR. If you already have a
spec, ks decompose --spec <spec.md> --project-name <name> plans it into
components and ks factory --spec ... plans and builds it; ks init prints
both paths when it finishes.
ks check runs the mechanical checks (tests, typecheck, lint, diff scope, bad patterns) against any tree by hand, with no PRD, branch, worktree or agent spend; ks check --json prints the same measurement as one JSON document for scripts.
It is the standalone entry point to the checks the factory runs in Phase 1, so a threshold can be measured before it is automated.
The measurement is read-only: it runs against your live checkout, so it never edits, stages, commits or leaves bytecode behind, and it exits 2 rather than reporting a pass when git cannot produce the diff against --base.
ks check --write-baseline records those measurements to scripts/kstrl/baseline.json and ks check --compare-baseline reports what a branch added to it - new failure signatures, signatures whose count rose, signatures the branch fixed, and any check that measured on the baseline and has stopped measuring here. That is the baseline: it stops concurrent work undoing the loop's progress while the loop runs. It is advisory (exit 0 either way) until you pass --fail-on-regression. How to adopt it and how to graduate to blocking: docs/baseline.md.
Every long-running command opens its live dashboard on a terminal automatically (--no-tui opts out), and bare ks opens a home shell with a run browser and command launcher; ks dash attaches a read-only view to any run - in flight or finished - from another terminal, and ks status prints the same state for scripts and CI (and opens the dashboard when you run it interactively).
You need at least one AI coding agent CLI:
| Agent | Install | Example models |
|---|---|---|
| Claude Code (recommended) | claude.ai/code | opus, sonnet, haiku, claude-fable-5 |
| OpenAI Codex | github.com/openai/codex | gpt-5.5, gpt-5.4 |
| Custom | Any command that reads stdin | - |
Model names current as of 2026-07: the claude aliases track the newest release in each tier (today Opus 4.8, Sonnet 5, Haiku 4.5) with claude-fable-5 as the top-end frontier model, and codex defaults to gpt-5.5 (the older gpt-5.x-codex ids are being retired).
kstrl does not validate model names: [agent].model is passed straight through to the CLI (claude --model / codex -m), so any model the installed CLI accepts works.
There is also an opt-in in-process adapter, [agent] type = "claude-sdk", that drives Claude through the Claude Agent SDK instead of a CLI subprocess and supports an in-loop USD budget ceiling ([agent].budget_usd). It requires the sdk extra (uv sync --extra sdk) and is never chosen by auto-detect. [agent].budget_usd is adapter-internal and bounds a single turn; it is not [factory].max_cost_usd, which is the run-level spend ceiling across every phase and component (details).
No default stack: kstrl assumes no language, toolchain or test runner. It runs the checks of the [stack] in kstrl.toml once a person confirms it, and ks doctor files a proposed [stack] from an old [verify]. ks init writes the same files in every repository.
kstrl builds a module map of the tree (every directory git lists, with its file and line counts) and injects it into the prompt. No LLM calls, no token cost. kstrl reads no source language; the agent reads the code itself.
When the agent signals completion, kstrl treats that as a claim and measures the work with checks the agent did not write:
- Mechanical checks (fast, computational): tests, typecheck, lint, no changes outside allowed paths, no leaked secrets.
- Adversarial review (LLM): an independent reviewer checks the diff against the acceptance criteria, then a security reviewer hunts vulnerabilities - in
hardmode their failures block. - Contract testing (multi-component runs): component branches merge tier-by-tier with integration tests at each tier.
A check passes when its command exits 0, and kstrl parses none of its output. When a check fails, the retry prompt carries its exit status and the lines of its output around every file location inside the worktree, or its last five lines when it names none. The model reads the tool's own message.
After each component, a distiller writes durable facts about what was built, and later components receive those facts in their context. That loop is closed and its uptake is measured (as a lower bound) on every run.
After each factory run, kstrl also records structured failure signatures, costs and finding categories to an evolution journal, and ks evolve reports the patterns that recur, sorted into inbox traffic, candidate lessons and neither.
ks evolve # analyze recent runs, find patterns
ks evolve --status # show experiment trends (retry rate over time)Be clear about what that is today: a record, not a closed loop. Nothing writes a candidate lesson anywhere, and nothing reads one back into the next run. The design that closes it, with attribution (a lesson that keeps failing to prevent the failure it targets retires itself) and a shared playbook across every project you run, is docs/continuous-learning-design.md, tracked as R9. Until it lands, this README does not claim the harness improves itself.
For large features, kstrl decomposes a spec into independent components and runs them in parallel:
kstrl decompose --spec features.md --project-name myproject
kstrl factory --manifest scripts/kstrl/manifest.json --max-parallel 4Each component runs in an isolated git worktree (.kstrl/worktrees/<run>/<component>) with its own PRD. ks run is actually factory mode with a single component - the same verification pipeline runs whether you're building one feature or twenty.
Scheduling, worktree isolation, merge gating, and the contract-testing bisect are described in ARCHITECTURE.md.
Enable the mirror and every factory run shows up in Linear without you lifting a finger:
[linear]
enabled = true
team_id = "your-team-uuid"export KSTRL_LINEAR_TOKEN="lin_api_..."ks decompose creates a project and one issue per component (user stories as a checklist in the issue body); non-blocker spec findings are filed into Triage. Status transitions ride Linear's own GitHub integration: the component branch carries the issue identifier and the PR body carries a Fixes ENG-42 trailer, so In Progress on PR-open and Done on merge cost kstrl zero API calls. Failures and budget halts land as a comment on the issue, and retries update the same issues - never duplicates, because issue ids persist in the manifest.
The mirror is observability-only by design: every Linear failure warns and degrades, and nothing in the pipeline ever fails because Linear did. Setup, per-team automation settings, and rate-limit behavior: docs/linear-integration.md.
The long-running kstrl surface is available through a terminal UI. Bare ks opens the home shell: project identity, a browser over every recorded run of every kind (with folded outcome/token/cost summaries and honest · cells while they compute), and a command launcher - factory and decompose launch from forms right there, retry picks a failed component off the manifest, and config, evolve, and the init wizard open as full screens. Non-TTY invocations are byte-identical to before (ks prints help, exit 2); KSTRL_NO_TUI=1 opts out everywhere.
Every long-running command (ks factory, ks run, ks retry, ks understand, ks feature, ks decompose) records a replayable event stream and opens its embedded dashboard on a terminal; plain output remains the default for non-TTY use. The screenshots are real, captured from a live toy-project factory run.
The overview board shows every component's status, authoritative phase, attempt, last-event age, and spend - here with auth-core and api-routes running in parallel workers while ui-shell waits on its dependency:
enter drills into a component: phase timeline with verdicts, the typed findings stream with reviewing-model attribution, the live engineer transcript, and evidence paths.
With pause_before_pr_merge enabled, the E6 human checkpoint opens as a real inspection surface - the diff excerpt, both finding streams, and what the attempt cost - instead of a y/n prompt. a approves, r rejects (fails the component and skips dependents), t consumes a retry, escape leaves it pending while you look around the dashboard:
Keys: enter detail, escape back, f follow transcript, c reopen a checkpoint, q quit (graceful stop when embedded; a second q escalates).
The TUI is a view, never the record: every run appends typed, schema-versioned events to .kstrl/runs/<run_id>/events.jsonl, which is why attaching mid-run, replaying a finished run, and surviving a dashboard crash all work by construction (details). Cost figures are CLI self-reports: any unreported call renders a + marker and the total becomes a lower bound - the dashboard never turns an honest number into a false one.
Agent-generated tests can be written to pass trivially. An acceptance plan is a directory outside the repository that holds plan.json and the files its checks use. ks factory --acceptance <dir> copies it under the control directory and runs every check on the base before the first engineer, on each component's head, and in Phase 3 on the merged tree. A check is an argv judged by its exit status alone. It runs in its own copy of the plan, with KSTRL_TREE naming the checkout under test, so no file the engineer writes is on its path. A held-out check is never shown to an engineer.
The PRD fixtures key and the [fixtures] section are retired (#700 slice 8). A PRD or kstrl.toml that still carries one is refused, with a line that says to write each fixture as an acceptance check.
kstrl is built for one person running it unattended, and it is designed around the difference between being in the loop and being on it. In the loop, you are a required step and nothing proceeds without you. On the loop, the factory runs itself and comes to you only at boundaries it cannot judge, while you adjust how it behaves between runs.
Today the boundaries that reach you are: a spec the architect halts on (there is no override flag; the spec gets fixed), the optional checkpoint before a merge, a budget or breaker that stopped a run, and an inbox of decisions the daemon could not take alone. How much the factory may do without asking is one ordered level, earned by evidence plus your recorded acknowledgement and revoked automatically. The levers you turn between runs are the project prompt, CLAUDE.md, kstrl.toml, and the acceptance criteria themselves; the R10 work adds a way for your pull-request comments to reach the next run. One more lever is scripts/kstrl/golden-patterns.md, which ks init scaffolds and you fill in: what a good change looks like in this repository, with a file to copy from for each pattern, carried into every engineer prompt of ks factory, ks run, ks retry, ks serve and ks feature (ks understand does not read it). It is yours to write and is read verbatim like CLAUDE.md, never generated for you. Until you edit it nothing is injected, because kstrl recognises its own scaffold and treats it as empty; past about 1500 tokens it keeps the START of the file and drops the end, cut at a line boundary where that still delivers most of the budget and at the budget boundary otherwise, with how much arrived and which end went announced in the prompt and once on your terminal.
scripts/kstrl/memory.md is the standing-feedback lever, and ks init
scaffolds it beside the golden patterns. Where golden patterns say what good
code looks like here and change rarely, memory carries your standing
corrections and grows as you review: permanent scope exclusions, areas whose
findings are known false positives, and the feedback you have given more than
once. What does not belong in it is a one-off instruction or a run log; an
entry changes future runs rather than one pull request, which is the whole
point of the file.
The position is the mechanism. It is read into the engineer prompt AFTER the
retry context and before CLAUDE.md, so your correction frames how this
attempt's failures get acted on rather than the other way round. It reaches
every engineer prompt of ks factory, ks run, ks retry, ks serve and
ks feature (ks understand does not read it), from the repo root and never from a
component's worktree. Until you edit it nothing is injected, because kstrl
recognises its own scaffold; past about 1000 tokens it keeps the END of the
file and drops the start, which is the opposite of golden patterns and is
because this is the file that grows at the end. Your newest corrections
survive, pruning from the top is what preserves them, and how much arrived and
which end went are announced in the prompt and once on your terminal.
/memory <text> and /iterate <text> comments on an open kstrl pull request
are appended under ## Guidance by ks serve when
[intake_github] steer_enabled is on; kstrl finds that heading rather than
appending to the end of the file, and adds the heading if the file has none.
The daemon writes the file and does not commit it.
You can, and for small tasks you should. kstrl is for when you want to:
- Define success criteria before starting - acceptance criteria, path restrictions - not just "make it work"
- Walk away - kstrl runs unattended with structured verification, not just a completion marker
- Watch it without babysitting it - a live dashboard over a replayable event log; attach, detach, or inspect after the fact
- Give the agent context - codebase scan injection means fewer wasted iterations discovering the codebase
- Get focused retries - each failing check's own output around the locations it names
- Build multiple components in parallel - factory mode with worktree isolation and contract testing
- Carry knowledge forward - facts distilled from each component reach the next one, and the journal records every failure signature for the learning work that follows
- Red-team the spec before building - the architect pass halts on blocker-severity spec ambiguities instead of guessing
ks autonomy demote Drop the autonomy level by one and start the cool-down.
ks autonomy history Show every recorded level transition.
ks autonomy promote Raise the autonomy level by one.
ks autonomy replay Replay the ladder's thresholds over recorded run history.
ks autonomy status Show the current level, its flag bundle, and what promotion needs.
ks check Run the mechanical checks against a tree and print the measurement.
ks ci poll Read the CI state of every recorded merge commit and record it.
ks config show Print the fully resolved config with the source of each value.
ks dash Live dashboard over a factory run (observe-only).
ks decompose Decompose a spec into components and generate PRDs.
ks doctor Assess whether this repository is ready to point kstrl at.
ks evolve Analyze factory runs and report recurring failure patterns.
ks factory Run the software factory - decompose and execute a spec.
ks feature Run feature understanding, then implementation.
ks health Trend the factory's own run metrics against its own history (R8.4).
ks inbox approve ITEM_ID Accept the exception and close the item.
ks inbox ls List items awaiting a decision.
ks inbox reject ITEM_ID Refuse the exception, recording why.
ks inbox retry ITEM_ID Requeue the item's component and close the item.
ks inbox show ITEM_ID Show one item in full, including its evidence.
ks inbox snooze ITEM_ID Defer an item; it returns when the TTL lapses.
ks init [DIRECTORY] Initialize kstrl in a project directory.
ks learn playbook Print the folded global playbook and its ledger's line count and SHA-256.
ks learn repair List every global playbook line the fold refuses; --yes voids each one in the ledger.
ks queue add SPEC Enqueue a spec file.
ks queue answer ITEM_ID SPEC Answer an escalated item: replace its spec and send it back to queued.
ks queue ls List queue items in run order.
ks queue pause Stop admitting queued work.
ks queue priority ITEM_ID Change a queued item's priority, keeping its id and history.
ks queue resume Start admitting queued work again.
ks queue retry ITEM_ID Send a failed or poisoned item back to queued.
ks queue rm ITEM_ID Delete an item and its spec.
ks queue show ITEM_ID Show one item in full, with its transition history.
ks queue sync Pull labelled GitHub issues into the queue (R8.6).
ks recheck RECORD Run an acceptance record's saved checks again and compare (#700).
ks retry COMPONENT_ID Retry a FAILED component from the factory manifest (R3.3).
ks run [MAX_ITERATIONS] Run the agentic loop as a single-component factory invocation.
ks serve Drain the continuous-intake queue (R8.6).
ks signals ls Print the signal ledger, oldest first.
ks signals poll Fetch one tracker page, classify it against the ledger, print the tally.
ks status Show per-component status from the manifest + progress log.
ks understand [MAX_ITERATIONS] Run codebase understanding loop (read-only mode).
Run ks COMMAND --help for the full option list of any command.
kstrl reads kstrl.toml at the project root; copy kstrl.toml.example to start. Precedence: CLI flags > environment variables > kstrl.toml > dataclass defaults. ks config show prints the fully resolved config with the source of each value.
# Agent selection
[agent]
type = "" # "claude-code" | "claude-sdk" | "codex"; empty/"auto" = auto-detect
command = "" # custom agent shell command; overrides type
model = "" # e.g. "sonnet" (claude) or "gpt-5.5" (codex); empty = agent default
reasoning_effort = "" # low | medium | high | max (model-dependent)
budget_usd = "" # in-loop USD ceiling; claude-sdk adapter only; empty/0 = unlimited (R7.6)
# Loop behavior
[run]
max_iterations = 10 # iteration budget per component
sleep_seconds = 2.0 # pause between iterations
interactive = false # human-in-the-loop mode for the legacy loop
# File locations
[paths]
prompt = "scripts/kstrl/prompt.md" # engineer prompt file
prd = "scripts/kstrl/prd.json" # PRD file
progress = "" # progress log the agent appends to; empty = each factory component writes beside its own PRD (inside its allowedPaths), set = that one path is forced on every component
codebase_map = "scripts/kstrl/codebase_map.md" # brownfield codebase notes
golden_patterns = "scripts/kstrl/golden-patterns.md" # operator-authored golden patterns, injected into every factory engineer prompt and every `ks run` prompt (R10.8)
memory = "scripts/kstrl/memory.md" # operator-authored standing feedback, injected into every factory engineer prompt and every `ks run` prompt after the retry context (R10.9)
allowed = [] # diff-scope allowlist, e.g. ["src/", "tests/"]; empty = unrestricted
# Branch handling
[git]
branch = "" # branch override; empty = use PRD branchName
auto_checkout = true # check the branch out automatically
# Output rendering
[ui]
ascii = false # ASCII separators only (no box-drawing characters)
# Timeout limits (seconds; 0 disables)
[timeout]
agent_iteration = 0.0 # one engineer iteration; 0 = no limit
component_total = 0.0 # wall clock per component across iterations; 0 = no limit
scheduler_backstop_margin = 60.0 # extra slack before the scheduler declares a worker dead
# Factory orchestration (Phase 0-3 pipeline)
[factory]
max_parallel = 4 # concurrent component workers
max_retries = 3 # per-component retry budget across all phases
retry_delay = 5.0 # seconds between retry attempts
use_worktrees = true # isolate each component in .kstrl/worktrees/<run>/<id>
single_pr = false # one PR for the whole run instead of per-component
create_prs = true # push + merge PRs via gh
review_mode = "hard" # hard | advisory | skip (Phase 2)
review_timeout_seconds = 0.0 # code and integration reviewer call timeout; 0 = no limit. A call killed at it is an infrastructure error, never a verdict
architect_timeout_seconds = 0.0 # architect call timeout (ks decompose, ks factory --spec); 0 = no limit. A call killed at it is one failed attempt
claim_agreement = "advisory" # advisory | block: what to do when the reviewer does not confirm a story the engineer marked passes=true (R10.3)
merge_timeout = 300.0 # hang guard: seconds to wait for PR merge confirmation
max_adversarial_calls = 0 # cap on review+security+distill LLM calls; 0 = no limit. At the cap a hard-mode review or security phase HALTS the component rather than merging it unreviewed; an advisory one skips. Budget 3 calls per component for hard review + hard security + knowledge (R10.5, docs/runbook.md)
max_total_tokens = 0 # run-level token budget; 0 = no limit. Counts cache reads at par, so it is a poor proxy for cost - prefer max_cost_usd. Halts before the next engineer iteration or phase, never mid-call (docs/env-vars.md)
max_cost_usd = 0.0 # run-level USD budget; 0 = no limit. Same halt granularity as max_total_tokens (between iterations, not mid-call), so NOT a hard cap. Not [agent] budget_usd (docs/env-vars.md)
pause_before_pr_merge = false # human checkpoint before each PR (E6)
progress_log_enabled = true # JSONL event log at .kstrl/progress.jsonl (R3.2); usage accounting is written either way (docs/env-vars.md)
keep_worktrees_on_failure = false # keep failed components' worktrees for post-mortem (R3.3)
integration_review = true # record-only review of the merged feature after Phase 3 (#482); never gates
integration_blocking = false # a non-clean integration stop fails the run; every open finding is handed off (#696)
convergence_attempts = 0 # fail a component whose gate failure count has not fallen for this many consecutive attempts; 0 = off (#233)
worktree_setup_timeout = 0.0 # seconds before the worktree setup's process group is killed; 0 = no limit (#624)
# No-progress circuit breaker (R7.5; 0 iterations disables)
[breaker]
no_progress_iterations = 3 # halt after N consecutive no-progress iterations; 0 disables (R7.5)
# OS-level agent sandboxing (R7.5; claude-code, claude-sdk and codex only)
[sandbox]
enabled = false # OS-sandbox the engineer, reviewer, understand and repair agent CLIs (writes scoped to the worktree); a custom agent command in one of those roles is refused (exit 2)
allow_network = false # re-open outbound network inside the sandbox (off = deny)
# Phase 1 mechanical verification
[verify]
check_diff_scope = true # fail on changes outside allowed paths
check_bad_patterns = true # scan the diff for secret-like patterns
subprocess_timeout = 0.0 # seconds per verification subprocess; 0 = no limit
require_self_critique = false # fail Phase 1 if the ## Self-Critique block is missing/sparse
self_critique_min_bullets = 3 # minimum substantive bullets in the block
progress_file_path = "" # progress file the self-critique check reads; empty = the log beside the component's PRD
fast_iteration_checks = [] # gates run after each unfinished engineer iteration, failures shown in the next prompt: any of "test_suite", "typecheck", "linter"; empty = off (#233)
# Phase 1 policy envelope (R8.1; opt-in)
[policy]
enabled = false # enforce the [policy] envelope in Phase 1 (opt-in)
paths_deny = [".github/workflows/**", "kstrl.toml", "ralph.toml", ".kstrl/**", "**/*.pem", "**/.env*"] # globs no change may touch (gitignore-style **)
max_files_changed = 40 # max files in a change; negative disables
max_lines_changed = 1500 # max added+removed lines; negative disables
secret_patterns = ["AKIA[0-9A-Z]{16}", "-----BEGIN (?:RSA |EC )?PRIVATE KEY-----", "sk-[a-zA-Z0-9]{20,}", "ghp_[a-zA-Z0-9]{36}", "xox[bpoas]-[a-zA-Z0-9-]+"] # regexes flagged in added diff lines
enforcement_paths_extra = [] # extra halt paths; additive - cannot shrink the built-in set
deploy = false # reserved for the R8.7 release gate; stored + hashed
# Autonomy ladder (R8.2; opt-in)
[autonomy]
enabled = false # derive run permissions from the ladder level (opt-in)
max_level = 4 # hard ceiling: never run above this level (1-4)
demote_on_calibration_regression = false # demote one level when `calibration compare` finds a regression (an inbox item is opened either way)
demote_on_health_breach = false # demote one level on an R8.4 health control-limit breach (an inbox item is opened either way); run `ks health` first - the measured false-alarm rate at the n floor is in roadmap R8.4
# Exception inbox (R8.3)
[inbox]
enabled = true # record exceptions awaiting a human decision
open_item_cap = 50 # open items after which queue intake pauses; 0 = unbounded
snooze_hours = 24.0 # default snooze TTL in hours; snoozed items return
notify_action_required = true # notify on action-required items and demotions only
# GitHub Issues remote inbox (R8.6)
[intake_github]
enabled = false # poll GitHub Issues for labelled work (opt-in outbound poller)
repo = "" # owner/name to poll; empty resolves from the checkout's remote
queued_label = "kstrl:queued" # the trigger label; allowed_actors decides who may apply it
label_prefix = "kstrl:" # prefix for the state labels written back to the issue
max_items_per_sync = 5 # upper bound on items admitted per sync, and on steering comments acted on per cycle
default_priority = 0 # queue priority given to remote-sourced items
comment_on_result = true # post the queue's verdict back to the source issue
dry_run = false # poll and log, but send no labels or comments; also suppresses steering's memory-file write
timeout_seconds = 60.0 # per-gh-invocation timeout in seconds
allowed_actors = [] # logins allowed to apply the trigger label; empty = anyone who can label
steer_enabled = false # act on /memory and /iterate comments on open kstrl PRs (writes the memory file)
# Continuous-intake daemon (R8.6)
[serve]
poll_interval_seconds = 60.0 # seconds between poll cycles when ks serve runs as a daemon
daily_budget_usd = 0.0 # hard stop on reported spend per local day; 0 = no cap
max_consecutive_poison = 3 # consecutive poisoned items that pause the whole queue
caffeinate = true # hold caffeinate -i for each run so the machine cannot sleep mid-factory
factory_timeout_seconds = 0.0 # kill a factory run after this long; 0 = no timeout
allow_uncovered_cost = false # run unattended even when no adapter reports cost, making the budget unenforceable
max_open_prs = 1 # scheduled admission stops while this many kstrl-authored PRs are open; 0 = no bound
# Continuous intake queue (R8.6)
[queue]
max_attempts = 3 # execution attempts per queue item before it is poisoned
lease_ttl_seconds = 3600.0 # claim validity in seconds; the reaper recovers anything past this
# Across-attempt review divergence detector (#265)
[divergence]
mode = "advisory" # skip | advisory | block; advisory records a diverging retry loop, block fails the component instead of paying for another retry (#265)
growth_steps = 2 # consecutive steps required before the detector fires, so it needs growth_steps + 1 measured attempts; must be >= 1 (use mode = "skip" to disable); the default is a structural minimum, not a measured number
# Phase 2.5 security review
[security]
mode = "advisory" # skip | advisory | hard
agent_cmd = "" # empty = inherit [agent]
agent_type = "" # empty = inherit [agent]
model = "" # empty = inherit [agent]
timeout_seconds = 0.0 # reviewer call timeout; 0 = no limit
fail_threshold = "high" # critical | high | medium | low (hard mode)
# Phase 3 cross-component contract testing
[contract]
mode = "tier" # tier | final | skip
timeout = 0.0 # seconds per contract test run; 0 = no limit
# Phase 4 release (R8.7 slice 1: records the release ref; deploys nothing)
[release]
enabled = false # record a release ref; still deploys nothing (R8.7 slice 1)
environment = "" # deploy environment name, e.g. staging or prod
# Phase 0 codebase scan (computational, no LLM)
[codebase_scan]
enabled = true # inject structural context into the prompt
module_map = true # directory tree with file and LOC counts
max_context_tokens = 4000 # cap to avoid prompt bloat
# Per-component knowledge layer
[knowledge]
enabled = true # distill + inject durable facts
max_core_tokens = 2000 # current component's facts (full text)
max_dependency_tokens = 1000 # dependency facts (full text)
max_sibling_tokens = 500 # other components' facts (first sentence)
distill_timeout_seconds = 0.0 # distiller call timeout; 0 = no limit
distill_model = "" # empty = falls back to [agent].model
max_facts_per_distill = 7 # cap on facts written per component
dependency_scope = "direct" # direct | transitive (E8)
# Continuous-learning journal
[evolution]
enabled = true # record run outcomes
journal_path = ".kstrl/evolution.jsonl" # JSONL journal location
experiments_path = ".kstrl/experiments.tsv" # experiment tracker location
min_pattern_frequency = 2 # pattern must recur in N runs to be reported
lookback_runs = 10 # past runs to analyze
# Run-milestone notification hooks (R3.2)
[notify]
on_complete = "" # shell hook fired once when the run finishes; empty = disabled
on_first_failure = "" # shell hook fired once on the first component failure
on_inbox_item = "" # shell hook fired per R8.3 inbox item kind; empty = disabled
hook_timeout = 30.0 # seconds before a hook command is killed
# Linear integration (R7.4; default off)
[linear]
enabled = false # mirror runs into Linear (project/issues/status via GitHub linking)
team_id = "" # Linear team UUID (required when enabled)
token_env = "KSTRL_LINEAR_TOKEN" # NAME of the env var holding the API token
auth_mode = "auto" # auto | api_key | oauth (auto sniffs the lin_api_ prefix)
api_url = "https://api.linear.app/graphql" # GraphQL endpoint
dry_run = false # record mutations instead of sending them
timeout_seconds = 30.0 # per-request timeout
min_request_interval = 0.5 # client-side throttle between requests (seconds)
# Runtime signal poller (R8.8 slice 1)
[signals]
enabled = false # poll the configured tracker
product = "" # the product name recorded on every ledger row
base_url = "http://127.0.0.1:8000" # the tracker's root URL, no trailing slash needed
project_id = "" # Bugsink's numeric project id, as a string
token_env = "KSTRL_SIGNALS_TOKEN" # NAME of the env var holding the tracker's bearer token
http_timeout = 10.0 # per-request timeout
new_issue_events = 3 # advisory threshold; labels a new issue, gates nothing
repeat_growth_events = 10 # advisory threshold; labels a repeat, gates nothing
# Cross-project learning: the global playbook (#217)
[learning]
contribute = true # append this project's lessons to the global playbook
consume = true # read global playbook lessons into this project's promptsEnvironment variables override kstrl.toml, and CLI flags override both. See docs/env-vars.md for the full env-var mapping.
The PRD (prd.json) is a list of user stories with testable acceptance criteria:
{
"branchName": "kstrl/login-feature",
"allowedPaths": ["src/", "tests/"],
"userStories": [
{
"id": "US-001",
"title": "User can log in with email",
"acceptanceCriteria": [
"Login form accepts email and password",
"Invalid credentials show error message",
"Tests pass: uv run pytest tests/test_auth.py"
],
"priority": 1,
"passes": false,
"notes": ""
}
]
}The agent updates passes and notes as it works, and kstrl reads them between iterations to decide whether to continue. Treat passes as the agent's claim, not the verdict: mechanical verification checks the flag is set, and the reviewer independently judges every criterion. Acceptance criteria should be concrete and testable - commands the agent can run, behavior it can verify - because they are the claim every check measures against.
allowedPaths is optional for a hand-written PRD (it feeds the Phase 1 diff-scope check); the architect is required to emit it for every decomposed component.
git clone https://github.com/0xfauzi/kstrl.git
cd kstrl
uv sync
uv tool install -e .
uv run pytest tests/ # 3200 passed, 28 skipped at the time of writing (2026-08)
uv run mypy kstrl/ --strict
uv run ruff check kstrl/ tests/The CLI reference and config reference sections of this README are generated: edit the source (click commands / config dataclasses) or scripts/gen_docs.py, then run uv run python scripts/gen_docs.py. CI fails if the generated sections are stale.
Contributions are welcome, including AI-assisted ones - but AI-generated code is never gated by AI self-review, so every change is reviewed by a human and PRs should declare whether an agent wrote them. Start with CONTRIBUTING.md for the setup, the process rules, and how to pick up roadmap work; the project wiki covers the vision, architecture, and roadmap in depth. To report a vulnerability, see SECURITY.md. Release history is in CHANGELOG.md.
MIT