Skip to content
ryangu00Public

About

Prove every "done." — verifies AI agent completion claims against evidence declared before the work. Observe-mode by default, zero-config, four runtimes.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

1 watching

Forks

Axiom

Axiom — Prove every "done."

CI License Python Plugin Ask DeepWiki

Trust nothing unaudited — including yourself.

Declare the evidence before the work begins. When your agent says "done," Axiom checks the claim against it — not against the conversation.

Axiom is the working discipline for long-running coding agents: one act of verification at each station of the loop, on the official Claude Code hook API — plus native adapters for Codex CLI, hermes-agent, and OpenClaw. The part that ships as deterministic enforcement today is the Execute station; every other station is labeled below with exactly the form it ships in, and nothing is labeled as installed behavior until it is.

Axiom itself makes no network calls, needs no API key, sends no telemetry, and puts no model in the verification path — Python stdlib only, all local. (The one thing that reaches outward is a cmd_succeeds predicate: it runs the command you declared, with your permissions — fresh execution, not a sandbox.) What it deliberately does not catch is enumerated in What this won't catch.

Update (2026-10)

As of 2026-10, the hook refinements below retain observe mode by default. They derive from predecessor incident, replay, review, and design history; they are not measured improvements on the public hooks.

  • Failure accounting: user-interrupted failures leave clusters untouched. Intentional polling defaults to until loops, while loops containing sleep, tail -f, and watch. The default threshold remains 3 similar failures in a cluster with a 30-minute inactivity window. rules.stuck-search.polling_patterns replaces the default regular expressions. A leading sleep N && exemption is opt-in and off by default: enabling it can hide a real failure after the sleep. No labeled sample has resolved whether that broader exemption is appropriate.
  • Advice and reporting: enforce mode writes one advice_injected event per injection. rules.stuck-search.cooldown_minutes defaults to 10 minutes per cluster, suppressing repeated injections but not observe-mode would_have_blocked records at or above the threshold. The report displays advice counts separately; only observe findings populate recent incidents and calibration notices. Interrupt handling and cooldown are design-level changes, with no measured public benefit.
  • Search tracking: the first successful configured search or fetch within 30 minutes of the latest advice in the same session records one search_after_trigger event with lag_seconds. The pending measurement survives a successful command clearing the cluster; a new injection in the session replaces it. rules.stuck-search.search_tools contains full-match, case-insensitive regular expressions, defaulting to WebSearch, WebFetch, and mcp__.*(?:search|fetch).*. A search is not evidence of relevance, causation, or compliance by the particular agent that received the advice.
  • Temporary shell targets: redirections, tee, and the first non-option sqlite3 argument use the shared temporary-root resolver and configured persistent-name patterns. Observe mode records findings; enforce mode injects an advisory and never denies a shell command. Quoted paths are supported, numeric and &> descriptor redirects are skipped, and unparsable commands fail open. Shell parsing remains heuristic.
  • Scratch directories: on POSIX, the host-managed per-user directory under system temporary roots is exempt by default. Additional directories come from rules.schema-guard.exempt_paths. Both candidate and root paths are resolved before directory containment is checked, so sibling filenames, similarly prefixed directories, and symlink escapes are not exempt.

Still open as of 2026-10: cancelled parallel calls are not filtered; their error-field text still needs live verification. Clusters remain shared within a project; whether sub-agents share their parent's session identifier is unverified. The effect of cooldown and proposed session scoping on existing reports has not been evaluated. No candidate refinement has been measured on a replay set against the public hooks, and the calibration commitment remains open. Scratch naming is a host convention, not a guarantee for other hosts or adapters. See known limitations and the configuration schema.

See it catch a lie

./scripts/demo.sh runs this in a throwaway directory in about 30 seconds — no Claude Code session needed, nothing of yours touched. Same hook, same evaluator, same decision JSON Claude Code acts on:

1. You declare the evidence BEFORE the work — a goal file in the project.

    ## acceptance
    [{"type": "file_exists", "path": "src/auth.py"},
     {"type": "cmd_succeeds", "cmd": ["python3", "-m", "unittest", "discover", "-s", "tests"]}]

2. SessionStart registers the claim (baseline snapshotted now, not later).

3. The agent does some work and says: "Done — auth is fixed, tests pass."

4. The turn tries to end. Axiom re-runs the declared evidence itself:

    decision: block
    reason:   AXIOM write verification failed: cmd_succeeds ['python3', '-m', 'unittest', 'discover', '-s', 'tests']: expected fresh command exits 0, actual exit 1. Fix the declared artifact or verification command, then stop again. Escape hatch: /axiom:enforce write-verify off

    The turn does not end. The agent gets the failure and keeps working.
    Note: the file EXISTS and the agent SAID tests pass — Axiom ran them.

5. The agent actually fixes it. Same claim, same evidence, re-run:

    no decision — the claim passed, the turn ends, the claim is cleared.

The demo forces enforce mode to show the block; on a real install that finding would be recorded, not blocked, until you say otherwise.

Install

/plugin marketplace add ryangu00/axiom
/plugin install axiom@axiom

Or point the marketplace at a local clone: /plugin marketplace add /path/to/axiom.

Every rule installs in observe mode: it records what it would have blocked and blocks nothing. You turn on enforcement per rule, when its findings have earned it — Axiom never switches itself on. This holds on every runtime, not just Claude Code: the shared CLI reads the same rule mode the Stop hook reads and tells each adapter whether it may act (CONTRACTS §5).


The one idea

A coding agent's output is testimony, not proof. In a single interactive session you are the check — you read what came back. In a loop — where the agent runs unattended, prompted by a schedule or a /goal condition instead of by you — nobody is reading. The agent's "done" becomes the premise of the next step, and a false one compounds silently.

Claude Code ships two ways to gate the end of a turn, and both ask a model. /goal closes a loop on a stopping condition with a small model judging whether you're done, and a Stop hook can be declared as type: "prompt" or type: "agent", which hands the same decision to a model with a prompt you write. They read the conversation. Done is a claim, not a proof.

Axiom is the other half of that, not a replacement for it: it checks the claim against the environment. A prompt hook asks a model whether the work looks finished; Axiom asks the filesystem whether the evidence you named before the work is there now. The model-based gate needs nothing declared up front and will form an opinion about anything — that is its advantage and the reason it cannot be relied on when the answer matters. Axiom needs you to say in advance what would count as done, and in exchange its verdict does not vary with phrasing, model version, or how convincingly the agent narrated its work.

They compose. Run a prompt hook for the judgement calls that have no crisp predicate, and let Axiom hold the ones that do.

The predicates themselves are decades-old primitives — exists, regex, hash, exit code — on purpose. What's new is where they live: declared before the work, snapshotted into a baseline at registration, held across sessions as a claim with an identity, and re-verified through a fresh evidence channel at the loop boundary. Old checks, new custody chain.

The evaluator and the claim lifecycle are host-agnostic Python; the Claude Code hooks are the first adapter, not the product. The adapter contract — three verbs any agent runtime can wire — is frozen in docs/ADAPTERS.md, and four runtimes ship against it today: Claude Code, Codex CLI, hermes-agent, and OpenClaw. The Codex CLI, hermes-agent and OpenClaw adapters were verified against the host's real consumption seam — the function or process boundary the host actually calls — not against its documented hook shape; for Claude Code the evidence is in-repo seam tests, not a live host run. The evidence table and the verified host versions are in ADAPTERS.md. The replay set behind the stuck-search threshold already draws on transcripts from two runtimes (the write-verify calibration is one runtime's, and neither has produced a published rate yet — see docs/CALIBRATION.md). Multi-runtime is where this tool came from, and the adapters are the packaging catching up to the data.

What it does, across the loop

Axiom is not an orchestrator — Claude Code already ships the loop primitives (scheduling, worktrees, skills, subagents, /goal). Axiom's design places one act of verification at each station of that loop. The last column is the honest part: it says what each row ships as today — hook is running code that acts on your turn, template is a convention you follow, library is opt-in and not wired into the runtime, roadmap is not written. docs/ROADMAP.md says what is deliberately not built and why — including the one gap that is the most obvious thing to ask for.

Loop station The unverified claim Axiom's check Ships as
Plan (forge a goal) "this plan is right" the goal template forces a risk rating, a rollback answer, executable done-criteria, and a per-task "if I delete this, does a criterion fail?" test before work starts; grilling settles the open decisions with the human in one round before the goal is forged, so nothing in done_criteria came from inference template
Execute "I finished it" write-verify — completion is checked against declared evidence predicates (files, git, fresh command runs), never inferred from a dirty working tree hook
"one more fix will work" (x8) stuck-search — failures are fingerprinted across attempts; at threshold it injects stop-retrying + search-externally guidance hook
Review "the code is fine" (said by the coder) the producer never signs off on itself; risk-rated work gets an independent, cross-family reviewer roadmap
Evolve "the machine learned a better rule" routing/threshold changes are proposed to a ledger a human approves — never written by an unattended loop roadmap
Remember "this recalled memory is current & safe" every lesson carries a timestamp + source and an unverified-memory prefix; instruction-shaped imports are quarantined library
Record "we'll remember why we did this" every re-plan is one changelog line in the goal file (timestamp / change / why) and closeout appends a route-outcome line; not left to the context window template

cmd_succeeds is fresh execution: its child process inherits the invoking user's permissions, environment, PATH, network, and filesystem. Argv-only execution, the executable allowlist, metacharacter rejection, and the timeout reduce injection surface; they are not a security boundary.

What actually ships in v1 (expanding the table's last column — the design is the whole table; the installed behavior is exactly this):

  • Ships now, as runtime hooks: the Execute checks (write-verify, stuck-search) and the guardrails (schema-guard, preflight).
  • Ships now, as discipline + templates: Plan (the goal template's risk / rollback / done-criteria discipline) and Record (goal-file changelog + route-outcome convention).
  • Ships now, as an opt-in library, not a runtime backend: the provider layer for write verification and Memory; predicate evaluation is shared with the runtime hooks.
  • Roadmap (not in v1): independent-reviewer wiring, and Evolve — the human-approved self-calibration engine (the observe-mode ledger already collects its input; the engine itself is staged). See Capability tiers.

Disciplined self-evolution

The reason "self-improving agents" tend to rot is that they rewrite their own rules unattended — one bad generalization poisons every later decision. Frameworks that do this at scale exist and are popular; that does not make it safe.

Axiom's stance: paths are free, goals are human-locked, rules are human-approved. The agent may change how it reaches a fixed goal (that's optimization, and every change leaves a changelog line). It may not silently change what the goal is (that's a collapsed premise — it escalates to you), nor rewrite its own routing/thresholds (that goes through a proposal you approve). Self-evolution, with a discipline on it.

Human-in-the-loop, concretely

"Human approval" is worthless as an adjective, so here is exactly where the human is in v1 — no more than that:

  • Nothing enforces until you say so. Every rule starts in observe mode. That is the gate: the default is record, don't act, and Axiom never promotes itself.
  • The decision is one command, and it is logged. /axiom:enforce write-verify on flips one rule and writes a mode_changed event to the ledger with who decided and what it changed from. The tool's decisions and its operator's decisions land in the same grep-able file — auditing only the machine's half would be auditing the wrong half.
  • What it would have done is on disk before it does anything. /axiom:report reads the ledger; would_have_blocked events carry the failed predicate and a timestamp. You approve enforcement against evidence from your own loops, not against this README.
  • Goal files are yours, on disk, in git. git diff is the audit trail, with no separate system to trust. What the runtime reads from a *.goal.md is exactly one thing: the ## acceptance block. Any other structure you keep there (done_criteria prose, route, a changelog) is a convention for humans and templates — Axiom does not parse it and does not enforce it.

Approval has to cost near-zero (read one line, type one command) or people route around it. That constraint is why there is no approval queue in v1.

Not shipped, and not claimed: a proposals queue for machine-suggested rule changes (that is the Evolve station — the ledger collects its input, the engine is not written), and the fail-closed egress gate sketched in docs/privacy-egress-design.md (a design note; the working implementation is coupled to a private knowledge base and is not part of v1). preflight advises, it does not block — it injects the recovery/scope questions as context and records the finding.

Zero-risk trial

/plugin marketplace add ryangu00/axiom
/plugin install axiom@axiom

On install, every hook is in observe mode: it records what it would have done and blocks nothing. After a few days:

/axiom:report
== Findings by rule ==

[write-verify] would-have-blocked: 3
  last incidents:
    fix auth bug | file_exists src/auth.py: expected file exists, actual missing | 2026-07-10T09:14Z

[schema-guard] would-have-blocked: 1

== Coverage ==
heartbeat days: 6

Each would-have-blocked is a real incident with the failed predicate and timestamp — enable blocking once you trust them: /axiom:enforce write-verify on.

If it caught nothing, that's an honest result — your loops are clean. Remove it and move on: /axiom:uninstall enumerates and deletes the files this plugin manages. (The host keeps plugin cache copies for a grace period; Claude Code manages those, not us — we don't claim to erase what we don't control.)

Born from incidents

Each hook exists because a specific failure cost real time in months of daily long-horizon agent operation:

  • write-verify ← agents reporting "done" on work that never touched disk, including a memory system that reported healthy for 13 days while silently not writing.
  • stuck-search ← a full night burned retrying an environment failure that had a verbatim fix sitting in a forum thread.
  • preflight ← an unmemory-verified destructive command that took a machine down.

The full taxonomy is in docs/failure-modes.md. What open-source ships in this space is either a methodology essay or a feature list; this is the residue of post-mortems.

Capability tiers

Axiom unlocks with your setup — nothing is forced on:

  • L0 (zero-config): the verification hooks, observe-mode by default. Works the moment you install.
  • L1 (cost-routing): if you run a multi-model setup (e.g. CCR) or an external agent CLI, the routing table + dispatch discipline apply — today as the template and convention in templates/, not as runtime code.
  • L2 (evolve): once the ledger has enough samples, threshold self-calibration proposes changes for you to approve. (Not shipped. The observe-mode ledger already collects its input; the engine is not written, and there is no config surface for it in v1.)

Memory: bring your own

Axiom's default memory is a local lessons.md the plugin manages. Writing to Claude Code's own on-disk memory is an explicit opt-in (memory_provider = "memory_md"), not the default. The provider layer is the extension point — it is not wired into the runtime hooks by default; point it at a real knowledge base — e.g. GBrain, an open-source brain layer — for graph-backed recall and write-verification receipts. The provider interface is not theoretical: the private predecessor these hooks derive from runs in production against exactly such a self-hosted store. The shipped hooks have no live data yet (see docs/CALIBRATION.md).

Why not just /goal?

/goal tells you when to stop. Axiom tells you what happened. /goal's completion judge reads the conversation and resets its baselines on --resume; Axiom's evidence lives in an on-disk goal file, is snapshotted into a baseline at registration, survives session loss, and is re-evaluated against the filesystem at the loop boundary. They stack: run /goal inside a task; let the goal file's ## acceptance block own what done means. Single-session, throwaway work? /goal alone is enough — writing a goal file for it is overhead, and Axiom says so.

Prior art & related work

We ran a competitive scan before first release. Claims about neighbors are held to the same evidence bar as claims about ourselves — each entry cites the project's own docs with an access date:

  • groundtruth — a Stop-hook completion-claim gate, the closest project to our flagship, and ahead of us on empirical calibration (1,272 real turns). It detects evidence in the same turn; Axiom pre-declares predicates with a baseline and a cross-session claim lifecycle.
  • claimcheck — a post-hoc CLI that auto-extracts claims from transcripts. Its extraction is broader than our declared-predicate contract; that method is weighed, and for now declined, in docs/ROADMAP.md.
  • tdd-guard — enforces a different discipline (TDD) at the same hook level, and marks the other side of a design split: it asks a model whether the work complies; we re-run predicates with no model in the path. Different costs, not a ranking.
  • nah — deterministic permissions at PreToolUse. Complementary station (should this run? vs did what you said happen actually happen?), and the bar we haven't met: it calibrates on a public corpus (101,194 tool calls) where ours is private and n=1.
  • oh-my-claudecode — much larger, sells multi-agent orchestration, and does ship a narrow version of this: a Stop hook that spots completion wording and blocks if the changed diff still holds TODO/stub/skipped-test markers. Same station, different question — it asks "did you leave junk in what you touched?", Axiom asks "did the specific thing you declared actually happen?" They compose.
  • planning-with-files — durable on-disk planning with an opt-in Stop gate that blocks on a plan still marked in_progress; the status string is agent-authored, which is the testimony we decline to trust. It also ships adapters for five-plus runtimes, so multi-runtime coverage is not our differentiator either.

Axiom independently derives from our own production incidents; where we took a method from a neighbor, it is credited by name. Full comparison, access dates, and what we adopted from whom: docs/PRIOR-ART.md.

Honest limits

  • Thresholds were set from one operator's workload (months of daily use, four execution lanes — varied, but n=1) and are not yet calibrated in the sense docs/CALIBRATION.md describes. A neighbor, nah, calibrates on a public corpus; that is the better standard and we say so. Your mileage will differ — that's what observe mode is for. Published false-positive/false-negative rates are not claimed. The commitment is tied to outside usage rather than a release number. At minimum it needs firings of the shipped hooks from operators other than the author — there are none yet. The full criteria, and what exists so far, are in docs/CALIBRATION.md. For write-verify: a replay of the private predecessor's trigger logic (70 labeled cases; 4 positives, 3 of them synthetic; no false positive on 55 negatives; mined and labeled 2026-07-10 from about six weeks of transcripts) and one batch of 50 labeled live firings from its first three days. For stuck-search: a replay set whose figures are not yet in a state to publish. None of it measures the hooks this repository ships, and none of it is an end-to-end error rate.
  • Claude Code's own roadmap is moving into this territory (hooks, /goal, /code-review). Axiom is designed to ride that roadmap, not race it — the hooks sit on the official hook API, the goal files sit above /goal. Where the platform absorbs a piece, you lose nothing you were depending on.

What this won't catch

Axiom catches the careless false "done" — the common one, where an agent reports success it never checked. It is not a seal against an agent actively trying to get past it. If you need that, you need a sandbox; this is for the loops you run outside one. The short version:

  • No claim, no check. Nothing registered a claim? Stop records unverified_completion and the turn ends. Axiom verifies evidence you declared — it does not infer claims from the transcript.
  • A predicate is a letter, not a spirit. file_exists passes on an empty file. Your predicates are the specification.
  • State sits at your agent's permission level. An agent with write access can delete its own active claim. Axiom raises the cost of a false "done" from free to deliberate; it does not make it impossible.
  • One block per stop cycle, on purpose. A failed claim blocks the stop once; if the agent immediately stops again (stop_hook_active), Axiom fails open and logs an escalation. A verifier that can wedge your agent forever is worse than no verifier. The claim stays active, so a later turn that still fails is blocked again — the cap is per re-entry, not per claim.
  • Observe mode blocks nothing. That's the default, and the point.

The full threat model, the audited implementation boundaries, the remaining hardening targets, and the n=1 calibration caveat: docs/KNOWN-LIMITATIONS.md. The ways the history privacy scan reported clean when it should not have, each with its test: docs/HISTORY-SCAN-FALSE-GREENS.md. The rest of the skeptic's list — isn't this just a prompt? you didn't invent this. a sandbox is the real answer. won't Anthropic build this in? stop hooks aren't reliable. no benchmark, no evidence. — is answered in docs/FAQ.md.

How this was built

Axiom was built with AI agents, under the discipline it ships. Claude Code orchestrated; OpenAI Codex wrote a substantial share of the code and, on every release, reviewed it adversarially as an independent second family; Codex and DeepSeek scored the design against a frozen rubric — the scoreboard, including a round we voided for being ungrounded, is in that file. A human made every decision that mattered and reviewed every change.

The rule we hold ourselves to: the producer never signs off on itself. Before each release the whole diff goes to a cross-family reviewer whose error modes are uncorrelated with the author's, and its findings are adjudicated with evidence, not accepted on authority. That is not a courtesy — it is the same principle as the plugin: a claim from the party that produced the work is testimony, and testimony gets audited.

What that bought, concretely, and where to check it: a cross-family pass found six contract-fidelity defects the author's own tests did not cover — five fixed, one rejected with its reasons, each one written into the commit that resolved it (git log --grep="cross-family review"). Two false-success defects in the verification core itself — the exact failure class this tool exists to catch — were found by review and fixed with regression tests before first release (git log --grep="resolve predicate paths identically"). The same review process caught a wrong claim about a neighbor in this repo's own prior-art page, and the correction is recorded there rather than quietly edited out.

Testing

Runs on Linux, macOS, and Windows. CI covers all three platforms with Python 3.10-3.12. Every gate below runs on every pull request and on pushes to main: lint, types, the full test suite across all platforms and Python versions, the Node adapter tests, and a privacy scan over tracked files.

Run the complete test suite through its canonical discovery command:

python3 -m unittest discover -s tests -v   # the Python suite (CI fails on zero discovery)
node --test adapters/openclaw/*.test.js    # the OpenClaw adapter suite
./scripts/demo.sh                          # the 30-second end-to-end demo (not in CI)

The legacy python3 scripts/selftest.py and python3 scripts/selftest_providers.py commands remain as compatibility entry points for their corresponding suites. CI uses canonical discovery, fails if it collects zero tests, and fails if a shipped adapter's tests are not collected — a quiet skip must never read as green.

License

Apache-2.0. See LICENSE.


RyanAI Lab · Extracted from the verification discipline we run in production. Updated 2026-09. Issues welcome.

About

Prove every "done." — verifies AI agent completion claims against evidence declared before the work. Observe-mode by default, zero-config, four runtimes.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages