Skip to content

fix: cache & concurrency — an optimization must never change a verdict - #59

Merged
AshwinUgale merged 1 commit into
mainfrom
fix/cache-and-concurrency
Sep 28, 2026
Merged

AshwinUgale merged 1 commit into
mainfrom
fix/cache-and-concurrency

Conversation

@AshwinUgale

Copy link
Copy Markdown
Owner

PR 3 of the audit follow-ups. --cache, --concurrency and skip-unchanged are optimizations, and each one could change a verdict. Every repro below ran on main first. Runs without --cache / --concurrency score exactly as before. All keyless examples are identical to main.

Before → after (the audit's own repros, rerun)

Repro main this PR
Edit contains("X1") → contains("ZZZ"), re-run with --cache no cache: baseline_failed; with cache: valid both baseline_failed
Two evals sharing a label (GEval, GEval) no cache 23.5%; with cache 0% both 23.5%
Edit the model inside run() (prompt mode), with cache no cache baseline_failed; with cache valid both baseline_failed
Tighten a threshold 0.5 → 0.9, with cache true 1.0; cache 0.0 both 1.0
Stateful deepeval-shaped metric, --concurrency 8 5/5 runs differ from serial (0%–85%) 0/5 differ
run() writes case["used_context"], --concurrency 8 5/5 differ 0/5 differ
Same suite, skip-unchanged on (runs=1) vs off (runs=2) 0/11 vs 4/11 kills 4/11 both
deepeval adapter + --max-calls 20 80 paid calls 20
{"final","trace"} dict output + --cache baseline_errored (sqlite) cached, valid

How

  • Cache v2 keys. An output is keyed on System.key() + case + a fingerprint of run. An outcome is keyed on the output + case + a fingerprint of the eval, never the label. The new muteval/fingerprint.py hashes code (recursively), defaults, closure values, the simple globals an eval reads, and a callable object's type and state.
    • The rule is to prefer a miss over a collision.
    • It is stable across processes, and computed once per run.
    • cache_version forces a new fingerprint for an eval that depends on a file or a remote rubric.
    • Old entries are never read.
  • run() side effects are replayed. The post-run case is stored with the output and restored on a cache hit. Anything that can't round-trip through JSON just isn't cached.
  • Isolation. Each (mutant, case, run) cell gets a private deep copy of the case. The deepeval adapter measures a shallow copy of its metric per call (it falls back to a lock if the metric can't be copied). Queued mutants are cancelled once the budget is hit.
  • Skip-unchanged now requires the identical output and the identical post-run case state.
  • The deepeval/ragas adapters set is_llm.

Also fixed (found while testing)

  • System(prompt=..., context=[...], tools=[...]), the README's own form, crashed mutant generation with unhashable type: 'list'. This bug is on main too; lists are now normalized to tuples.
  • Eval labels are unique and aligned (GEval, GEval#2). A too-long eval_names list is an error. Zero-config labels use the full spec.
  • --max-mutants -1 silently dropped the last mutant; it's now rejected (exit 2), as is a negative --sample.
  • System.key() crashed on mixed extra key types. One unkeyable baseline sample marked its case undetermined. The ragas adapter split a string context into characters.

API note

Cache.get_outcome / set_outcome now take (output, case, eval_fingerprint). lookup_output / store_output carry the post-run case. Passing a Cache to run_mutation_testing is unchanged.

Verification

  • New tests/test_cache_concurrency.py (19 tests): fingerprints, every cache repro as "with cache == without cache", side-effect replay with zero calls, isolation, concurrency 1 vs 8, budget, labels, list context.
  • pytest 540 passed, including slow tests. ruff check / ruff format --check / mypy clean. The diff was scanned for token-shaped literals.
  • Docs: README (cache fingerprinting, thread-safety note), LIMITATIONS (cache limits, concurrency contract), CHANGELOG, CLAUDE.md.

🤖 Generated with Claude Code

- cache v2 keys: outputs on System.key() + case + fingerprint(run); outcomes on
  output + case + fingerprint(eval) — never the label. An edited eval / changed
  threshold / edited run() no longer serves stale verdicts; evals sharing a
  label no longer share a result. New muteval.fingerprint (code, closures,
  defaults, globals; cache_version override), computed once per run.
- run()'s writes into the case are stored with the output and replayed on a
  cache hit; dict outputs cached (was sqlite ProgrammingError).
- per-(mutant, case, run) private case copies; deepeval adapter measures a
  shallow copy of its metric per call (lock fallback): --concurrency no longer
  changes verdicts. Queued mutants cancelled on BudgetExceeded.
- skip-unchanged also requires identical post-run case state.
- deepeval/ragas adapters set is_llm (budgeted, ordered after cheap checks,
  never called by the canary).
- System(context=[...], tools=[...]) (the README's form) no longer crashes
  mutant dedup; System.key() handles mixed extra key types.
- unique/aligned eval labels; zero-config labels use the full spec; negative
  --max-mutants/--sample rejected; ragas string context not split into chars;
  unkeyable baseline samples ignored.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@AshwinUgale
AshwinUgale merged commit 345ef98 into main Sep 28, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant