Repository navigation
fix: cache & concurrency — an optimization must never change a verdict - #59
Merged
Merged
Conversation
- cache v2 keys: outputs on System.key() + case + fingerprint(run); outcomes on output + case + fingerprint(eval) — never the label. An edited eval / changed threshold / edited run() no longer serves stale verdicts; evals sharing a label no longer share a result. New muteval.fingerprint (code, closures, defaults, globals; cache_version override), computed once per run. - run()'s writes into the case are stored with the output and replayed on a cache hit; dict outputs cached (was sqlite ProgrammingError). - per-(mutant, case, run) private case copies; deepeval adapter measures a shallow copy of its metric per call (lock fallback): --concurrency no longer changes verdicts. Queued mutants cancelled on BudgetExceeded. - skip-unchanged also requires identical post-run case state. - deepeval/ragas adapters set is_llm (budgeted, ordered after cheap checks, never called by the canary). - System(context=[...], tools=[...]) (the README's form) no longer crashes mutant dedup; System.key() handles mixed extra key types. - unique/aligned eval labels; zero-config labels use the full spec; negative --max-mutants/--sample rejected; ragas string context not split into chars; unkeyable baseline samples ignored. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR 3 of the audit follow-ups.
--cache,--concurrencyand skip-unchanged are optimizations, and each one could change a verdict. Every repro below ran onmainfirst. Runs without--cache/--concurrencyscore exactly as before. All keyless examples are identical tomain.Before → after (the audit's own repros, rerun)
maincontains("X1")→contains("ZZZ"), re-run with--cachebaseline_failed; with cache:validbaseline_failedGEval,GEval)run()(prompt mode), with cachebaseline_failed; with cachevalidbaseline_failed--concurrency 8run()writescase["used_context"],--concurrency 8--max-calls 20{"final","trace"}dict output +--cachebaseline_errored(sqlite)How
System.key()+ case + a fingerprint ofrun. An outcome is keyed on the output + case + a fingerprint of the eval, never the label. The newmuteval/fingerprint.pyhashes code (recursively), defaults, closure values, the simple globals an eval reads, and a callable object's type and state.cache_versionforces a new fingerprint for an eval that depends on a file or a remote rubric.run()side effects are replayed. The post-run case is stored with the output and restored on a cache hit. Anything that can't round-trip through JSON just isn't cached.is_llm.Also fixed (found while testing)
System(prompt=..., context=[...], tools=[...]), the README's own form, crashed mutant generation withunhashable type: 'list'. This bug is onmaintoo; lists are now normalized to tuples.GEval,GEval#2). A too-longeval_nameslist is an error. Zero-config labels use the full spec.--max-mutants -1silently dropped the last mutant; it's now rejected (exit 2), as is a negative--sample.System.key()crashed on mixedextrakey types. One unkeyable baseline sample marked its case undetermined. The ragas adapter split a string context into characters.API note
Cache.get_outcome/set_outcomenow take(output, case, eval_fingerprint).lookup_output/store_outputcarry the post-run case. Passing aCachetorun_mutation_testingis unchanged.Verification
tests/test_cache_concurrency.py(19 tests): fingerprints, every cache repro as "with cache == without cache", side-effect replay with zero calls, isolation, concurrency 1 vs 8, budget, labels, list context.pytest540 passed, including slow tests.ruff check/ruff format --check/mypyclean. The diff was scanned for token-shaped literals.🤖 Generated with Claude Code