Ran mutation testing (muteval) on experiment 0016 #6012
Replies: 4 comments
|
We were just talking about a similar topic the other day -> when/how to run a large eval suite and when to run a small one, and how many samples do you need to get statistically relevant information. For the experiment (presumably to show up as a PR in fullsend-ai/experiments) I think including @astefanutti @jflowers and @jwm4 are starting to look more at the evals in @fullsend-ai/agents and may find this interesting. |
|
Thinking about it a bit more - I think I didn't understand at first is that something like If I'm adding a new eval scenario to the suite - how do I know if my eval is sufficiently close to the edge of the behavior I want to test or not - or would a model simply always pass it trivially. With something like |
|
Yeah @ralphbean , your second comment is the use I care about, it tells you whether your eval would notice the agent getting worse. A scenario that still passes after I've weakened the skill/agent definition isn't testing much. I'll open it as an experiment PR in fullsend-ai/experiments: the 0016 run, which prompt edits survived (weakened modals, dropped instruction lines the classifier didn't react to), and the coverage gaps that leaves in the golden set. |
|
@ralphbean Raised the experiment PR for this: fullsend-ai/experiments#58 · experiment: mutation testing for…. It covers the 0016 run, which prompt edits the eval suite didn't catch, and the coverage gaps that leaves in the golden set. |
Uh oh!
There was an error while loading. Please reload this page.
testing-agents.mdlists mutation testing as one of the approaches for checking whether an eval suite would catch a regression, so after the significance layer (0026) I tried it on a eval in one your experiments. I maintain a tool for it — muteval — so that's what I used. I'm not here to pitch it; the idea is already in your doc, I just used the tool I had to run it.Experiment 0016 (the promptfoo PR-scope classifier) was a convenient fixture — small, public, already here. Not picking on it, just needed a real eval to point at. muteval degrades the prompt (weakening modals, dropping/reordering/paraphrasing instructions) and reruns 0016's 8 cases against each mutated version, checking whether any case fails.
The result was what you'd expect: a fair amount of the prompt can be changed without any case failing, since the clear-cut cases pass on the model's priors either way. That's normal for a small golden set — it's just the kind of thing mutation testing makes visible.
A couple of caveats: I ran it on gpt-4o-mini rather than 0016's Vertex Sonnet (the promptfoo path wants an OpenAI-compatible endpoint), so the broad read holds across models but individual results don't. And on fullsend's real agentic evals this would only make sense as a periodic / release-cadence check, not a per-PR gate, since each mutant reruns the whole agent.
Sharing it here and asking where, if anywhere, it's worth taking. I'm agnostic between:
Happy to follow your lead. @rh-hemartin
All reactions