Skip to content

fix(agent): enforce approval gate on allowFinalResponse path and validate predicate args (#54) - #94

Open
LukasParke wants to merge 2 commits into
mainfrom
fix/54-approval-gate-validated-args
Open

fix(agent): enforce approval gate on allowFinalResponse path and validate predicate args (#54)#94
LukasParke wants to merge 2 commits into
mainfrom
fix/54-approval-gate-validated-args

Conversation

@LukasParke

Copy link
Copy Markdown
Contributor

Summary

Two independent ways the tool-approval gate could be bypassed, letting a tool the user was supposed to approve execute unguarded.

Bug 1 — allowFinalResponse skipped the approval gate entirely

When a stopWhen condition halted the loop on a turn that still carried tool calls, the final-response path called executeToolRound(pendingToolCalls, turnContext) directly, with no handleApprovalCheck — unlike the two in-loop call sites, which gate every round.

Consequences:

  • A tool marked requireApproval: true (or gated by a predicate) executed without approval.
  • Hook-based deny never fired on this path either: hookDeniedCalls is only populated inside handleApprovalCheck, and neither executeToolRound nor executeSingleToolCall partitions on approval internally. So a PermissionRequest hook returning deny was silently ignored.

This is reachable in ordinary use — any run with a stopWhen (including the default step limit) that halts on a turn carrying a gated call.

Bug 2 — the approval predicate saw different arguments than execute

A function-based requireApproval was invoked with toolCall.arguments, which at that point is only JSON.parsed (see extractToolCallsFromResponse). execute receives the arguments after validateToolInput (z4.parse) runs, which applies the schema's defaults, coercions, and transforms — so the two disagreed whenever the schema does any of those.

Concretely, with inputSchema: z.object({ dangerous: z.boolean().default(true) }) and a model emitting {}:

  • predicate saw dangerous: undefined → no approval required
  • execute then ran with dangerous: true

The predicate was deciding on values that were never the ones used. A stale comment in conversation-state.ts asserted the arguments were "already parsed and validated against the tool's Zod inputSchema" — that was false, and is corrected here.

The fix

  • model-result.ts: added if (await this.handleApprovalCheck(pendingToolCalls, turnNumber, currentResponse)) { return; } before the final-response executeToolRound, mirroring the in-loop call sites. On pause, handleApprovalCheck already persists pendingToolCalls + status: 'awaiting_approval', records auto-approved calls as unsent results, and sets finalResponse, so the early return is consistent: nothing executed, so there is no round to record, and it correctly skips both markStateComplete() and the final text-coercion request. sessionEndReason stays 'max_turns', matching the sibling HITL pause return in the same block.
  • conversation-state.ts: the predicate's arguments are now z4.safeParsed against the tool's inputSchema — the same zod entry point validateToolInput uses — so the predicate sees exactly what execute will receive. Fails closed: if the arguments don't satisfy the schema, approval is required rather than judging a value execute would never see. Zod is imported directly rather than reusing validateToolInput because tool-executor.ts imports conversation-state.ts; sharing the helper would create an import cycle. The false comment is replaced with one explaining the actual invariant.

Test coverage

New packages/agent/tests/unit/approval-gate-regressions.test.ts (5 tests). All 4 regression tests were verified red against unmodified code before implementing, then green after — confirmed by stashing the fixes and re-running (4 failed / 1 passed → 5 passed).

  • predicate sees schema defaults ({}{ dangerous: true }, approval required)
  • predicate sees schema coercions ({ amount: '500' }{ amount: 500 }, so > 100 compares numerically rather than lexicographically)
  • fails closed on invalid arguments, and the predicate is not called at all
  • allowFinalResponse gate: stepCountIs(1) firing on a turn carrying a requireApproval call — asserts the tool does not execute, the run pauses with awaiting_approval, the gated call is on pendingToolCalls, requiresApproval() is true, and no final text-coercion request is made. Structured so the first round completes with an ungated tool, ensuring the break lands on the post-loop path rather than being caught by the pre-loop gate.
  • control: non-gated tools still execute normally on the allowFinalResponse path and the final response is still produced (guards against over-blocking).

Per this repo's practice, I re-audited consumers after the contract change: both executeToolRound call sites are now gated, and partitionToolCalls / toolRequiresApproval have no other non-test callers.

Verification

pnpm turbo run build typecheck lint test --filter=@openrouter/agent — all 4 tasks pass. Full unit suite: 648 tests / 52 files passing, no type errors.

Judgment call

The call-level requireApproval override (options.requireApproval) was left as-is. It receives the whole ParsedToolCall, not just the arguments — a deliberately different public contract from the tool-level predicate — and normalizing its arguments would change a published signature's semantics. Worth a follow-up decision, but out of scope for a patch fix; flagging it rather than changing it silently.

Fixes #54

🤖 Generated with Claude Code

perry-the-pr-reviewer[bot]

This comment was marked as resolved.

@LukasParke
LukasParke marked this pull request as ready for review August 6, 2026 14:25

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

Open in Devin Review

Comment on lines +6130 to +6132
if (await this.handleApprovalCheck(pendingToolCalls, turnNumber, currentResponse)) {
return;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Permission prompt can be raised twice for the same tool call

The pending tool calls are re-checked for approval (handleApprovalCheck(pendingToolCalls, ...) at packages/agent/src/lib/model-result.ts:6130) even when the very same calls were already checked before the loop, so a permission handler can be asked about — and a user prompted for — the same call twice.
Impact: Integrations that log or surface an interactive prompt from the permission handler see duplicate prompts/audit entries for one tool call.

Why the initial response's calls get gated twice when the stop condition fires on the first iteration

The pre-loop gate at packages/agent/src/lib/model-result.ts:5858 runs handleApprovalCheck(toolCalls, 0, currentResponse) on the initial response's calls. If a PermissionRequest hook returns allow or deny, handleApprovalCheck returns false (see packages/agent/src/lib/model-result.ts:2892-2900) and the loop is entered. If shouldStopExecution() fires on the very first iteration (e.g. stepCountIs(0), or a token/time-based condition already satisfied), the loop breaks with stoppedByStopWhen = true before any follow-up request, so currentResponse is still the initial response. The post-loop path then extracts the same tool calls (packages/agent/src/lib/model-result.ts:6104-6106) and calls handleApprovalCheck again, which re-runs partitionToolCalls and emitPermissionRequest for each gated call — emitPermissionRequest has no per-call memoization (packages/agent/src/lib/model-result.ts:2543-2590). The second test in the new suite (stepCountIs(0)) exercises exactly this double-gating path, but with no gated tool so the duplication isn't observed.

Mid-loop rounds are unaffected because makeFollowupRequest replaces currentResponse with a response whose calls have not yet been gated.

Prompt for agents
In packages/agent/src/lib/model-result.ts, the new approval gate on the allowFinalResponse path (around line 6130) can re-gate tool calls that were already gated by the pre-loop gate at line 5858. This happens when the stop condition fires on the first loop iteration: the loop breaks before any follow-up request, so `currentResponse` is still the response whose calls were already passed through `handleApprovalCheck`. Re-gating re-emits the `PermissionRequest` hook for the same call (emitPermissionRequest has no memoization), producing duplicate permission prompts/audit records, and re-runs any function-based requireApproval predicate. Consider tracking which tool call ids (or which response id) have already been through handleApprovalCheck on this run and skipping the re-check for those, or only gating calls that have not yet been gated.
Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +319 to 322
const parsed = z4.safeParse(tool.function.inputSchema, toolCall.arguments);
if (!parsed.success || !isRecord(parsed.data)) {
return true;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Malformed tool arguments now stall or abort the run instead of returning an error to the model

When a tool's arguments don't match its schema the call is now unconditionally marked as needing human approval (return true at packages/agent/src/lib/conversation-state.ts:321), so a run that has no place to store a pause fails outright instead of letting the model see the validation error and retry.
Impact: A single badly-formed tool call from the model can abort the whole run, or pause it waiting for a human to approve a call that can never succeed.

Fail-closed path interacts badly with the pause requirements

For a tool whose requireApproval is a function, toolRequiresApproval now safeParses toolCall.arguments against tool.function.inputSchema and returns true on failure. That verdict flows into partitionToolCallshandleApprovalCheck, which, when no StateAccessor is configured, throws Tool(s) require approval but no state accessor is configured: ... (packages/agent/src/lib/model-result.ts:2902-2909) — the run rejects. Previously the predicate was consulted on the raw arguments and, if it returned false, the executor's own validateToolInput (packages/agent/src/lib/tool-executor.ts:265) threw a Zod error that was converted into a tool error output the model could recover from.

Even with a state accessor, the run pauses and asks the user to approve a call whose arguments will fail validation the moment it is executed.

A narrower fail-closed rule (e.g. only fail closed for calls that would otherwise be auto-approved and would actually execute, or letting schema-invalid calls fall through to the executor's normal validation error) would preserve the security intent without turning malformed model output into a hard failure.

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

LukasParke and others added 2 commits August 6, 2026 16:30
…date predicate args (#54)

Two ways the tool-approval gate could be bypassed.

The allowFinalResponse path executed pending tool calls with no approval
check. When a stopWhen condition halted the loop on a turn that still
carried tool calls, the final-response path called executeToolRound
directly, skipping the gate the normal loop applies on every round. A
tool marked requireApproval would execute unguarded, and since the
PermissionRequest hook's deny bookkeeping lives inside
handleApprovalCheck, hook-based deny never fired on this path either.

Function-based requireApproval also received unvalidated arguments: the
predicate got the raw JSON-parsed wire payload while execute receives
the values after the tool's Zod inputSchema runs, so any default,
coercion, or transform made the two disagree. The predicate now parses
with the same schema the executor uses, and fails closed (requires
approval) when the arguments don't satisfy it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a regression test for the PermissionRequest hook returning 'deny'
on the post-loop allowFinalResponse path: the denied tool must not
execute, the run must not pause for a human, and the hook's reason must
be recorded in state as a synthesized rejected output for the call.

Verified load-bearing: with the approval-gate fix in model-result.ts
reverted, the hook handler is never invoked (0 calls) because that path
had no approval check at all, which is where hookDeniedCalls is
populated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@LukasParke
LukasParke force-pushed the fix/54-approval-gate-validated-args branch from 808690d to 3ae034e Compare August 6, 2026 21:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Approval gate not enforced on the allowFinalResponse path; predicate also sees pre-normalization args

1 participant