Skip to content

fix(agent-runtime): keep image and blob bytes out of tool result fallback text - #1570

Open
xpeng5278-web wants to merge 1 commit into
vastsa:mainfrom
xpeng5278-web:fix/tool-result-binary-fallback
Open

xpeng5278-web wants to merge 1 commit into
vastsa:mainfrom
xpeng5278-web:fix/tool-result-binary-fallback

Conversation

@xpeng5278-web

Copy link
Copy Markdown
Contributor

Fixes #1569

What

In the host tool bridge (packages/agent-runtime/src/runtime.ts), the MCP content branch and the bare content-block-array branch now stringify their no-text fallback through a small withoutInlineBinary helper. It drops image data and resource blob bytes and keeps everything else (type, mimeType, uri, other fields), so the model is still told what was returned.

Image block extraction, details, the images branch, hostContentBlockText and toolResultFromUi are unchanged.

Docs: docs/spec/03-runtime/02-agent-runtime.md §3.3 now states that the no-text JSON fallback also keeps image and blob bytes out of the model-visible text.

Why

textParts only holds text, resource.text and resource_link content. An image-only or blob-only result therefore fell back to JSON.stringify(rawContent), base64 included. Vision models received the image twice (image block + base64 text). Text-only models received the base64 as plain text. One screenshot could add hundreds of KB to the context. The images branch already strips images before stringifying, and #1360 / #1551 already keep binary out of the text channel; this aligns the two remaining fallbacks.

How tested

  • New cases in packages/agent-runtime/src/runtime.test.ts:
    • image-only MCP content and image-only bare array → no base64 in text, image block kept, imageCount: 1
    • blob-only MCP resource → no blob bytes in text, uri and mimeType still present
    • text-only model + image-only result → no base64 in text, no image block
  • On unmodified main (eec988c0f): the 4 new cases fail (expected '…' not to contain 'cG5nLWJ5dGVz'). With this change: runtime.test.ts 311/311 pass.
  • packages/agent-runtime: vitest run → 1333 tests pass; tsc -p . --noEmit clean.
  • node scripts/check-pr-base-main.mjs, node scripts/check-architecture.mjs, pnpm lint:biome, docs/scripts/check-docs.mjs passed.
  • Electron E2E: not run on this Linux box.

Scope

Only the no-text JSON fallback in those two branches, plus the regression tests and one spec sentence. Results that have a text block, the images shape, and persisted-row restore are unchanged. Other binary-carrying block kinds (e.g. MCP audio) are not touched here.

…back text

An image-only or blob-only host result has no text block, so the MCP
content and bare content-block fallbacks stringified the raw payload.
That copied base64 into the model text channel, duplicating the image
block for vision models and dumping bytes as text for every other model.

Strip image data and resource blob bytes in that JSON fallback only.
Type, mimeType, and uri stay so the attachment is still described.

@muzimu217 muzimu217 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified locally on macOS (patch applied to a detached worktree at current HEAD eec988c0f): agent-runtime full suite 96 files / 1333 tests pass (the +4 over main are the new binary-stripping cases).

The fix is the right companion to #1551, and the helper is cleanly scoped:

  • withoutInlineBinary recurses the fallback JSON, dropping only two things — image data and resource blob bytes — while keeping type, mimeType, uri and every other field. So an image-only or blob-only tool result no longer floods the context with megabytes of base64, yet the model still sees what the tool returned and where (the resource uri survives), which is the actionable part.
  • Applied on both the bare content-block-array branch and the MCP content branch — the same two paths #1551 touched — so the two PRs stay symmetric.
  • Image block extraction and the images/details surfaces are untouched, so vision consumers lose nothing.

One boundary note for the record (non-blocking): stripping is on the fallback-text path only — a tool result that returns proper image blocks still feeds them to vision models as blocks. That's the intended split and worth keeping as-is.

Nothing blocking. fix scope, fits the window.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Image-only / blob-only MCP and plugin tool results put base64 into the model text

2 participants