Skip to content

feat(sdk)!: 导出 APIError,HTTP 错误和流内错误按 Kind 跨 provider 分类 - #64

Merged
HoneyBBQ merged 11 commits into
mainfrom
feat/errors-apierror
Sep 30, 2026
Merged

HoneyBBQ merged 11 commits into
mainfrom
feat/errors-apierror

Conversation

@HoneyBBQ

@HoneyBBQ HoneyBBQ commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

问题

main 上的 sdk.APIError(#63)只有状态码、状态文本、通用提取的 message、request ID 和原始 body,没有失败原因的分类。调用方要区分 key 无效、额度耗尽、provider 过载,仍然只能看状态码和文本,而同一个状态码在各 provider 下含义不同:OpenAI 返回 429 时可能是限流,也可能是额度耗尽。它的 Error() 还附带最多 4 KiB 的响应体,响应体可能回显终端用户的输入,打印错误就会写进日志。

流也有同样的问题。Anthropic 的 error 事件读错了字段,流中途的 overloaded 到调用方只剩 unknown error。流在终止事件之前断开,会被当作成功结束。

改动

  • 第一个提交撤销 chore(error): expose APIError to sdk #63(9b039e1)。NewAPIError、NewAPIErrorFromResponse、IsStatus 和 ClassifyProbeError 随之删除,由本 PR 的类型取代。
  • 导出 sdk.APIError,一律以指针使用。字段有 Provider、StatusCode,provider 自己的 Type、Code、Message,以及 RequestID、Kind、响应的 Header 和原始 Body。Error() 包含 provider、状态码、type/code、message 和 request ID。它不包含 body 和任何响应头,因为 body 可能回显终端用户的输入。Provider 是发出请求的 provider 的 Name()。OpenCode Go 经 Completions、Responses 或 Messages 发请求,错误里是这三者之一的名字。
  • RetryAfter() 先读 retry-after-ms,再读 Retry-After,支持秒数和 HTTP-date 两种写法。
  • sdk.ErrorKind 跨 provider 分类失败原因,取值有 authentication、permission_denied、quota_exhausted、rate_limited、server_error、unknown。这是开放集合,KindOf(err) 可以从任意错误链中读出 Kind。各 provider 先按自己的错误 type 或 code 映射,都匹配不上时再看 HTTP 状态码:401、402、403、429,以及 501、505 以外的 5xx。provider 的 code 优先于状态码,所以 OpenAI 返回 429 且 code 为 insufficient_quota 时,Kind 是 quota_exhausted。分类从不依据 message 文本。
  • FetchJSON、FetchRaw 和 FetchSSE 遇到非 2xx 响应时返回 *sdk.APIError,按 provider 使用的错误格式解码。多个包都会收到的格式在 internal/errorformat 中解码:OpenAI 格式(completions、responses、embedding、images,以及经这些包访问的 Bedrock)、Google 的 google.rpc.Status(generativeai、embedding、transcription)、OpenRouter 和阿里云 DashScope。Anthropic、Codex、Copilot 和方舟视频的解码器留在各自的包里。Google 先按 ErrorInfo 的 reason 分类,只认 google.api.ErrorReason 中的值,没有映射时再看 canonical status。Gemini 的每分钟配额和每日配额用尽都返回 429 RESOURCE_EXHAUSTED,所以 QuotaFailure 中有一条违反的是每日配额时,Kind 为 quota_exhausted。每日配额按 quotaId 含 PerDay 或 Daily 判断,与 Gemini CLI 相同。OpenRouter 的 RequestID 取自 x-generation-id 响应头。API 参考没有写这个头,但聊天补全的错误响应带有它,Access-Control-Expose-Headers 也列出了它。
  • 非 2xx 响应体最多读取 1 MiB,再丢弃至多 64 KiB 的剩余部分,让连接可以复用。更长的剩余部分不再读取,连接随 resp.Body 关闭。
  • provider 改用 %w 包装错误,不再格式化成字符串。删除 utils.APIError、parseAPIError 和 Detail()。
  • Copilot 集成测试按 *sdk.APIError 的 StatusCode 和 Code 判断 403 和 model_not_supported,不再匹配错误文本。
  • 流内错误事件返回 *sdk.APIError,StatusCode 为 0,Body 是事件数据。覆盖 Anthropic 的 error,以及 Responses 和 Codex 的 error 与 response.failed。Anthropic 事件按官方文档读 error.type 和 error.message。Responses 的 error 事件两种结构都接受:code 和 message 在顶层(API 参考的写法),或嵌套在 error 对象里(Codex CLI fixture 的写法)。
  • 失败的流只发一个 ErrorPart,随后是一个 FinishPart,结束原因为 FinishReasonError,带上已上报的 usage。此前 Anthropic、Completions、Copilot 和 Google 收到损坏的数据块时会发两个 ErrorPart。
  • 流在终止事件之前结束时,发出一个包装新 sentinel sdk.ErrStreamIncomplete 的 ErrorPart。各后端的终止事件是:Anthropic 为 message_stop;Responses 和 Codex 为 response.completed 或 response.incomplete;Completions 和 Copilot 为 [DONE] 或带 finish_reason 的数据块;Google 为候选的 finishReason 或 promptFeedback.blockReason。
  • Responses 在 200 响应体里带 error 对象、阿里云图片在 200 响应体里带业务错误码时,返回 StatusCode 为 0 的 *sdk.APIError。
  • 新增独立模块 internal/sdkdiff,官方 SDK 只是它的测试依赖,主模块的依赖不变。CI 对它单独运行 lint 和测试。

本 PR 不处理直接调用 http.Client 的路径(语音、转录、multipart 图片、视频下载、WebSocket)和 Provider.Test,它们仍返回无类型的错误。

验证

  • go build ./...、go vet ./...、go vet -tags integration ./...、go test ./... -short -count=1 -race 和 golangci-lint run ./... 均通过,go mod tidy 无 diff。
  • conformance 套件为六个聊天 provider 加了错误 fixture,取自各 provider 的文档或官方 SDK 的录制数据,每条都注明来源 URL。每个错误用例在 Generate 和 Stream 两条路径上都断言:errors.As、StatusCode、Type、Code、Message、RequestID、Kind,错误上保留的 body 和响应头,以及 Error() 不含 API key 和 body。故意写错期望的 Kind,套件会失败。
  • 套件还对每个聊天 provider 跑三种流:流内错误、终止事件前关闭的流、损坏的数据块。Anthropic 和 Codex 的 fixture 逐字取自官方文档和 Codex CLI。
  • OpenCode Go 的三个 delegate 收到 401 和 429 时,Generate、Stream 和 TestModel 都返回该 delegate 的 *sdk.APIError,StatusCode 和 Kind 正确。
  • 每个解码器都有表格测试,internal/errorformat 的每个分支都有用例。embedding、images(生成与 JSON 编辑)、videos 和 Google 转录有端到端的 httptest 用例。Codex 流内的 error 事件和带 usage 的 response.failed 经过完整的流路径测试,后者断言 FinishPart 保留了 usage。
  • Gemini 的 key 无效、缺少 key 和模型不存在三种错误体取自 2026-09-29 对 generativelanguage.googleapis.com 的实际请求,原文写入测试。key 无效为 authentication,模型不存在为 unknown。缺少 key 的响应是 403 PERMISSION_DENIED,没有 ErrorInfo,所以是 permission_denied。
  • Gemini 配额用尽的 429 取自 gemini-cli 两个 issue 中引用的 API 响应。只违反每日配额、同时违反每日和每分钟配额,两种都是 quota_exhausted。只违反每分钟配额是 rate_limited。
  • 2026-09-29 对 OpenRouter 和阿里云 DashScope 图片发了真实请求。OpenRouter 的 key 无效为 401 authentication,模型不存在为 400 unknown。后者的 RequestID 是 x-generation-id 的值,这条错误体和响应头原文写入测试。
  • DashScope 的 wan2.2-t2i-flash、qwen-image-plus 和 wan2.7-image-pro 生成成功。key 无效为 401 InvalidApiKey,Kind 是 authentication。尺寸越界在任务结果里报告,StatusCode 为 0,Code 为 InvalidParameter。
  • 经一个 OpenAI 与 Anthropic 兼容的中转网关,对 Responses、Completions、Codex 和 Anthropic Messages 发了真实请求。成功的生成与流式请求结果正常。401、403、404 和流内的 response.failed 都返回 *sdk.APIError,状态码和 Kind 正确。中转网关的错误体不是官方格式,没有用作 fixture。
  • 契约测试列出了五个枚举的全部取值和各自的 Kind:OpenAI ResponseErrorCode、google.rpc.Code、google.api.ErrorReason、OpenRouter ApiErrorType 和 Anthropic ErrorType。枚举取自固定提交的 OpenAPI 或 proto 文件。规格中的错误示例原样解码,OpenRouter 每个状态码的示例都在内。
  • 差分测试把同一份错误响应分别交给 twilight 和官方 Go SDK(openai-go、anthropic-sdk-go、go-genai)解析。两边读出的状态码、type、code、message 和 request ID 一致。覆盖范围是 Chat Completions、Responses 和 Anthropic Messages 的 HTTP 错误与流内错误,以及 Gemini 的 HTTP 错误。官方 SDK 不解析的字段不比较,例如 Anthropic 的 message。故意改错 request ID 的读取顺序或丢掉 message,测试会失败。
  • 官方 SDK 的错误测试数据都是手写的,没有可复用的真实流量。
  • Google 和方舟没有文档化的 request ID 响应头,它们的 RequestID 为空。OpenRouter 的视频接口和 401 响应不带 x-generation-id,RequestID 同样为空。Copilot 的错误格式没有公开,fixture 依照 VS Code Copilot 客户端的解析逻辑编写。阿里云 DashScope 没有机器可读的公开契约,只按文档测试。
  • 没有实际抓包确认 api.openai.com 发的是哪种 error 事件结构,两种结构都有测试。Gemini 流中途的错误对象不做解码。

⚠️ No human QA

@HoneyBBQ
HoneyBBQ force-pushed the feat/errors-apierror branch from e8b1681 to 4970d9d Compare September 28, 2026 12:31
@HoneyBBQ
HoneyBBQ marked this pull request as ready for review September 28, 2026 12:45
@HoneyBBQ HoneyBBQ changed the title feat(sdk)!: export APIError with a cross-provider Kind feat(sdk)!: export APIError with a cross-provider Kind for HTTP and stream errors Sep 28, 2026
@HoneyBBQ
HoneyBBQ added this pull request to stack #67 September 28, 2026 13:17
@HoneyBBQ HoneyBBQ changed the title feat(sdk)!: export APIError with a cross-provider Kind for HTTP and stream errors feat(sdk)!: 导出 APIError,HTTP 错误和流内错误按 Kind 跨 provider 分类 Sep 28, 2026
@HoneyBBQ
HoneyBBQ force-pushed the feat/errors-apierror branch from 5e73fb2 to fc05db5 Compare September 28, 2026 19:55
@HoneyBBQ
HoneyBBQ marked this pull request as draft September 28, 2026 20:18
@HoneyBBQ
HoneyBBQ force-pushed the feat/errors-apierror branch 4 times, most recently from 67e0352 to 0c00e36 Compare September 29, 2026 22:11
This reverts commit 9b039e1. The exported error type is replaced by the
one in the following commits. OpenCode Go, written against it in #49,
uses utils.APIError until then.
Provider failures reached callers only as text: the HTTP layer's error type
was internal, so status, provider codes and the response body could not be
recovered with errors.As, and downstream code matched error strings.

APIError is the exported type for a failure the provider reported. It keeps
the provider's type, code, message and request ID verbatim, the response
headers and raw body, and a cross-provider Kind. StatusCode is 0 when the HTTP
status line did not report the failure. Error() carries provider, status,
type/code, message and request ID, never the body or any header, because the
body may echo end-user input. RetryAfter parses retry-after-ms, delta-seconds
and HTTP-date. ErrorKind is an open set; KindOf reads it from an error chain.
FetchJSON, FetchRaw and FetchSSE now build the error with NewHTTPError: the
status gives a fallback Kind, and a per-provider ErrorDecoder fills type,
code, message and request ID from the provider's own error body and headers
and refines Kind from its type or code, which win over the status (OpenAI's
429 insufficient_quota is a quota failure, not a rate limit). Nothing is
classified from message text.

Each package passes its name and decoder through RequestOptions. OpenAI's
format, shared by the OpenAI-compatible packages, and Google's
google.rpc.Status live in internal/errorformat; Anthropic, Codex, Copilot,
DashScope, Ark and OpenRouter decode their own. The type/code-to-Kind
functions are separate from the decoders so stream error events can reuse
them.

The chat backends wrap the error with %w instead of flattening it through
Detail(), so errors.As reaches the APIError on both the generate and the
stream path. utils.APIError, parseAPIError and Detail() are removed.

The conformance suite now checks, on both paths, that a provider's error
reply surfaces as the APIError it describes, with the reply's body and
headers, and that the error text contains neither the credential nor the
body. Error fixtures use the providers' documented formats.
Behaviour changes:

- In-stream error events are returned as *sdk.APIError with StatusCode 0,
  the event data as Body, and Type, Code, Message, RequestID and Kind
  decoded from the event: the Anthropic "error" event, the Responses and
  Codex "error" event (top-level or nested error object) and
  "response.failed". Codex now handles response.failed.
- A failed stream carries at most one ErrorPart, followed by a FinishPart
  with FinishReasonError and the usage reported so far. Malformed chunks no
  longer produce two ErrorParts (Anthropic, Completions, Copilot, Google).
- A stream that ends cleanly before its terminal event reports an
  ErrorPart wrapping the new sdk.ErrStreamIncomplete instead of a
  successful FinishPart. Terminal events: message_stop (Anthropic),
  response.completed or response.incomplete (Responses, Codex), a
  finish_reason or [DONE] (Completions, Copilot), a candidate finishReason
  or promptFeedback.blockReason (Google).
- A Responses non-stream 2xx body with an error object and an Alibaba Cloud
  images business code in a 2xx body are returned as *sdk.APIError.
- The DashScope decoder reads the code and message of a failed task's
  output.
- providertest checks in-band errors, incomplete and malformed streams, and
  RunErrorCases accepts Status 0 cases.
… test errors by field

A non-2xx body is read up to 1 MiB and a further 64 KiB is drained so the
connection stays reusable. The Copilot integration helpers match
*sdk.APIError fields instead of Error() text.
…eam errors

Decoder tables gain the Google reason and status branches, OpenRouter's
typed kinds, OpenAI's type-only classification and non-JSON bodies. The
Gemini API key and model-not-found bodies are verbatim from
generativelanguage.googleapis.com on 2026-09-29. Codex streams an error
event and a response.failed that reports usage; images decodes errors
from the JSON edit request.
…lues only

API_KEY_MISSING, QUOTA_EXCEEDED and RESOURCE_EXHAUSTED are not ErrorInfo
reasons of the googleapis.com domain. Map the enum's other credential and
access reasons instead; unlisted reasons still fall back to the canonical
status.
internal/errorformat and the Anthropic decoder list every value of the
providers' published error enums (OpenAI ResponseErrorCode, google.rpc.Code,
google.api.ErrorReason, OpenRouter ApiErrorType, Anthropic ErrorType) with
its Kind, and decode the specs' error examples verbatim.

internal/sdkdiff is a separate module that serves each error response to
both twilight and the official Go SDK (openai-go, anthropic-sdk-go,
go-genai) and requires both to read the same status, type, code, message
and request ID. CI lints and runs it.
OpenRouter sends x-generation-id on chat completion errors and exposes it through Access-Control-Expose-Headers. DecodeOpenRouter left RequestID empty, so callers had no ID to give OpenRouter support.
The status fallback reported every 5xx as server_error, which callers
retry. 501 Not Implemented and 505 HTTP Version Not Supported mean the
server does not support the request, and retrying does not change that.
They are now unknown; 524, 529 and the other 5xx stay server_error.
Gemini answers 429 RESOURCE_EXHAUSTED for per-minute and per-day quotas
alike, so a used-up daily quota was rate_limited, and callers retried it
after a RetryInfo delay of seconds. DecodeGoogle now reads the quota IDs
of the google.rpc.QuotaFailure detail and reports quota_exhausted when
one of the violated quotas is a daily one. The KindQuotaExhausted
documentation now includes daily quotas.
@HoneyBBQ
HoneyBBQ force-pushed the feat/errors-apierror branch from 0c00e36 to 546c207 Compare September 30, 2026 06:53
@HoneyBBQ
HoneyBBQ marked this pull request as ready for review September 30, 2026 06:54
@HoneyBBQ
HoneyBBQ merged commit d0ae284 into main Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant