Repository navigation
feat(sdk)!: 导出 APIError,HTTP 错误和流内错误按 Kind 跨 provider 分类 - #64
Merged
Merged
Conversation
HoneyBBQ
force-pushed
the
feat/errors-apierror
branch
from
September 28, 2026 12:31
e8b1681 to
4970d9d
Compare
HoneyBBQ
marked this pull request as ready for review
September 28, 2026 12:45
HoneyBBQ
added this pull request to stack #67
September 28, 2026 13:17
HoneyBBQ
force-pushed
the
feat/errors-apierror
branch
from
September 28, 2026 19:55
5e73fb2 to
fc05db5
Compare
HoneyBBQ
marked this pull request as draft
September 28, 2026 20:18
HoneyBBQ
force-pushed
the
feat/errors-apierror
branch
4 times, most recently
from
September 29, 2026 22:11
67e0352 to
0c00e36
Compare
Provider failures reached callers only as text: the HTTP layer's error type was internal, so status, provider codes and the response body could not be recovered with errors.As, and downstream code matched error strings. APIError is the exported type for a failure the provider reported. It keeps the provider's type, code, message and request ID verbatim, the response headers and raw body, and a cross-provider Kind. StatusCode is 0 when the HTTP status line did not report the failure. Error() carries provider, status, type/code, message and request ID, never the body or any header, because the body may echo end-user input. RetryAfter parses retry-after-ms, delta-seconds and HTTP-date. ErrorKind is an open set; KindOf reads it from an error chain.
FetchJSON, FetchRaw and FetchSSE now build the error with NewHTTPError: the status gives a fallback Kind, and a per-provider ErrorDecoder fills type, code, message and request ID from the provider's own error body and headers and refines Kind from its type or code, which win over the status (OpenAI's 429 insufficient_quota is a quota failure, not a rate limit). Nothing is classified from message text. Each package passes its name and decoder through RequestOptions. OpenAI's format, shared by the OpenAI-compatible packages, and Google's google.rpc.Status live in internal/errorformat; Anthropic, Codex, Copilot, DashScope, Ark and OpenRouter decode their own. The type/code-to-Kind functions are separate from the decoders so stream error events can reuse them. The chat backends wrap the error with %w instead of flattening it through Detail(), so errors.As reaches the APIError on both the generate and the stream path. utils.APIError, parseAPIError and Detail() are removed. The conformance suite now checks, on both paths, that a provider's error reply surfaces as the APIError it describes, with the reply's body and headers, and that the error text contains neither the credential nor the body. Error fixtures use the providers' documented formats.
Behaviour changes: - In-stream error events are returned as *sdk.APIError with StatusCode 0, the event data as Body, and Type, Code, Message, RequestID and Kind decoded from the event: the Anthropic "error" event, the Responses and Codex "error" event (top-level or nested error object) and "response.failed". Codex now handles response.failed. - A failed stream carries at most one ErrorPart, followed by a FinishPart with FinishReasonError and the usage reported so far. Malformed chunks no longer produce two ErrorParts (Anthropic, Completions, Copilot, Google). - A stream that ends cleanly before its terminal event reports an ErrorPart wrapping the new sdk.ErrStreamIncomplete instead of a successful FinishPart. Terminal events: message_stop (Anthropic), response.completed or response.incomplete (Responses, Codex), a finish_reason or [DONE] (Completions, Copilot), a candidate finishReason or promptFeedback.blockReason (Google). - A Responses non-stream 2xx body with an error object and an Alibaba Cloud images business code in a 2xx body are returned as *sdk.APIError. - The DashScope decoder reads the code and message of a failed task's output. - providertest checks in-band errors, incomplete and malformed streams, and RunErrorCases accepts Status 0 cases.
… test errors by field A non-2xx body is read up to 1 MiB and a further 64 KiB is drained so the connection stays reusable. The Copilot integration helpers match *sdk.APIError fields instead of Error() text.
…eam errors Decoder tables gain the Google reason and status branches, OpenRouter's typed kinds, OpenAI's type-only classification and non-JSON bodies. The Gemini API key and model-not-found bodies are verbatim from generativelanguage.googleapis.com on 2026-09-29. Codex streams an error event and a response.failed that reports usage; images decodes errors from the JSON edit request.
…lues only API_KEY_MISSING, QUOTA_EXCEEDED and RESOURCE_EXHAUSTED are not ErrorInfo reasons of the googleapis.com domain. Map the enum's other credential and access reasons instead; unlisted reasons still fall back to the canonical status.
internal/errorformat and the Anthropic decoder list every value of the providers' published error enums (OpenAI ResponseErrorCode, google.rpc.Code, google.api.ErrorReason, OpenRouter ApiErrorType, Anthropic ErrorType) with its Kind, and decode the specs' error examples verbatim. internal/sdkdiff is a separate module that serves each error response to both twilight and the official Go SDK (openai-go, anthropic-sdk-go, go-genai) and requires both to read the same status, type, code, message and request ID. CI lints and runs it.
OpenRouter sends x-generation-id on chat completion errors and exposes it through Access-Control-Expose-Headers. DecodeOpenRouter left RequestID empty, so callers had no ID to give OpenRouter support.
The status fallback reported every 5xx as server_error, which callers retry. 501 Not Implemented and 505 HTTP Version Not Supported mean the server does not support the request, and retrying does not change that. They are now unknown; 524, 529 and the other 5xx stay server_error.
Gemini answers 429 RESOURCE_EXHAUSTED for per-minute and per-day quotas alike, so a used-up daily quota was rate_limited, and callers retried it after a RetryInfo delay of seconds. DecodeGoogle now reads the quota IDs of the google.rpc.QuotaFailure detail and reports quota_exhausted when one of the violated quotas is a daily one. The KindQuotaExhausted documentation now includes daily quotas.
HoneyBBQ
force-pushed
the
feat/errors-apierror
branch
from
September 30, 2026 06:53
0c00e36 to
546c207
Compare
HoneyBBQ
marked this pull request as ready for review
September 30, 2026 06:54
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
问题
main 上的
sdk.APIError(#63)只有状态码、状态文本、通用提取的 message、request ID 和原始 body,没有失败原因的分类。调用方要区分 key 无效、额度耗尽、provider 过载,仍然只能看状态码和文本,而同一个状态码在各 provider 下含义不同:OpenAI 返回 429 时可能是限流,也可能是额度耗尽。它的Error()还附带最多 4 KiB 的响应体,响应体可能回显终端用户的输入,打印错误就会写进日志。流也有同样的问题。Anthropic 的
error事件读错了字段,流中途的 overloaded 到调用方只剩unknown error。流在终止事件之前断开,会被当作成功结束。改动
APIErrorto sdk #63(9b039e1)。NewAPIError、NewAPIErrorFromResponse、IsStatus和ClassifyProbeError随之删除,由本 PR 的类型取代。sdk.APIError,一律以指针使用。字段有Provider、StatusCode,provider 自己的Type、Code、Message,以及RequestID、Kind、响应的Header和原始Body。Error()包含 provider、状态码、type/code、message 和 request ID。它不包含 body 和任何响应头,因为 body 可能回显终端用户的输入。Provider是发出请求的 provider 的Name()。OpenCode Go 经 Completions、Responses 或 Messages 发请求,错误里是这三者之一的名字。RetryAfter()先读retry-after-ms,再读Retry-After,支持秒数和 HTTP-date 两种写法。sdk.ErrorKind跨 provider 分类失败原因,取值有authentication、permission_denied、quota_exhausted、rate_limited、server_error、unknown。这是开放集合,KindOf(err)可以从任意错误链中读出 Kind。各 provider 先按自己的错误 type 或 code 映射,都匹配不上时再看 HTTP 状态码:401、402、403、429,以及 501、505 以外的 5xx。provider 的 code 优先于状态码,所以 OpenAI 返回 429 且 code 为insufficient_quota时,Kind 是quota_exhausted。分类从不依据 message 文本。FetchJSON、FetchRaw和FetchSSE遇到非 2xx 响应时返回*sdk.APIError,按 provider 使用的错误格式解码。多个包都会收到的格式在internal/errorformat中解码:OpenAI 格式(completions、responses、embedding、images,以及经这些包访问的 Bedrock)、Google 的google.rpc.Status(generativeai、embedding、transcription)、OpenRouter 和阿里云 DashScope。Anthropic、Codex、Copilot 和方舟视频的解码器留在各自的包里。Google 先按 ErrorInfo 的 reason 分类,只认google.api.ErrorReason中的值,没有映射时再看 canonical status。Gemini 的每分钟配额和每日配额用尽都返回 429RESOURCE_EXHAUSTED,所以QuotaFailure中有一条违反的是每日配额时,Kind 为quota_exhausted。每日配额按 quotaId 含PerDay或Daily判断,与 Gemini CLI 相同。OpenRouter 的RequestID取自x-generation-id响应头。API 参考没有写这个头,但聊天补全的错误响应带有它,Access-Control-Expose-Headers也列出了它。resp.Body关闭。%w包装错误,不再格式化成字符串。删除utils.APIError、parseAPIError和Detail()。*sdk.APIError的StatusCode和Code判断 403 和model_not_supported,不再匹配错误文本。*sdk.APIError,StatusCode为 0,Body是事件数据。覆盖 Anthropic 的error,以及 Responses 和 Codex 的error与response.failed。Anthropic 事件按官方文档读error.type和error.message。Responses 的error事件两种结构都接受:code 和 message 在顶层(API 参考的写法),或嵌套在error对象里(Codex CLI fixture 的写法)。ErrorPart,随后是一个FinishPart,结束原因为FinishReasonError,带上已上报的 usage。此前 Anthropic、Completions、Copilot 和 Google 收到损坏的数据块时会发两个ErrorPart。sdk.ErrStreamIncomplete的ErrorPart。各后端的终止事件是:Anthropic 为message_stop;Responses 和 Codex 为response.completed或response.incomplete;Completions 和 Copilot 为[DONE]或带finish_reason的数据块;Google 为候选的finishReason或promptFeedback.blockReason。StatusCode为 0 的*sdk.APIError。internal/sdkdiff,官方 SDK 只是它的测试依赖,主模块的依赖不变。CI 对它单独运行 lint 和测试。本 PR 不处理直接调用
http.Client的路径(语音、转录、multipart 图片、视频下载、WebSocket)和Provider.Test,它们仍返回无类型的错误。验证
go build ./...、go vet ./...、go vet -tags integration ./...、go test ./... -short -count=1 -race和golangci-lint run ./...均通过,go mod tidy无 diff。errors.As、StatusCode、Type、Code、Message、RequestID、Kind,错误上保留的 body 和响应头,以及Error()不含 API key 和 body。故意写错期望的Kind,套件会失败。TestModel都返回该 delegate 的*sdk.APIError,StatusCode和Kind正确。internal/errorformat的每个分支都有用例。embedding、images(生成与 JSON 编辑)、videos 和 Google 转录有端到端的httptest用例。Codex 流内的error事件和带 usage 的response.failed经过完整的流路径测试,后者断言FinishPart保留了 usage。authentication,模型不存在为unknown。缺少 key 的响应是 403PERMISSION_DENIED,没有 ErrorInfo,所以是permission_denied。quota_exhausted。只违反每分钟配额是rate_limited。authentication,模型不存在为 400unknown。后者的RequestID是x-generation-id的值,这条错误体和响应头原文写入测试。InvalidApiKey,Kind 是authentication。尺寸越界在任务结果里报告,StatusCode为 0,Code为InvalidParameter。response.failed都返回*sdk.APIError,状态码和 Kind 正确。中转网关的错误体不是官方格式,没有用作 fixture。ResponseErrorCode、google.rpc.Code、google.api.ErrorReason、OpenRouterApiErrorType和 AnthropicErrorType。枚举取自固定提交的 OpenAPI 或 proto 文件。规格中的错误示例原样解码,OpenRouter 每个状态码的示例都在内。RequestID为空。OpenRouter 的视频接口和 401 响应不带x-generation-id,RequestID同样为空。Copilot 的错误格式没有公开,fixture 依照 VS Code Copilot 客户端的解析逻辑编写。阿里云 DashScope 没有机器可读的公开契约,只按文档测试。error事件结构,两种结构都有测试。Gemini 流中途的错误对象不做解码。