Skip to content

feat(backend): add websocket voice streaming pipeline - #116

Merged
LUPENGHAN merged 3 commits into
1024XEngineer:MVPfrom
yyy-router:feature/backend-ws-voice-stream
Jul 30, 2026
Merged

LUPENGHAN merged 3 commits into
1024XEngineer:MVPfrom
yyy-router:feature/backend-ws-voice-stream

Conversation

@yyy-router

Copy link
Copy Markdown
Contributor

概述

实现后端 WebSocket 实时语音流处理闭环。客户端通过 Binary Frame 持续发送麦克风音频,后端实时转发给阿里云 ASR,并在识别完成后调用 LLM 生成前端可确认的结构化日程草稿。

主要变更

  • 扩展 WebSocket 端点以区分 JSON Text Frame 和 Binary Frame。
  • 新增音频流开始、二进制分片接收、结束确认和异步结果推送处理。
  • 使用有界队列限制内存占用,并支持最大录音时长校验。
  • 阿里云 ASR 默认使用 server_vad,并发发送音频和接收识别事件。
  • 聚合 VAD 多段最终文本,同时保留 Manual 模式兼容能力。
  • 将 ASR 和 LLM 异常映射为可区分处理阶段的统一失败结果。
  • 缩短已完成 ASR 会话的 WebSocket 关闭等待,移除额外 10 秒延迟。
  • 在应用组装入口注册语音处理服务,并在断连时清理后台任务。

验证

  • 后端完整测试:100 个通过。
  • Ruff 检查通过。
  • 本次改动文件 Mypy 检查通过。
  • 已使用浏览器麦克风完成真实 ASR + LLM 端到端验证。

Closes #115

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found in this review.

@LUPENGHAN
LUPENGHAN merged commit 1855b7b into 1024XEngineer:MVP Jul 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants