Skip to content

Wire AlsaAudioSink into SinkRecovery so a replugged device recovers - #46

Merged
chrisuthe merged 2 commits into
mainfrom
chrisuthe/task/wire-alsaaudiosink-into-sinkrecovery-so-a
Aug 30, 2026
Merged

chrisuthe merged 2 commits into
mainfrom
chrisuthe/task/wire-alsaaudiosink-into-sinkrecovery-so-a

Conversation

@chrisuthe

Copy link
Copy Markdown
Member

Fixes #45. Reported downstream as music-assistant/local-audio-addon#28.

The defect

AlsaAudioSink was the only sink with a real device not wired into SinkRecovery. recover_() handled -EINTR/-EAGAIN, -EPIPE and -ESTRPIPE, and everything else -- -ENODEV among them -- fell through to a bare cli_log(ERROR, "alsa: %s", ...) and return false, leaving pcm_ open and failed_ clear.

All three callers are in write()'s inner loop and treat false as break, so write() returned 0, the sync task re-presented the same buffer, and the unthrottled ERROR repeated once per retry. write()'s pcm_ == nullptr discard path -- the code that would have coped -- was unreachable, because the handle outlived the hardware. Nothing before the next stream retired it, so only a restart cleared it.

The fix

Every error the three transient branches cannot clear in place is now device loss. handle_device_loss_() closes the device and spends SinkRecovery's inline attempt without making it, escalating to a new poll() override that reopens at last_format_ behind the existing 2 s -> 30 s delay and five-attempt budget. No new policy and no new constants.

Two calls the issue deliberately left to this PR:

1. The inline attempt is spent rather than made. snd_pcm_open() cannot be bounded the way PULSE_RECOVERY_TIMEOUT_MS bounds a Pulse reconnect, and on a plugin PCM -- default on a PipeWire host, or the alsa:pulse route -- it parses config and waits on a daemon socket with no timeout at all. Making that call on the sync task's thread, under device_mutex_, would break the rule the sink's threading model rests on: that write() never blocks unboundedly while holding it. SND_PCM_NONBLOCK is not the way out -- it changes the opened stream's semantics, which write()'s snd_pcm_wait() model depends on, and bounds neither the config parse nor a plugin's connect. Mechanically this is PulseAudioSink::reopen_in_place_()'s context-down branch, for the same reason.

2. Device loss is the residual, not a list of errnos. The issue asked which set counts; the answer here is that an allowlist is the wrong shape. The two mistakes are not the same size -- an errno wrongly left out restores the forever-spin for anything we guessed wrong about, while one wrongly taken in costs a close and a couple of seconds of discard before poll() reopens a device that was there all along, bounded by SINK_RESCAN_ATTEMPTS. This supersedes the -ENODEV/-ENXIO/-EIO question rather than answering it.

That inversion also closes a second door an allowlist would have left open: the failed prepare() after an underrun and the failed prepare() after a suspend both used to return false with the handle open and nothing armed. Both are reachable, and the first is how this very bug presents -- a device pulled while the ring drains is seen as -EPIPE first, with the -ENODEV only surfacing from the prepare() after it. All three failure points now route through the one helper.

Consequences, not scope

Named explicitly so the diff is legible:

  • configure() sets last_format_ before anything can fail and calls reset() at both success exits, never in open_device_() -- poll() calls that between rescan_due() and rescan_done(), where a reset() would refill the budget from inside the attempt spending it. Restructuring the fast path so both exits share a tail means a fast-path reopen failure now sets failed_, which it did not before.
  • write()'s discard path takes its frame size from last_format_ once close_device_() has zeroed bytes_per_frame_, matching pulse_sink.cpp. A sub-frame buffer therefore now returns 0 rather than length -- unreachable in practice, since the player hands over whole frames.
  • The same-format fast path needed no failure record to consult: a lost device leaves pcm_ == nullptr, so the fast-path condition is simply false. A null pcm_ is the record.

Behaviour worth knowing

The budget is per configured stream. A device that dies, recovers via poll(), then dies again in the same stream gets nothing until the next configure() -- rescan_done(true) latches, and reset() only runs on a stream that really opened. That is SinkRecovery working as designed and matches PulseAudioSink, but "recovered once, then silent" looks like a bug if you do not know to expect it.

Testing

Green, with an honest limit. Builds with zero warnings and 412/412 ctest pass on Debian bookworm with libasound2-dev (backends null, stdout, alsa) -- ALSA does not compile on macOS, so this was run in a container.

No new tests, and deliberately: the escalation sequence handle_device_loss_() performs is exactly tests/sink_recovery_test.cpp's escalate() helper, already covered by AFailedReopenEscalatesToTheRescan and EveryFurtherWriteOfTheOutageIsToldToDiscard. That is the split SinkRecovery is device-free for. The ALSA wiring itself is not reachable without hardware.

Nobody has yet pulled a DAC out of a running player and watched it come back. This is reasoned from the code and from the reporter's logs, not measured. docs/ROADMAP.md records it as an owed four-case hardware pass rather than claiming otherwise.

(If you re-run the suite as root, three StateStore tests fail -- root ignores the directory permission they rely on. As a non-root user all 412 pass.)

Docs

docs/ROADMAP.md item 14 claimed AlsaAudioSink "has nothing to do with a main-loop tick" -- the assumption that let ALSA out of that item, and so the root cause. Corrected in place rather than split into a new item, and positioned after item 14's PortAudio-only "What remains after this" paragraph, which is the opposite of true for ALSA. Item 20's caller list went from three to four; item 2's ALSA bullet cross-references item 14.

Out of scope

Ask 3 on the issue -- a sink-health field on format_status() -- is left alone. It is a separate change with a separate justification and applies to all four sinks.

AlsaAudioSink was the only sink with a real device not wired into
SinkRecovery. recover_() handled -EINTR/-EAGAIN, -EPIPE and -ESTRPIPE and
treated everything else -- -ENODEV among them -- as a failed write: it
logged, returned false, and left pcm_ open on hardware that was gone.
Nothing before the next stream retired the handle, so write() returned 0,
the sync task re-presented the same buffer, and the unthrottled ERROR
repeated once per retry until the process restarted.

Every error the three transient branches cannot clear in place is now
device loss. handle_device_loss_() closes the device and spends
SinkRecovery's inline attempt without making it, which escalates to a new
poll() override that reopens at last_format_ behind the existing delay and
budget. Taking device loss as the residual rather than as a list of errnos
is deliberate: an errno wrongly left out restores the forever-spin, while
one wrongly taken in costs a close and a couple of seconds of discard
before poll() reopens a device that was there all along.

The inline attempt is spent rather than made because snd_pcm_open() cannot
be bounded the way PULSE_RECOVERY_TIMEOUT_MS bounds a Pulse reconnect, and
on a plugin PCM it parses config and waits on a daemon socket with no
timeout at all. Making that call on the sync task's thread, under
device_mutex_, would break the rule the sink's threading model rests on.

The failed prepare() following an underrun and the failed prepare()
following a suspend route to the same helper. Both are reachable and both
stranded the handle: a device pulled while the ring drains is seen as
-EPIPE first, with the -ENODEV only surfacing from the prepare() after it.

Also here, as consequences rather than scope:

- configure() sets last_format_ before anything can fail and calls reset()
  at both success exits, never in open_device_() -- poll() calls that
  between rescan_due() and rescan_done(), where a reset() would refill the
  budget from inside the attempt spending it. Restructuring the fast path
  so both exits share a tail means a fast-path reopen failure now sets
  failed_, which it did not before.
- write()'s discard path takes its frame size from last_format_ once
  close_device_() has zeroed bytes_per_frame_, so a sub-frame buffer now
  returns 0 rather than length.

No new policy and no new constants. The escalation sequence is exactly
tests/sink_recovery_test.cpp's escalate() helper, already covered there;
the ALSA wiring itself needs hardware, and is recorded in docs/ROADMAP.md
as an owed hardware pass rather than claimed as verified.

Closes #45
@chrisuthe
chrisuthe marked this pull request as ready for review August 30, 2026 16:31
@chrisuthe
chrisuthe requested a balanced review from Copilot August 30, 2026 16:31

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Recovery performs an unbounded open while holding a mutex needed by writes, potentially blocking playback and the main loop indefinitely.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Integrates ALSA device-loss recovery with the shared SinkRecovery mechanism.

Changes:

  • Detects device loss and schedules delayed reopen attempts.
  • Preserves stream format for recovery.
  • Updates recovery documentation and roadmap.
File summaries
File Description
src/alsa_sink.h Declares ALSA recovery state and polling.
src/alsa_sink.cpp Implements device-loss handling and reopening.
docs/ROADMAP.md Documents ALSA recovery behavior.
Review details
  • Files reviewed: 3/3 changed files
  • Comments generated: 3
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/alsa_sink.cpp
Comment thread docs/ROADMAP.md Outdated
Comment thread src/alsa_sink.cpp
The item 14 correction said a device pulled out mid-track "discarded audio
and one ERROR per buffer". It discarded nothing: write() broke out of its
loop and returned 0, so the sync task re-presented the same buffer against
the dead handle and made no progress at all. Discarding is what the sink
does now, and the distinction is the whole of the defect -- so getting it
backwards in the one paragraph explaining it was worth fixing.

The hardware pass likewise promised one ERROR for a powered-off DAC. An
outage logs twice: recover_() names the device and the errno, and write()'s
discard branch says what happens next, each latched once per outage. The
property that matters is that neither repeats per buffer.

Also records what holding device_mutex_ across poll()'s snd_pcm_open()
costs -- a real pause in protocol handling, and a concurrent write()
blocked on the mutex rather than honouring its timeout_ms -- along with why
that is not a new risk: configure() already makes the same unbounded open
on the main loop under the same mutex, on every stream rather than only on
a lost device.
@chrisuthe

Copy link
Copy Markdown
Member Author

Went through Copilot's three. One was a real error in my prose and is fixed; the other two are accurate observations whose proposed remedies I've declined, with reasoning below so it's on the record rather than silently dropped.

Accepted — the ROADMAP misdescribed the old behaviour (docs/ROADMAP.md:1820)

Correct, and wrong in exactly the way this issue turns on. I'd written that a device pulled mid-track "discarded audio and one ERROR per buffer". It discarded nothing: write() broke out of its loop and returned 0, so the sync task re-presented the same buffer against a dead handle and made no progress at all. Discarding is what the sink does now. Fixed in bb1186d.

Declined — latching failed_ in handle_device_loss_() (src/alsa_sink.cpp:520)

The observation is right: an outage logs twice, not once. recover_() names the device and the errno, then write()'s discard branch says what happens next.

But failed_ is only a log-once latch — nothing in the sink reads it for control flow (alsa_sink.cpp:434, :575, :600 are its only uses). Latching it here wouldn't change behaviour, it would just delete the second message, and that message is a different fact from the first: what happened versus what the sink is doing about it now. An operator would otherwise see "is gone" and then silence, with nothing saying audio is being dropped. This is the split PulseAudioSink already makes — open_stream_() says why once at ERROR, report_failed_recovery_() says what follows.

Both lines are latched, so the property that actually matters holds: neither repeats per buffer, which was the reported symptom. What was wrong was my hardware-pass note promising one ERROR; that now says two and why. Fixed in bb1186d, doc side only.

Declined — moving snd_pcm_open() off device_mutex_ onto a recovery worker (src/alsa_sink.cpp:800)

The property is real and I've documented it rather than argue it away. Two corrections to the framing, though, and then the scope call.

poll() and stop() both run on the main loop (audio_sink.h:169, main.cpp:991 and :1036), so shutdown can't be blocked by this in the sense of a stall it can't get out of — the loop is simply busy, exactly as it is elsewhere. The write() half is accurate: a concurrent write() blocks on the mutex rather than honouring timeout_ms.

It isn't a new risk. configure() already makes an identical unbounded snd_pcm_open() on the main loop under the same mutex (alsa_sink.cpp:536 + :571) — on every stream, not just on a lost device. poll() adds at most SINK_RESCAN_ATTEMPTS more per configured stream, behind a doubling delay.

The case that actually stalls is narrow. An absent hw: PCM fails fast, which is the scenario in #45 — paying real time needs a plugin PCM waiting on a daemon socket, or an exclusive device another process has taken in the gap.

The remedy is a background thread, plus restructuring open_device_() to build into a local handle and publish under the lock. This sink is deliberately built without one — a single mutex is its entire threading model, and the ROADMAP treats every no-background-thread argument here as a threading argument. That's a large, risk-bearing change to make speculatively on a fix PR, against a hazard that already exists on the shipped path.

So I've named the cost in item 14 instead — the same way that item already does for PortAudio's rescan — and the hardware pass includes measuring how long snd_pcm_open() really stalls the loop on an absent hw: device. If that measurement comes back bad, it's the thing that justifies revisiting this, and it should be its own change rather than riding along here.

No code changed for any of the three; bb1186d is docs only. Build still clean with zero warnings in our sources, 412/412 green.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

AlsaAudioSink is not wired into SinkRecovery, so a replugged device never comes back

2 participants