Audio capture
What captures the audio on your machine, and the one question about it that this project answered wrong for four days.
Distro support
Section titled “Distro support”Audio capture works two ways, picked automatically:
- PipeWire (
pw-record), preferred, and the only backend that can tap an individual application’s playback stream, which is what--targetis for. - PulseAudio (
parec), the fallback. Everything works except per-application capture;--list-targetsthere lists monitor sources and says why.
pactl is used for the default sink and mute state where present (on PipeWire
too, via pipewire-pulse); without it, PipeWire’s own default.audio.sink
metadata answers the same question. On a machine running both stacks where the
automatic pick is wrong, VINOWHISPER_CAPTURE_BACKEND=pipewire or
=pulseaudio forces it.
Package names per family live in one table in
vinowhisper/distro.py, so vinowhisper-doctor and
the wizard both speak your distro:
| Family | Covered | Confidence |
|---|---|---|
| Fedora / RHEL / Alma / Rocky / Nobara / Bazzite / Silverblue | ✓ | Built and run here |
| Debian / Ubuntu / Pop / Mint / elementary / Raspbian | ✓ | From the package index, not from use |
| Arch / CachyOS / EndeavourOS / Manjaro / Garuda | ✓ | Same |
| openSUSE / SLES | ✓ | Same |
| Void, Gentoo, Alpine | ✓ | Same |
| NixOS | ✓ | Configuration advice, not nix-env lines: imperative installs there do not persist |
| Anything else | degrades | Generic advice, and it says so |
If a package name is wrong for your distro, that is expected, and it is the fastest thing here to fix. See CONTRIBUTING.md or the distro-support issue template.
That is also why the wizard and the doctor print these commands rather than
run them unasked. They carry -y or --needed so they work exactly as
printed, and every NPU entry points at the upstream release as the
authoritative fallback. The family is matched on ID in /etc/os-release
first, and on ID_LIKE only when the ID is unknown, since ID_LIKE is often
missing or unhelpful.
“Captions stop when you mute the system”
Section titled ““Captions stop when you mute the system””They don’t, and this section used to say the opposite at length. Measured with
vinowhisper-doctor on 2026-08-07, always against a Chrome stream playing the
same audio as a control:
| Sink state | Default sink monitor | Chrome stream | Monitor / app |
|---|---|---|---|
| 100% volume | 0.01419 | 0.02040 | 0.70 |
| 20% volume | 0.05739 | 0.06610 | 0.87 |
| Muted | 0.08578 | 0.08781 | 0.98 |
The ratio is the measurement, not the absolute levels, since the content
differed between runs. It doesn’t move with the volume slider and it doesn’t
move on mute. The sink monitor is both pre-volume and pre-mute here.
pw-dump agrees on the volume half: monitor.channel-volumes is unset, and
PipeWire defaults it to false.
So the original diagnosis was wrong, and so was every mitigation built on it. Lowering the volume costs nothing. Muting the system costs nothing. The earlier threshold-lowering work was chasing a mechanism that isn’t there.
What actually silences the capture
Section titled “What actually silences the capture”- Muting the application rather than the system. YouTube’s own mute
button, or Chrome’s slider in the KDE mixer. The app then writes silence
into its own PipeWire stream, and there is no tap upstream of that.
--targetdoes not help, because the stream it would capture is the silence. Note the node stays listed in--list-targetsthe whole time, because that lists nodes that exist, not nodes carrying signal. This is the likeliest explanation for the original report. --target effect_output.bass_eq. It shows up in--list-targetsas the most obvious-looking choice and is the worst one: it sits on the output side of the EQ chain, downstream of both volume and mute. It measured 0.00034 at 20% volume and 0.00000 while muted, in the same runs where the sink monitor read 0.05739 and 0.08578. It is the one node here that genuinely is post-everything. Aim at a real application instead.- Nothing playing.
vinowhisper-doctorsays so explicitly when every target reads zero.
Quiet-but-not-silent audio is handled separately: every window is boosted
toward a speech-like level (config.TARGET_RMS, up to 20x) before it reaches
the model, since Whisper’s accuracy degrades on quiet input. Ordinary web
video lands around 0.014 rms, so this earns its place on source material
alone, independent of the volume question above.
Silence
Section titled “Silence”True digital silence on a sink monitor measures about 0.0 to 0.004 rms, from
dither and the EQ chain’s noise. The 0.002 gate (config.SILENCE_RMS_THRESHOLD)
only saves NPU cycles on dead air. It is not the defence against
hallucinations; that is the two-cycle commit policy plus collapse_repeats.
Raise it and quiet speech, which really does sit near the noise floor, gets
dropped.
Some silence is normal, so the terminal only prints its “no signal” notice after 45 seconds of it. That much, while someone expects captions, is a symptom, and the notice points at the causes above.