Installing
The short version is in the README. This is what the installer actually does, and the dependency pin behind it.
curl -fsSL https://raw.githubusercontent.com/karanshukla/vinoWhisper/main/scripts/install.sh | bashThat installs uv (a pinned version, only if you
have none), clones the latest vX.Y.Z release tag rather than main
(--ref main for unreleased code, --ref v0.6.1 for a specific one), builds
the environment with uv sync --locked, so exactly what uv.lock pins and
never a fresh resolve, and hands over to vinowhisper-setup, which is where every
machine-specific decision happens: your capture tool, your NPU driver, the
model export your device needs, and systemd units generated against the paths
that actually exist. It prints every command before running it and asks first.
--yes answers yes to everything and --dry-run no to everything, and piped
stdin without --yes refuses rather than run sudo commands nobody agreed to.
A step that is already done says so and is skipped. It never invents a package
command where a distro has none (the NPU userspace driver, on most), and points
at Intel’s releases instead. The ~/.local/bin entries are symlinks into the
checkout, so they follow a git pull, and bash completion gets one link per
command, because bash-completion loads a file lazily on the first Tab for a
command of the same name.
The whole script is one main function called on its last line, so a
download cut short by the network is a syntax error that runs nothing, rather
than half an install. The release tag comes from git ls-remote on the
repository, not the GitHub API, so there is no rate limit and no JSON to
parse. Re-running it on an existing checkout moves that checkout to the
requested ref (detached), so a checkout installed from main goes back to the
latest release unless you pass --ref main again.
From a checkout, or to see what it would do without doing it:
git clone https://github.com/karanshukla/vinoWhisper && cd vinoWhisperuv sync --extra export # --extra export: the one-time model exportuv run vinowhisper-setup --dry-run # the whole plan, nothing changeduv run vinowhisper-setup # for real, one prompt per stepFrom PyPI
Section titled “From PyPI”pip install vinowhisperThis works, and until 2026-08-31 it could not. The dependency set resolved
openvino from PyPI, where the builds that can construct the NPU static
Whisper pipeline did not exist, so a pip install would have succeeded and then
failed to load a model, which is worse than not shipping at all. Stable
2026.3.1 removed that constraint (see the version floor below).
What it gets you: the five commands and the Python dependencies. What it cannot
get you: an NPU driver, a model export, or systemd units. Run
vinowhisper-setup afterwards for those, exactly as the installer script would.
Re-run it after upgrading from 0.6.x or earlier too: the server moved from
127.0.0.1:8099 to a Unix socket in $XDG_RUNTIME_DIR, and the wizard is
what rewrites the socket unit and restarts it.
The export tooling (optimum, and torch with it) is the export extra, which
the wizard offers to install when it gets to the model step;
pip install 'vinowhisper[export]' does it up front.
Python must be 3.11-3.13. 3.14 made functools.partial a descriptor, which
breaks optimum’s NORMALIZED_CONFIG_CLASS = SomeConfig.with_args(...)
class-attribute idiom outright. Version-independent root cause, confirmed
2026-08-03 across every optimum/transformers pairing tried. requires-python
enforces it, so pip will refuse rather than install something broken.
The OpenVINO version floor
Section titled “The OpenVINO version floor”openvino>=2026.3.1, and the floor is exact rather than cautious: 2026.3.1 is
the first stable release that can build the NPU static Whisper pipeline.
From 2026-08-03 to 2026-08-31 this project pinned nightly wheels from
storage.openvinotoolkit.org, with prereleases allowed globally, because
stable 2026.2.1 could not build that pipeline at all: its
pipeline_static.cpp pattern-matcher did not recognise the optimum-intel
export’s SDPA attention-mask node shape, and
OPENVINO_ASSERT(!self_attn_nodes.empty()) failed. That pin was the project’s
standing dependency risk, since upstream prunes nightly builds on its own
schedule.
Re-measured 2026-08-31 on the Wildcat Lake NPU, against the same
--disable-stateful export, with WhisperPipeline(..., STATIC_PIPELINE=True):
| Version | Source | Pipeline build | generate() |
|---|---|---|---|
| 2026.3.1 | PyPI stable | 2.0s | ok |
| 2026.4.0.dev20260805 | nightly | 2.3s | ok |
| 2026.5.0.dev20260831 | nightly | 1.9s | ok |
Stable caught up, so the nightly index, the prerelease = "allow" policy and
the [tool.uv.sources] routing are all gone and these resolve from PyPI like
anything else. .github/workflows/deps-canary.yml still runs the full resolve
weekly, off the pull-request path, but now for ordinary upstream churn rather
than for wheels aging out from under the lock.
STATIC_PIPELINE=True is not optional and does not degrade. Omitting it
sends the NPU down the generic stateful path, which fails with Stateful models without 'beam_idx' input are not supported in StatefulToStateless transformation. That reads like a bad export and is not one.
The model export, and what verifies it
Section titled “The model export, and what verifies it”Exporting downloads ~1GB from Hugging Face and converts it to OpenVINO IR that
then runs on your hardware. The download comes first and is checked first:
vinowhisper/model_sources.json pins the repository to a commit and records
the sha256 of each file the export reads, python -m vinowhisper.source
fetches exactly those into models/source/ under the data directory, and any
difference fails before optimum-cli starts. The export then runs from that
directory with HF_HUB_OFFLINE=1. optimum-cli has no --revision flag
(optimum-intel 2.2.0), which is why the snapshot is local rather than in the
Hugging Face cache. A --model with no pinned source is exported unchecked,
with a warning.
vinowhisper/model_digests.json pins the sha256 of
every file in the export this project has actually run, and both
scripts/convert_model.sh and vinowhisper-setup check what came down against
it. vinowhisper-doctor re-checks it on demand, at about 1.2s for 1.5GB, and
only for an export that is present and the right shape for its device, since
“hashes don’t match” is noise next to “wrong export entirely”.
The pin is on the exported IR, not on the upstream safetensors, because the IR
is what WhisperPipeline loads and the export is not a pure function of the
weights. Six statuses, and only three of them stop anything:
| Status | Means | Blocks setup |
|---|---|---|
verified |
every pinned file matches | |
unpinned |
no pin for this model and variant | no |
drift |
bytes and export toolchain both moved | no |
mismatch |
the pinned toolchain produced different bytes | yes |
incomplete |
pinned files are missing, so the export is partial | yes |
known_bad |
exported by a toolchain measured to produce a broken export | yes |
An unpinned export warning is not a problem to fix. It is what
--model openai/whisper-base.en looks like, and what the stateful export looks
like until someone produces one. Re-pin with
./scripts/update_digests.py --variant npu once you trust an export, and
commit the diff.
The export is bit-reproducible, which is what makes any of this work.
Measured 2026-09-04: two independent optimum-cli exports of
whisper-small.en on the same toolchain produced all 16 files byte-identical.
Across toolchains it is not: against the 2026-08-03 export
(OpenVINO 2026.2.1, optimum-intel 2.0.0, transformers 5.0.0), an export under
OpenVINO 2026.3.1 / optimum-intel 2.1.0 / transformers 5.5.4 changed 9 of 16
files, including both decoder .bin weights. openvino_encoder_model.bin came
out identical across both. That is why drift is reported separately from a
real mismatch, and the versions are read out of the export’s own rt_info
block rather than from whatever happens to be installed. That makes drift
the export’s own claim about itself: something able to rewrite the export can
rewrite every rt_info block to match and get drift instead of mismatch.
The source check is what guards the download; this one catches accidents.
That block is read from every .xml in the export and merged, and any version
two files disagree on makes the answer unknown, never fine: reading just one
file would let a single edited graph buy the softer drift verdict (editing
all of them consistently still does). The
digest covers every file, not just the weights, because generation_config.json
decides how decoding behaves and tokenizer.json decides the text. A missing
or truncated pin file downgrades to unpinned rather than breaking every
export. Verification is deliberately not part of loading the model, since
hashing takes about 1.2s and would land on the socket-activated cold start.
known_bad entries come in two shapes: an exact combination measured broken,
or a floor (at_least, as for transformers 5.4.0). A version the export does
not report never satisfies a floor, and update_digests.py rewrites the
hashes without touching the list.
transformers 5.4.0 breaks the NPU export
Section titled “transformers 5.4.0 breaks the NPU export”Bisected on hardware 2026-09-04. Export with transformers<5.4. Anything
from 5.4.0 on produces a Whisper decoder the NPU static pipeline compiles and
then cannot run:
RuntimeError: Port for tensor name cache_position was not found. (src/inference/src/cpp/infer_request.cpp:191)It fails at generate(), not at load, so nothing complains until the first
transcription.
The bisect held optimum-intel 2.1.0, optimum 2.3.0, openvino 2026.3.1,
openvino-genai 2026.3.1.0 and torch 2.13.0 fixed, exported whisper-small.en
with --disable-stateful in a clean venv per version, and loaded each on the
NPU:
| transformers | Pipeline build | generate() |
|---|---|---|
| 5.0.0 | ok | 0.96s |
| 5.2.0 | ok | 0.86s |
| 5.3.0 | ok | 0.71s |
| 5.4.0 | ok | fails |
| 5.5.4 | ok | fails |
Two control runs rule out the rest of the stack. optimum-intel 2.1.0 with transformers 5.0.0 works, so the optimum pair is not at fault; the full 2026-08-03 package set re-run under openvino 2026.3.1 also works, so the runtime is not either.
The mechanism is a tensor name. cache_position appears exactly once in
each transformers 5.3.0 decoder graph, on the output port of
__module.model.model.decoder/aten::arange/Range, and zero times in the 5.4.0
graphs. It is neither a model input nor a model output in either export, so the
static pipeline is resolving an internal traced tensor by name, and 5.4.0
stopped emitting that name. The exported input and output signatures are
otherwise identical between the two.
Since 2026-09-12 the export’s dependencies are their own extra,
vinowhisper[export], which holds transformers<5.4. Before that a fresh
install resolved 5.5.4 and vinowhisper-setup failed at the model step. The
extra is separate because only the export uses optimum and transformers, and
optimum brings torch with it, while the runtime imports none of them.
vinowhisper-setup offers to install it when it needs to export, as
uv sync --extra export in a checkout and pip install 'vinowhisper[export]'
otherwise.
The pin’s known_bad entry still carries a floor at transformers 5.4.0, for
an export made some other way, so vinowhisper-setup and convert_model.sh
report a known_bad export rather than handing over a model that fails later.
The stateful (CPU/GPU) export is unaffected: it builds and decodes under 5.5.4
(measured 2026-09-12).
Raising the cap does not upgrade anything (2026-09-20). Every released
optimum-intel, 2.2.0 included, requires transformers<5.6,>=4.51. Ask for
anything past that and the resolver satisfies it by backtracking optimum to a
pre-transformers-5.x release instead, at which point optimum-cli fails on
import rather than at export time. Two open transformers advisories,
GHSA-fgcw-684q-jj6r (fixed in 5.5.0) and GHSA-xrqw-3rrv-vx5w (fixed in
5.10.0), are unreachable for that reason and the bisection above. Nothing at
runtime imports transformers; the only call site is the one-time
optimum-cli export openvino in scripts/convert_model.sh, against a
hardcoded openai/whisper-small.en unless you pass your own --model.
The desktop overlay (optional)
Section titled “The desktop overlay (optional)”vinowhisper-setup --gui installs vinowhisper-gui, a floating caption box
that also does dictation, with a tray icon and global shortcuts. It downloads the binary from the
GitHub release and checks it against the sha256 pinned in the Python package,
so no Rust toolchain is needed. It is never on PyPI; from a checkout it can be
built with cargo instead (./scripts/install.sh --gui). See gui.md.
NPU platform data (optional)
Section titled “NPU platform data (optional)”vinowhisper-setup --ovfetch installs ovfetch,
which gives vinowhisper-doctor Intel’s per-platform NPU driver data: the
first driver verified on your NPU, and the OpenVINO range recorded as working
with the driver you have. Setup offers it only when there is an Intel NPU,
downloads the static release binary and checks it against the sha256 pinned in
this package, and leaves an ovfetch that is already current alone. Nothing
needs it; without it the doctor just has two fewer lines. See
hardware.md.
Pinning the terminal on top
Section titled “Pinning the terminal on top”The overlay above stays on top by itself. For the terminal UI instead: the status bar is Rich in an ordinary terminal, so keeping it above other windows is a window-manager job, not the app’s. On KWin: System Settings > Window Management > Window Rules, match the terminal window, set Keep Above Other Windows to Force/Yes, plus Skip Taskbar and Skip Pager if you want it out of the way. No titlebar and a small fixed size make it read like an overlay rather than a terminal.