Skip to content

Installing

The short version is in the README. This is what the installer actually does, and the dependency pin behind it.

Terminal window
curl -fsSL https://raw.githubusercontent.com/karanshukla/vinoWhisper/main/scripts/install.sh | bash

That installs uv (a pinned version, only if you have none), clones the latest vX.Y.Z release tag rather than main (--ref main for unreleased code, --ref v0.6.1 for a specific one), builds the environment with uv sync --locked, so exactly what uv.lock pins and never a fresh resolve, and hands over to vinowhisper-setup, which is where every machine-specific decision happens: your capture tool, your NPU driver, the model export your device needs, and systemd units generated against the paths that actually exist. It prints every command before running it and asks first. --yes answers yes to everything and --dry-run no to everything, and piped stdin without --yes refuses rather than run sudo commands nobody agreed to. A step that is already done says so and is skipped. It never invents a package command where a distro has none (the NPU userspace driver, on most), and points at Intel’s releases instead. The ~/.local/bin entries are symlinks into the checkout, so they follow a git pull, and bash completion gets one link per command, because bash-completion loads a file lazily on the first Tab for a command of the same name.

The whole script is one main function called on its last line, so a download cut short by the network is a syntax error that runs nothing, rather than half an install. The release tag comes from git ls-remote on the repository, not the GitHub API, so there is no rate limit and no JSON to parse. Re-running it on an existing checkout moves that checkout to the requested ref (detached), so a checkout installed from main goes back to the latest release unless you pass --ref main again.

From a checkout, or to see what it would do without doing it:

Terminal window
git clone https://github.com/karanshukla/vinoWhisper && cd vinoWhisper
uv sync --extra export # --extra export: the one-time model export
uv run vinowhisper-setup --dry-run # the whole plan, nothing changed
uv run vinowhisper-setup # for real, one prompt per step
Terminal window
pip install vinowhisper

This works, and until 2026-08-31 it could not. The dependency set resolved openvino from PyPI, where the builds that can construct the NPU static Whisper pipeline did not exist, so a pip install would have succeeded and then failed to load a model, which is worse than not shipping at all. Stable 2026.3.1 removed that constraint (see the version floor below).

What it gets you: the five commands and the Python dependencies. What it cannot get you: an NPU driver, a model export, or systemd units. Run vinowhisper-setup afterwards for those, exactly as the installer script would. Re-run it after upgrading from 0.6.x or earlier too: the server moved from 127.0.0.1:8099 to a Unix socket in $XDG_RUNTIME_DIR, and the wizard is what rewrites the socket unit and restarts it. The export tooling (optimum, and torch with it) is the export extra, which the wizard offers to install when it gets to the model step; pip install 'vinowhisper[export]' does it up front.

Python must be 3.11-3.13. 3.14 made functools.partial a descriptor, which breaks optimum’s NORMALIZED_CONFIG_CLASS = SomeConfig.with_args(...) class-attribute idiom outright. Version-independent root cause, confirmed 2026-08-03 across every optimum/transformers pairing tried. requires-python enforces it, so pip will refuse rather than install something broken.

openvino>=2026.3.1, and the floor is exact rather than cautious: 2026.3.1 is the first stable release that can build the NPU static Whisper pipeline.

From 2026-08-03 to 2026-08-31 this project pinned nightly wheels from storage.openvinotoolkit.org, with prereleases allowed globally, because stable 2026.2.1 could not build that pipeline at all: its pipeline_static.cpp pattern-matcher did not recognise the optimum-intel export’s SDPA attention-mask node shape, and OPENVINO_ASSERT(!self_attn_nodes.empty()) failed. That pin was the project’s standing dependency risk, since upstream prunes nightly builds on its own schedule.

Re-measured 2026-08-31 on the Wildcat Lake NPU, against the same --disable-stateful export, with WhisperPipeline(..., STATIC_PIPELINE=True):

Version Source Pipeline build generate()
2026.3.1 PyPI stable 2.0s ok
2026.4.0.dev20260805 nightly 2.3s ok
2026.5.0.dev20260831 nightly 1.9s ok

Stable caught up, so the nightly index, the prerelease = "allow" policy and the [tool.uv.sources] routing are all gone and these resolve from PyPI like anything else. .github/workflows/deps-canary.yml still runs the full resolve weekly, off the pull-request path, but now for ordinary upstream churn rather than for wheels aging out from under the lock.

STATIC_PIPELINE=True is not optional and does not degrade. Omitting it sends the NPU down the generic stateful path, which fails with Stateful models without 'beam_idx' input are not supported in StatefulToStateless transformation. That reads like a bad export and is not one.

Exporting downloads ~1GB from Hugging Face and converts it to OpenVINO IR that then runs on your hardware. The download comes first and is checked first: vinowhisper/model_sources.json pins the repository to a commit and records the sha256 of each file the export reads, python -m vinowhisper.source fetches exactly those into models/source/ under the data directory, and any difference fails before optimum-cli starts. The export then runs from that directory with HF_HUB_OFFLINE=1. optimum-cli has no --revision flag (optimum-intel 2.2.0), which is why the snapshot is local rather than in the Hugging Face cache. A --model with no pinned source is exported unchecked, with a warning.

vinowhisper/model_digests.json pins the sha256 of every file in the export this project has actually run, and both scripts/convert_model.sh and vinowhisper-setup check what came down against it. vinowhisper-doctor re-checks it on demand, at about 1.2s for 1.5GB, and only for an export that is present and the right shape for its device, since “hashes don’t match” is noise next to “wrong export entirely”.

The pin is on the exported IR, not on the upstream safetensors, because the IR is what WhisperPipeline loads and the export is not a pure function of the weights. Six statuses, and only three of them stop anything:

Status Means Blocks setup
verified every pinned file matches
unpinned no pin for this model and variant no
drift bytes and export toolchain both moved no
mismatch the pinned toolchain produced different bytes yes
incomplete pinned files are missing, so the export is partial yes
known_bad exported by a toolchain measured to produce a broken export yes

An unpinned export warning is not a problem to fix. It is what --model openai/whisper-base.en looks like, and what the stateful export looks like until someone produces one. Re-pin with ./scripts/update_digests.py --variant npu once you trust an export, and commit the diff.

The export is bit-reproducible, which is what makes any of this work. Measured 2026-09-04: two independent optimum-cli exports of whisper-small.en on the same toolchain produced all 16 files byte-identical. Across toolchains it is not: against the 2026-08-03 export (OpenVINO 2026.2.1, optimum-intel 2.0.0, transformers 5.0.0), an export under OpenVINO 2026.3.1 / optimum-intel 2.1.0 / transformers 5.5.4 changed 9 of 16 files, including both decoder .bin weights. openvino_encoder_model.bin came out identical across both. That is why drift is reported separately from a real mismatch, and the versions are read out of the export’s own rt_info block rather than from whatever happens to be installed. That makes drift the export’s own claim about itself: something able to rewrite the export can rewrite every rt_info block to match and get drift instead of mismatch. The source check is what guards the download; this one catches accidents.

That block is read from every .xml in the export and merged, and any version two files disagree on makes the answer unknown, never fine: reading just one file would let a single edited graph buy the softer drift verdict (editing all of them consistently still does). The digest covers every file, not just the weights, because generation_config.json decides how decoding behaves and tokenizer.json decides the text. A missing or truncated pin file downgrades to unpinned rather than breaking every export. Verification is deliberately not part of loading the model, since hashing takes about 1.2s and would land on the socket-activated cold start.

known_bad entries come in two shapes: an exact combination measured broken, or a floor (at_least, as for transformers 5.4.0). A version the export does not report never satisfies a floor, and update_digests.py rewrites the hashes without touching the list.

Bisected on hardware 2026-09-04. Export with transformers<5.4. Anything from 5.4.0 on produces a Whisper decoder the NPU static pipeline compiles and then cannot run:

RuntimeError: Port for tensor name cache_position was not found.
(src/inference/src/cpp/infer_request.cpp:191)

It fails at generate(), not at load, so nothing complains until the first transcription.

The bisect held optimum-intel 2.1.0, optimum 2.3.0, openvino 2026.3.1, openvino-genai 2026.3.1.0 and torch 2.13.0 fixed, exported whisper-small.en with --disable-stateful in a clean venv per version, and loaded each on the NPU:

transformers Pipeline build generate()
5.0.0 ok 0.96s
5.2.0 ok 0.86s
5.3.0 ok 0.71s
5.4.0 ok fails
5.5.4 ok fails

Two control runs rule out the rest of the stack. optimum-intel 2.1.0 with transformers 5.0.0 works, so the optimum pair is not at fault; the full 2026-08-03 package set re-run under openvino 2026.3.1 also works, so the runtime is not either.

The mechanism is a tensor name. cache_position appears exactly once in each transformers 5.3.0 decoder graph, on the output port of __module.model.model.decoder/aten::arange/Range, and zero times in the 5.4.0 graphs. It is neither a model input nor a model output in either export, so the static pipeline is resolving an internal traced tensor by name, and 5.4.0 stopped emitting that name. The exported input and output signatures are otherwise identical between the two.

Since 2026-09-12 the export’s dependencies are their own extra, vinowhisper[export], which holds transformers<5.4. Before that a fresh install resolved 5.5.4 and vinowhisper-setup failed at the model step. The extra is separate because only the export uses optimum and transformers, and optimum brings torch with it, while the runtime imports none of them. vinowhisper-setup offers to install it when it needs to export, as uv sync --extra export in a checkout and pip install 'vinowhisper[export]' otherwise.

The pin’s known_bad entry still carries a floor at transformers 5.4.0, for an export made some other way, so vinowhisper-setup and convert_model.sh report a known_bad export rather than handing over a model that fails later. The stateful (CPU/GPU) export is unaffected: it builds and decodes under 5.5.4 (measured 2026-09-12).

Raising the cap does not upgrade anything (2026-09-20). Every released optimum-intel, 2.2.0 included, requires transformers<5.6,>=4.51. Ask for anything past that and the resolver satisfies it by backtracking optimum to a pre-transformers-5.x release instead, at which point optimum-cli fails on import rather than at export time. Two open transformers advisories, GHSA-fgcw-684q-jj6r (fixed in 5.5.0) and GHSA-xrqw-3rrv-vx5w (fixed in 5.10.0), are unreachable for that reason and the bisection above. Nothing at runtime imports transformers; the only call site is the one-time optimum-cli export openvino in scripts/convert_model.sh, against a hardcoded openai/whisper-small.en unless you pass your own --model.

vinowhisper-setup --gui installs vinowhisper-gui, a floating caption box that also does dictation, with a tray icon and global shortcuts. It downloads the binary from the GitHub release and checks it against the sha256 pinned in the Python package, so no Rust toolchain is needed. It is never on PyPI; from a checkout it can be built with cargo instead (./scripts/install.sh --gui). See gui.md.

vinowhisper-setup --ovfetch installs ovfetch, which gives vinowhisper-doctor Intel’s per-platform NPU driver data: the first driver verified on your NPU, and the OpenVINO range recorded as working with the driver you have. Setup offers it only when there is an Intel NPU, downloads the static release binary and checks it against the sha256 pinned in this package, and leaves an ovfetch that is already current alone. Nothing needs it; without it the doctor just has two fewer lines. See hardware.md.

The overlay above stays on top by itself. For the terminal UI instead: the status bar is Rich in an ordinary terminal, so keeping it above other windows is a window-manager job, not the app’s. On KWin: System Settings > Window Management > Window Rules, match the terminal window, set Keep Above Other Windows to Force/Yes, plus Skip Taskbar and Skip Pager if you want it out of the way. No titlebar and a small fixed size make it read like an overlay rather than a terminal.