mic audio
→ mono float samples
→ resample to 16 kHz
→ split into 30 ms frames
→ VAD keeps speech frames
→ concatenate speech samples
→ Parakeet ONNX inference
→ trim/filter text
→ save history/output text
Out of scope for this spec: hotkeys, triggers, UI shell, app updater, system autostart, desktop integration.
Audio contract
Capture PCM from selected input device.
Convert all input to mono float32.
Target sample rate is 16_000 Hz.
Multi-channel input is averaged to mono.
If device runs at another rate, resample before VAD/inference.
VAD contract
VAD is separate from Parakeet. It is a pre-inference speech gate.
Input: one continuous float32 sample buffer, mono, 16 kHz, speech-only.
Reject empty buffers.
Optional: prepend small silence for timestamp/model stability.
Timestamp granularity can be segment-level; minimal Noko only needs text.
Output: result.text.trim().
Postprocess
Minimal:
trim whitespace
collapse repeated spaces
drop empty result
remove known junk tokens if observed
Optional:
custom word correction
punctuation/capitalization cleanup
History
Save after successful transcription.
Fields:
timestamp
transcript
duration or sample count
model id/version
quantization
runtime provider
optional audio path
optional error record for failed transcription
Failure cases
Handle explicitly:
no input device
unsupported sample format
resampler failure
VAD model missing
Parakeet model dir incomplete
runtime provider unavailable
empty speech buffer after VAD
inference error/crash
CPU instruction incompatibility
Main lesson
Parakeet should receive clean, normalized, speech-only 16 kHz mono audio. Resampling, VAD, smoothing, history, and output are app responsibilities outside Parakeet.
Noko-owned assets
Noko does not depend on another app's model cache. By default it owns assets under $NOKO_DATA_DIR, or $XDG_DATA_HOME/nuguland/noko, or ~/.local/share/nuguland/noko.
check-assets-json uses remote HEAD metadata and local file sizes to report whether VAD/model updates are available. install-assets installs/fetches missing files. update-assets redownloads model/VAD files from the configured source URLs. The *-progress variants emit NDJSON progress events for the popup progress bar.
install-assets installs/fetches:
onnxruntime-node into runtime/node_modules/
Silero VAD ONNX into vad/silero_vad.onnx
Parakeet TDT int8 ONNX files into parakeet-tdt-0.6b-v3-int8/
System runtime deps are checked by doctor and shown in setup:
pw-record for PipeWire mic capture
wtype for typing finished transcripts into the focused Wayland control
During install-assets, Noko installs missing system deps through pkexec and the host package manager when a supported plan exists. This deliberately triggers Nuguland's Polkit prompt instead of hiding privilege escalation.
NOKO_VAD_ONNX and NOKO_PARAKEET_DIR are overrides only.
Implementation status
noko.js: core audio normalization, 16 kHz resampling, 30 ms framing, VAD smoothing, postprocess, history records, injected Parakeet pipeline.
noko.mjs: CLI doctor, setup installer, real PipeWire capture to 16 kHz mono float JSON, raw indefinite capture transcription, real ONNX VAD probabilities, real VAD-gated speech JSON, one-command mic→VAD→Parakeet transcription, plus JSON fixture transcription path.
noko-vad.mjs: real Silero/equivalent ONNX VAD bridge; accepts 480-sample Noko frames and adapts to Silero state/context feeds.
Default doctor points at Noko-owned asset paths and reports missing runtime/assets until install-assets has populated them. sentencepiece-js is only required for model directories that use tokenizer.model; the verified Parakeet directory uses vocab.txt.
Bar/popup widget states:
checking: transient after refresh
blocked: missing VAD, Parakeet assets, or runtime; chip label noko! with warning role
installing: setup is downloading runtime/models; warning strong chip
ready: all required local assets/runtime are readable; chip label noko
listening: indefinite capture is running; subtle animated accent line in the bar chip
transcribing: inference is running; subtle animated accent line in the bar chip
error: invalid doctor output or process failure; warning role
Popup actions:
refresh doctor
automatically check setup updates
install/update setup with progress bar
start/stop toggle capture
show capture waveform and scrollable local transcript history
Global trigger:
niri owns Super+\`` because global compositor keybinds belong in niri/cfg/keybinds.kdl`.
Nuguland owns what that trigger does through qs ipc -c nuguland call noko trigger.
Finished transcripts are typed into the focused Wayland control with wtype -- <text>.
Toggle mode is the only supported mode: first press starts capture, second press stops and transcribes.
Push-to-talk is intentionally unsupported because niri's normal keybind action does not provide key-release events.
# Noko
Minimal local speech-to-text tool using Parakeet.
## Scope
Noko owns the core transcription flow only:
```text
mic audio
→ mono float samples
→ resample to 16 kHz
→ split into 30 ms frames
→ VAD keeps speech frames
→ concatenate speech samples
→ Parakeet ONNX inference
→ trim/filter text
→ save history/output text
```
Out of scope for this spec: hotkeys, triggers, UI shell, app updater, system autostart, desktop integration.
## Audio contract
- Capture PCM from selected input device.
- Convert all input to mono `float32`.
- Target sample rate is `16_000 Hz`.
- Multi-channel input is averaged to mono.
- If device runs at another rate, resample before VAD/inference.
## VAD contract
VAD is separate from Parakeet. It is a pre-inference speech gate.
- Use Silero VAD ONNX or equivalent.
- Input frames are exactly 30 ms.
- At 16 kHz, one frame is `480` samples.
- Output only speech samples to Parakeet.
VAD defaults:
- threshold: `0.3`
- frame size: `30 ms`
- pre-roll: `15` frames
- onset: `2` frames
- hangover: `15` frames
Smoothing state machine:
```text
not_speech + speech frame
→ count onset
→ if onset threshold met: enter speech, flush pre-roll
speech + speech frame
→ keep frame, reset hangover
speech + noise frame
→ keep frame while hangover remains
not_speech + noise frame
→ drop frame
```
## Parakeet runtime
- Parakeet is directory-based, not a single model file.
- Expected model contents conceptually:
- preprocessor ONNX
- encoder ONNX
- decoder/joint ONNX
- vocabulary/tokenizer file
- Prefer int8 model directories for size and speed.
- Use ONNX Runtime or a wrapper around it.
- Load model once and reuse it.
- Serialize inference unless runtime thread-safety is proven.
Engine states:
```text
unloaded → loading → ready → transcribing → ready
↘ failed
```
## Inference contract
- Input: one continuous `float32` sample buffer, mono, 16 kHz, speech-only.
- Reject empty buffers.
- Optional: prepend small silence for timestamp/model stability.
- Timestamp granularity can be segment-level; minimal Noko only needs text.
- Output: `result.text.trim()`.
## Postprocess
Minimal:
- trim whitespace
- collapse repeated spaces
- drop empty result
- remove known junk tokens if observed
Optional:
- custom word correction
- punctuation/capitalization cleanup
## History
Save after successful transcription.
Fields:
- timestamp
- transcript
- duration or sample count
- model id/version
- quantization
- runtime provider
- optional audio path
- optional error record for failed transcription
## Failure cases
Handle explicitly:
- no input device
- unsupported sample format
- resampler failure
- VAD model missing
- Parakeet model dir incomplete
- runtime provider unavailable
- empty speech buffer after VAD
- inference error/crash
- CPU instruction incompatibility
## Main lesson
Parakeet should receive clean, normalized, speech-only 16 kHz mono audio. Resampling, VAD, smoothing, history, and output are app responsibilities outside Parakeet.
## Noko-owned assets
Noko does not depend on another app's model cache. By default it owns assets under `$NOKO_DATA_DIR`, or `$XDG_DATA_HOME/nuguland/noko`, or `~/.local/share/nuguland/noko`.
Commands:
```sh
node noko.mjs assets-json
node noko.mjs check-assets-json
node noko.mjs install-assets
node noko.mjs update-assets
node noko.mjs install-assets-progress
node noko.mjs update-assets-progress
node noko.mjs transcribe-capture-json 5 history.jsonl
node noko.mjs transcribe-raw-json capture.f32 history.jsonl
```
`check-assets-json` uses remote `HEAD` metadata and local file sizes to report whether VAD/model updates are available. `install-assets` installs/fetches missing files. `update-assets` redownloads model/VAD files from the configured source URLs. The `*-progress` variants emit NDJSON progress events for the popup progress bar.
`install-assets` installs/fetches:
- `onnxruntime-node` into `runtime/node_modules/`
- Silero VAD ONNX into `vad/silero_vad.onnx`
- Parakeet TDT int8 ONNX files into `parakeet-tdt-0.6b-v3-int8/`
System runtime deps are checked by `doctor` and shown in setup:
- `pw-record` for PipeWire mic capture
- `wtype` for typing finished transcripts into the focused Wayland control
During `install-assets`, Noko installs missing system deps through `pkexec` and the host package manager when a supported plan exists. This deliberately triggers Nuguland's Polkit prompt instead of hiding privilege escalation.
`NOKO_VAD_ONNX` and `NOKO_PARAKEET_DIR` are overrides only.
## Implementation status
- `noko.js`: core audio normalization, 16 kHz resampling, 30 ms framing, VAD smoothing, postprocess, history records, injected Parakeet pipeline.
- `noko.mjs`: CLI doctor, setup installer, real PipeWire capture to 16 kHz mono float JSON, raw indefinite capture transcription, real ONNX VAD probabilities, real VAD-gated speech JSON, one-command mic→VAD→Parakeet transcription, plus JSON fixture transcription path.
- `noko-vad.mjs`: real Silero/equivalent ONNX VAD bridge; accepts 480-sample Noko frames and adapts to Silero state/context feeds.
- `noko-parakeet.mjs`: Parakeet model-directory discovery, ONNX preprocessor→encoder→TDT decoder/joint inference, vocab decode, optional SentencePiece token decode, and serialized inference guard.
- `noko.test.js`: unit, injected end-to-end, VAD bridge, asset-path, and model-directory tests.
- `Noko.qml` + bar wiring: compact Nuguland status chip, automatic update checks, setup/update popup, system-runtime checks, toggle capture, capture waveform, focused-window text insertion via `wtype`, scrollable transcript history, and IPC.
Verified end to end with:
- real captured/fixture 16 kHz mono audio
- temporary Silero VAD ONNX + `onnxruntime-node`
- Noko-owned Parakeet TDT int8 model directory outside any other app cache
- default `NOKO_DATA_DIR` asset resolution with no other-app path in `assets-json` or `doctor`
- default Node `onnxruntime-node` Parakeet backend running preprocessor→encoder→TDT decoder/joint→vocab decode
- output transcript `Hello world`
- persisted history JSONL record
Default `doctor` points at Noko-owned asset paths and reports missing runtime/assets until `install-assets` has populated them. `sentencepiece-js` is only required for model directories that use `tokenizer.model`; the verified Parakeet directory uses `vocab.txt`.
Bar/popup widget states:
- `checking`: transient after refresh
- `blocked`: missing VAD, Parakeet assets, or runtime; chip label `noko!` with warning role
- `installing`: setup is downloading runtime/models; warning strong chip
- `ready`: all required local assets/runtime are readable; chip label `noko`
- `listening`: indefinite capture is running; subtle animated accent line in the bar chip
- `transcribing`: inference is running; subtle animated accent line in the bar chip
- `error`: invalid doctor output or process failure; warning role
Popup actions:
- refresh doctor
- automatically check setup updates
- install/update setup with progress bar
- start/stop toggle capture
- show capture waveform and scrollable local transcript history
Global trigger:
- niri owns `Super+\`` because global compositor keybinds belong in `niri/cfg/keybinds.kdl`.
- Nuguland owns what that trigger does through `qs ipc -c nuguland call noko trigger`.
- Finished transcripts are typed into the focused Wayland control with `wtype -- <text>`.
- Toggle mode is the only supported mode: first press starts capture, second press stops and transcribes.
- Push-to-talk is intentionally unsupported because niri's normal keybind action does not provide key-release events.