# Noko Minimal local speech-to-text tool using Parakeet. ## Scope Noko owns the core transcription flow only: ```text mic audio → mono float samples → resample to 16 kHz → split into 30 ms frames → VAD keeps speech frames → concatenate speech samples → Parakeet ONNX inference → trim/filter text → save history/output text ``` Out of scope for this spec: hotkeys, triggers, UI shell, app updater, system autostart, desktop integration. ## Audio contract - Capture PCM from selected input device. - Convert all input to mono `float32`. - Target sample rate is `16_000 Hz`. - Multi-channel input is averaged to mono. - If device runs at another rate, resample before VAD/inference. ## VAD contract VAD is separate from Parakeet. It is a pre-inference speech gate. - Use Silero VAD ONNX or equivalent. - Input frames are exactly 30 ms. - At 16 kHz, one frame is `480` samples. - Output only speech samples to Parakeet. VAD defaults: - threshold: `0.3` - frame size: `30 ms` - pre-roll: `15` frames - onset: `2` frames - hangover: `15` frames Smoothing state machine: ```text not_speech + speech frame → count onset → if onset threshold met: enter speech, flush pre-roll speech + speech frame → keep frame, reset hangover speech + noise frame → keep frame while hangover remains not_speech + noise frame → drop frame ``` ## Parakeet runtime - Parakeet is directory-based, not a single model file. - Expected model contents conceptually: - preprocessor ONNX - encoder ONNX - decoder/joint ONNX - vocabulary/tokenizer file - Prefer int8 model directories for size and speed. - Use ONNX Runtime or a wrapper around it. - Load model once and reuse it. - Serialize inference unless runtime thread-safety is proven. Engine states: ```text unloaded → loading → ready → transcribing → ready ↘ failed ``` ## Inference contract - Input: one continuous `float32` sample buffer, mono, 16 kHz, speech-only. - Reject empty buffers. - Optional: prepend small silence for timestamp/model stability. - Timestamp granularity can be segment-level; minimal Noko only needs text. - Output: `result.text.trim()`. ## Postprocess Minimal: - trim whitespace - collapse repeated spaces - drop empty result - remove known junk tokens if observed Optional: - custom word correction - punctuation/capitalization cleanup ## History Save after successful transcription. Fields: - timestamp - transcript - duration or sample count - model id/version - quantization - runtime provider - optional audio path - optional error record for failed transcription ## Failure cases Handle explicitly: - no input device - unsupported sample format - resampler failure - VAD model missing - Parakeet model dir incomplete - runtime provider unavailable - empty speech buffer after VAD - inference error/crash - CPU instruction incompatibility ## Main lesson Parakeet should receive clean, normalized, speech-only 16 kHz mono audio. Resampling, VAD, smoothing, history, and output are app responsibilities outside Parakeet. ## Noko-owned assets Noko does not depend on another app's model cache. By default it owns assets under `$NOKO_DATA_DIR`, or `$XDG_DATA_HOME/nuguland/noko`, or `~/.local/share/nuguland/noko`. Commands: ```sh node noko.mjs assets-json node noko.mjs check-assets-json node noko.mjs install-assets node noko.mjs update-assets node noko.mjs install-assets-progress node noko.mjs update-assets-progress node noko.mjs transcribe-capture-json 5 history.jsonl node noko.mjs transcribe-raw-json capture.f32 history.jsonl ``` `check-assets-json` uses remote `HEAD` metadata and local file sizes to report whether VAD/model updates are available. `install-assets` installs/fetches missing files. `update-assets` redownloads model/VAD files from the configured source URLs. The `*-progress` variants emit NDJSON progress events for the popup progress bar. `install-assets` installs/fetches: - `onnxruntime-node` into `runtime/node_modules/` - Silero VAD ONNX into `vad/silero_vad.onnx` - Parakeet TDT int8 ONNX files into `parakeet-tdt-0.6b-v3-int8/` System runtime deps are checked by `doctor` and shown in setup: - `pw-record` for PipeWire mic capture - `wtype` for typing finished transcripts into the focused Wayland control During `install-assets`, Noko installs missing system deps through `pkexec` and the host package manager when a supported plan exists. This deliberately triggers Nuguland's Polkit prompt instead of hiding privilege escalation. `NOKO_VAD_ONNX` and `NOKO_PARAKEET_DIR` are overrides only. ## Implementation status - `noko.js`: core audio normalization, 16 kHz resampling, 30 ms framing, VAD smoothing, postprocess, history records, injected Parakeet pipeline. - `noko.mjs`: CLI doctor, setup installer, real PipeWire capture to 16 kHz mono float JSON, raw indefinite capture transcription, real ONNX VAD probabilities, real VAD-gated speech JSON, one-command mic→VAD→Parakeet transcription, plus JSON fixture transcription path. - `noko-vad.mjs`: real Silero/equivalent ONNX VAD bridge; accepts 480-sample Noko frames and adapts to Silero state/context feeds. - `noko-parakeet.mjs`: Parakeet model-directory discovery, ONNX preprocessor→encoder→TDT decoder/joint inference, vocab decode, optional SentencePiece token decode, and serialized inference guard. - `noko.test.js`: unit, injected end-to-end, VAD bridge, asset-path, and model-directory tests. - `Noko.qml` + bar wiring: compact Nuguland status chip, automatic update checks, setup/update popup, system-runtime checks, toggle capture, capture waveform, focused-window text insertion via `wtype`, scrollable transcript history, and IPC. Verified end to end with: - real captured/fixture 16 kHz mono audio - temporary Silero VAD ONNX + `onnxruntime-node` - Noko-owned Parakeet TDT int8 model directory outside any other app cache - default `NOKO_DATA_DIR` asset resolution with no other-app path in `assets-json` or `doctor` - default Node `onnxruntime-node` Parakeet backend running preprocessor→encoder→TDT decoder/joint→vocab decode - output transcript `Hello world` - persisted history JSONL record Default `doctor` points at Noko-owned asset paths and reports missing runtime/assets until `install-assets` has populated them. `sentencepiece-js` is only required for model directories that use `tokenizer.model`; the verified Parakeet directory uses `vocab.txt`. Bar/popup widget states: - `checking`: transient after refresh - `blocked`: missing VAD, Parakeet assets, or runtime; chip label `noko!` with warning role - `installing`: setup is downloading runtime/models; warning strong chip - `ready`: all required local assets/runtime are readable; chip label `noko` - `listening`: indefinite capture is running; subtle animated accent line in the bar chip - `transcribing`: inference is running; subtle animated accent line in the bar chip - `error`: invalid doctor output or process failure; warning role Popup actions: - refresh doctor - automatically check setup updates - install/update setup with progress bar - start/stop toggle capture - show capture waveform and scrollable local transcript history Global trigger: - niri owns `Super+\`` because global compositor keybinds belong in `niri/cfg/keybinds.kdl`. - Nuguland owns what that trigger does through `qs ipc -c nuguland call noko trigger`. - Finished transcripts are typed into the focused Wayland control with `wtype -- `. - Toggle mode is the only supported mode: first press starts capture, second press stops and transcribes. - Push-to-talk is intentionally unsupported because niri's normal keybind action does not provide key-release events.