Luigit
repositories / dotfiles

dotfiles

bugabingas dorkfiles

owned by admin

quickshell/nuguland/NOKO.md

Raw
Rendered preview

Noko

Minimal local speech-to-text tool using Parakeet.

Scope

Noko owns the core transcription flow only:

mic audio
→ mono float samples
→ resample to 16 kHz
→ split into 30 ms frames
→ VAD keeps speech frames
→ concatenate speech samples
→ Parakeet ONNX inference
→ trim/filter text
→ save history/output text

Out of scope for this spec: hotkeys, triggers, UI shell, app updater, system autostart, desktop integration.

Audio contract

  • Capture PCM from selected input device.
  • Convert all input to mono float32.
  • Target sample rate is 16_000 Hz.
  • Multi-channel input is averaged to mono.
  • If device runs at another rate, resample before VAD/inference.

VAD contract

VAD is separate from Parakeet. It is a pre-inference speech gate.

  • Use Silero VAD ONNX or equivalent.
  • Input frames are exactly 30 ms.
  • At 16 kHz, one frame is 480 samples.
  • Output only speech samples to Parakeet.

VAD defaults:

  • threshold: 0.3
  • frame size: 30 ms
  • pre-roll: 15 frames
  • onset: 2 frames
  • hangover: 15 frames

Smoothing state machine:

not_speech + speech frame
→ count onset
→ if onset threshold met: enter speech, flush pre-roll

speech + speech frame
→ keep frame, reset hangover

speech + noise frame
→ keep frame while hangover remains

not_speech + noise frame
→ drop frame

Parakeet runtime

  • Parakeet is directory-based, not a single model file.
  • Expected model contents conceptually:
    • preprocessor ONNX
    • encoder ONNX
    • decoder/joint ONNX
    • vocabulary/tokenizer file
  • Prefer int8 model directories for size and speed.
  • Use ONNX Runtime or a wrapper around it.
  • Load model once and reuse it.
  • Serialize inference unless runtime thread-safety is proven.

Engine states:

unloaded → loading → ready → transcribing → ready
                         ↘ failed

Inference contract

  • Input: one continuous float32 sample buffer, mono, 16 kHz, speech-only.
  • Reject empty buffers.
  • Optional: prepend small silence for timestamp/model stability.
  • Timestamp granularity can be segment-level; minimal Noko only needs text.
  • Output: result.text.trim().

Postprocess

Minimal:

  • trim whitespace
  • collapse repeated spaces
  • drop empty result
  • remove known junk tokens if observed

Optional:

  • custom word correction
  • punctuation/capitalization cleanup

History

Save after successful transcription.

Fields:

  • timestamp
  • transcript
  • duration or sample count
  • model id/version
  • quantization
  • runtime provider
  • optional audio path
  • optional error record for failed transcription

Failure cases

Handle explicitly:

  • no input device
  • unsupported sample format
  • resampler failure
  • VAD model missing
  • Parakeet model dir incomplete
  • runtime provider unavailable
  • empty speech buffer after VAD
  • inference error/crash
  • CPU instruction incompatibility

Main lesson

Parakeet should receive clean, normalized, speech-only 16 kHz mono audio. Resampling, VAD, smoothing, history, and output are app responsibilities outside Parakeet.

Noko-owned assets

Noko does not depend on another app's model cache. By default it owns assets under $NOKO_DATA_DIR, or $XDG_DATA_HOME/nuguland/noko, or ~/.local/share/nuguland/noko.

Commands:

node noko.mjs assets-json
node noko.mjs check-assets-json
node noko.mjs install-assets
node noko.mjs update-assets
node noko.mjs install-assets-progress
node noko.mjs update-assets-progress
node noko.mjs transcribe-capture-json 5 history.jsonl
node noko.mjs transcribe-raw-json capture.f32 history.jsonl

check-assets-json uses remote HEAD metadata and local file sizes to report whether VAD/model updates are available. install-assets installs/fetches missing files. update-assets redownloads model/VAD files from the configured source URLs. The *-progress variants emit NDJSON progress events for the popup progress bar.

install-assets installs/fetches:

  • onnxruntime-node into runtime/node_modules/
  • Silero VAD ONNX into vad/silero_vad.onnx
  • Parakeet TDT int8 ONNX files into parakeet-tdt-0.6b-v3-int8/

System runtime deps are checked by doctor and shown in setup:

  • pw-record for PipeWire mic capture
  • wtype for typing finished transcripts into the focused Wayland control

During install-assets, Noko installs missing system deps through pkexec and the host package manager when a supported plan exists. This deliberately triggers Nuguland's Polkit prompt instead of hiding privilege escalation.

NOKO_VAD_ONNX and NOKO_PARAKEET_DIR are overrides only.

Implementation status

  • noko.js: core audio normalization, 16 kHz resampling, 30 ms framing, VAD smoothing, postprocess, history records, injected Parakeet pipeline.
  • noko.mjs: CLI doctor, setup installer, real PipeWire capture to 16 kHz mono float JSON, raw indefinite capture transcription, real ONNX VAD probabilities, real VAD-gated speech JSON, one-command mic→VAD→Parakeet transcription, plus JSON fixture transcription path.
  • noko-vad.mjs: real Silero/equivalent ONNX VAD bridge; accepts 480-sample Noko frames and adapts to Silero state/context feeds.
  • noko-parakeet.mjs: Parakeet model-directory discovery, ONNX preprocessor→encoder→TDT decoder/joint inference, vocab decode, optional SentencePiece token decode, and serialized inference guard.
  • noko.test.js: unit, injected end-to-end, VAD bridge, asset-path, and model-directory tests.
  • Noko.qml + bar wiring: compact Nuguland status chip, automatic update checks, setup/update popup, system-runtime checks, toggle capture, capture waveform, focused-window text insertion via wtype, scrollable transcript history, and IPC.

Verified end to end with:

  • real captured/fixture 16 kHz mono audio
  • temporary Silero VAD ONNX + onnxruntime-node
  • Noko-owned Parakeet TDT int8 model directory outside any other app cache
  • default NOKO_DATA_DIR asset resolution with no other-app path in assets-json or doctor
  • default Node onnxruntime-node Parakeet backend running preprocessor→encoder→TDT decoder/joint→vocab decode
  • output transcript Hello world
  • persisted history JSONL record

Default doctor points at Noko-owned asset paths and reports missing runtime/assets until install-assets has populated them. sentencepiece-js is only required for model directories that use tokenizer.model; the verified Parakeet directory uses vocab.txt.

Bar/popup widget states:

  • checking: transient after refresh
  • blocked: missing VAD, Parakeet assets, or runtime; chip label noko! with warning role
  • installing: setup is downloading runtime/models; warning strong chip
  • ready: all required local assets/runtime are readable; chip label noko
  • listening: indefinite capture is running; subtle animated accent line in the bar chip
  • transcribing: inference is running; subtle animated accent line in the bar chip
  • error: invalid doctor output or process failure; warning role

Popup actions:

  • refresh doctor
  • automatically check setup updates
  • install/update setup with progress bar
  • start/stop toggle capture
  • show capture waveform and scrollable local transcript history

Global trigger:

  • niri owns Super+\`` because global compositor keybinds belong in niri/cfg/keybinds.kdl`.
  • Nuguland owns what that trigger does through qs ipc -c nuguland call noko trigger.
  • Finished transcripts are typed into the focused Wayland control with wtype -- <text>.
  • Toggle mode is the only supported mode: first press starts capture, second press stops and transcribes.
  • Push-to-talk is intentionally unsupported because niri's normal keybind action does not provide key-release events.
# Noko

Minimal local speech-to-text tool using Parakeet.

## Scope

Noko owns the core transcription flow only:

```text
mic audio
→ mono float samples
→ resample to 16 kHz
→ split into 30 ms frames
→ VAD keeps speech frames
→ concatenate speech samples
→ Parakeet ONNX inference
→ trim/filter text
→ save history/output text
```

Out of scope for this spec: hotkeys, triggers, UI shell, app updater, system autostart, desktop integration.

## Audio contract

- Capture PCM from selected input device.
- Convert all input to mono `float32`.
- Target sample rate is `16_000 Hz`.
- Multi-channel input is averaged to mono.
- If device runs at another rate, resample before VAD/inference.

## VAD contract

VAD is separate from Parakeet. It is a pre-inference speech gate.

- Use Silero VAD ONNX or equivalent.
- Input frames are exactly 30 ms.
- At 16 kHz, one frame is `480` samples.
- Output only speech samples to Parakeet.

VAD defaults:

- threshold: `0.3`
- frame size: `30 ms`
- pre-roll: `15` frames
- onset: `2` frames
- hangover: `15` frames

Smoothing state machine:

```text
not_speech + speech frame
→ count onset
→ if onset threshold met: enter speech, flush pre-roll

speech + speech frame
→ keep frame, reset hangover

speech + noise frame
→ keep frame while hangover remains

not_speech + noise frame
→ drop frame
```

## Parakeet runtime

- Parakeet is directory-based, not a single model file.
- Expected model contents conceptually:
  - preprocessor ONNX
  - encoder ONNX
  - decoder/joint ONNX
  - vocabulary/tokenizer file
- Prefer int8 model directories for size and speed.
- Use ONNX Runtime or a wrapper around it.
- Load model once and reuse it.
- Serialize inference unless runtime thread-safety is proven.

Engine states:

```text
unloaded → loading → ready → transcribing → ready
                         ↘ failed
```

## Inference contract

- Input: one continuous `float32` sample buffer, mono, 16 kHz, speech-only.
- Reject empty buffers.
- Optional: prepend small silence for timestamp/model stability.
- Timestamp granularity can be segment-level; minimal Noko only needs text.
- Output: `result.text.trim()`.

## Postprocess

Minimal:

- trim whitespace
- collapse repeated spaces
- drop empty result
- remove known junk tokens if observed

Optional:

- custom word correction
- punctuation/capitalization cleanup

## History

Save after successful transcription.

Fields:

- timestamp
- transcript
- duration or sample count
- model id/version
- quantization
- runtime provider
- optional audio path
- optional error record for failed transcription

## Failure cases

Handle explicitly:

- no input device
- unsupported sample format
- resampler failure
- VAD model missing
- Parakeet model dir incomplete
- runtime provider unavailable
- empty speech buffer after VAD
- inference error/crash
- CPU instruction incompatibility

## Main lesson

Parakeet should receive clean, normalized, speech-only 16 kHz mono audio. Resampling, VAD, smoothing, history, and output are app responsibilities outside Parakeet.

## Noko-owned assets

Noko does not depend on another app's model cache. By default it owns assets under `$NOKO_DATA_DIR`, or `$XDG_DATA_HOME/nuguland/noko`, or `~/.local/share/nuguland/noko`.

Commands:

```sh
node noko.mjs assets-json
node noko.mjs check-assets-json
node noko.mjs install-assets
node noko.mjs update-assets
node noko.mjs install-assets-progress
node noko.mjs update-assets-progress
node noko.mjs transcribe-capture-json 5 history.jsonl
node noko.mjs transcribe-raw-json capture.f32 history.jsonl
```

`check-assets-json` uses remote `HEAD` metadata and local file sizes to report whether VAD/model updates are available. `install-assets` installs/fetches missing files. `update-assets` redownloads model/VAD files from the configured source URLs. The `*-progress` variants emit NDJSON progress events for the popup progress bar.

`install-assets` installs/fetches:

- `onnxruntime-node` into `runtime/node_modules/`
- Silero VAD ONNX into `vad/silero_vad.onnx`
- Parakeet TDT int8 ONNX files into `parakeet-tdt-0.6b-v3-int8/`

System runtime deps are checked by `doctor` and shown in setup:

- `pw-record` for PipeWire mic capture
- `wtype` for typing finished transcripts into the focused Wayland control

During `install-assets`, Noko installs missing system deps through `pkexec` and the host package manager when a supported plan exists. This deliberately triggers Nuguland's Polkit prompt instead of hiding privilege escalation.

`NOKO_VAD_ONNX` and `NOKO_PARAKEET_DIR` are overrides only.

## Implementation status

- `noko.js`: core audio normalization, 16 kHz resampling, 30 ms framing, VAD smoothing, postprocess, history records, injected Parakeet pipeline.
- `noko.mjs`: CLI doctor, setup installer, real PipeWire capture to 16 kHz mono float JSON, raw indefinite capture transcription, real ONNX VAD probabilities, real VAD-gated speech JSON, one-command mic→VAD→Parakeet transcription, plus JSON fixture transcription path.
- `noko-vad.mjs`: real Silero/equivalent ONNX VAD bridge; accepts 480-sample Noko frames and adapts to Silero state/context feeds.
- `noko-parakeet.mjs`: Parakeet model-directory discovery, ONNX preprocessor→encoder→TDT decoder/joint inference, vocab decode, optional SentencePiece token decode, and serialized inference guard.
- `noko.test.js`: unit, injected end-to-end, VAD bridge, asset-path, and model-directory tests.
- `Noko.qml` + bar wiring: compact Nuguland status chip, automatic update checks, setup/update popup, system-runtime checks, toggle capture, capture waveform, focused-window text insertion via `wtype`, scrollable transcript history, and IPC.

Verified end to end with:

- real captured/fixture 16 kHz mono audio
- temporary Silero VAD ONNX + `onnxruntime-node`
- Noko-owned Parakeet TDT int8 model directory outside any other app cache
- default `NOKO_DATA_DIR` asset resolution with no other-app path in `assets-json` or `doctor`
- default Node `onnxruntime-node` Parakeet backend running preprocessor→encoder→TDT decoder/joint→vocab decode
- output transcript `Hello world`
- persisted history JSONL record

Default `doctor` points at Noko-owned asset paths and reports missing runtime/assets until `install-assets` has populated them. `sentencepiece-js` is only required for model directories that use `tokenizer.model`; the verified Parakeet directory uses `vocab.txt`.

Bar/popup widget states:

- `checking`: transient after refresh
- `blocked`: missing VAD, Parakeet assets, or runtime; chip label `noko!` with warning role
- `installing`: setup is downloading runtime/models; warning strong chip
- `ready`: all required local assets/runtime are readable; chip label `noko`
- `listening`: indefinite capture is running; subtle animated accent line in the bar chip
- `transcribing`: inference is running; subtle animated accent line in the bar chip
- `error`: invalid doctor output or process failure; warning role

Popup actions:

- refresh doctor
- automatically check setup updates
- install/update setup with progress bar
- start/stop toggle capture
- show capture waveform and scrollable local transcript history

Global trigger:

- niri owns `Super+\`` because global compositor keybinds belong in `niri/cfg/keybinds.kdl`.
- Nuguland owns what that trigger does through `qs ipc -c nuguland call noko trigger`.
- Finished transcripts are typed into the focused Wayland control with `wtype -- <text>`.
- Toggle mode is the only supported mode: first press starts capture, second press stops and transcribes.
- Push-to-talk is intentionally unsupported because niri's normal keybind action does not provide key-release events.