--- id: NL-SPEC-37DAB33A type: spec title: System popup --- # System popup ## Behavior - The popup shows machine identity and live resources, then the top five CPU consumers, followed by current issues and recent incidents. - Identity includes host, model, distribution, kernel, window manager, package count, shell, uptime, and battery facts when available. - Live statistics include CPU, memory, swap, temperature and fan when available, network, disk, and load. - Metrics use graphical fill bars and bounded sparklines rather than a terminal aesthetic. - Cheap metrics refresh about every two seconds. - Expensive disk, failed-service, and process checks run about every five seconds only while open. - Identity and update checks are cached and refreshed on open or infrequently. - Issues cover failed services, thermal state, low disk, memory with growing swap, battery health, updates, and load relative to core count. - Current conditions are severity-ordered under `needs attention`; recorded crashes and OOM kills are newest-first under `incident history`. History rows use neutral cards and subdued clock icons, never warning/error card emphasis, because a past event does not establish a current failure. Current failures retain their severity styling; history retains counts, times, open/`arrrgh` actions, and existing journal-window expiry. No acknowledgment/dismissal state is added and no journal entries or coredumps are deleted. Dispatch errors remain visibly actionable even on neutral history rows. - Repeated crashes with the same exact executable path and UID share one row, showing occurrence count and the latest crash time. Missing or lossy executable/UID identity leaves crashes separate; different applications/users and OOM events never merge. Groups are derived from retained journal records without dismissal state or additional persistence, and keep stable IDs as occurrences enter or leave the retained window. Groups are newest-first by their latest occurrence; the header counts records, not rows. The existing open action inspects the latest crash, while `arrrgh` receives every retained occurrence in the selected group, including exact boot/cursor/time and bounded metadata. Pending, error/retry, keyboard focus/reveal, and close-on-dispatch remain owned by the existing resource service and popup controls. - Incidents are limited to journal-visible records in the current boot and last 24 hours, up to the latest 100 matching records. Missing records do not imply that no application crashed; unavailable or incomplete history is explicit and retains last-good data. - Failed services, thermal conditions, disk/memory/load pressure, crashes, and OOM kills have an `arrrgh` button with the Nerd Font pirate icon and `assess and diagnose with pi` tooltip. Available updates and battery wear have no agent action, including at the dispatch boundary. Existing open actions remain available, including the update action. - App/service identity leads each row in emphasized text; issue type and metadata are secondary, with relative and local calendar times for incidents. Untrusted metadata is plain text, never markup. - `arrrgh` starts a fresh interactive Pi session in a new WezTerm window rooted at `$XDG_STATE_HOME/nuguland/repairs/`. Missing or relative `XDG_STATE_HOME` falls back to `~/.local/state`; the working directory stays stable across issues and sessions. Launch preflight creates the directory privately if absent, preserves existing contents and permissions, and refuses dispatch when it is unusable. The shared prompt lives in `quickshell/nuguland/prompts/system-repair.md`, not global agent configuration. Exactly one literal `{{context}}` slot receives JSON facts: issue identity/type/severity, observation time, machine identity, relevant measurements, service-manager scopes, and incident boot/time/cursor/process metadata. It contains no prescribed diagnostic commands, inferred cause, raw journal payloads, environment dumps, or core memory. The agent selects its own diagnostic tools. The template is read asynchronously on each launch; missing or malformed templates block dispatch with an error instead of reusing stale instructions. Each launch saves a private, uniquely named `.request-*.md` prompt snapshot in the repair directory and passes it to Pi as an `@file` argument. This preserves every retained occurrence even when the prompt exceeds Linux's per-argument size limit; write failures block dispatch with a visible retryable error. Request snapshots are repair evidence, never widget state, and earlier snapshots are not overwritten. - The prompt requires per-issue Markdown reports under `issues/` inside the stable repairs directory. Reports identify the exact issue and machine, and append dated attempts without replacing history. They contain the problem, evidence-backed root cause, attempted solutions and outcomes, verified fix, verification steps/results, learnings/prevention, blockers, and next steps. Only verified successful fixes are marked resolved; unsuccessful or partially verified attempts remain unresolved. Reports exclude secrets and raw sensitive logs; the agent verifies saving and reports the log path/status or save failure. Report writing is a prompt obligation, not an automatic resolution detector. - The shared instruction is to assess, diagnose, and recommend or apply an appropriate remedy, not to force a fix for every warning. Monitoring, no change, or a recommendation can be justified outcomes; an already-gone incident is not an invented successful repair. - Diagnosis comes first; privileged, destructive, package, service restart, and desktop-session changes require explicit human approval. This is a prompt constraint, not a new OS sandbox or permission mechanism. - Resource service owns dispatch state: pending disables duplicate launches, missing Pi/WezTerm leaves a visible error and retry, dispatch closes the popup to release keyboard focus. Dispatch means a launch request, not proof that Pi or its provider initialized successfully. - Pointer, Tab/Shift+Tab, Enter/Space, Escape, outside dismissal, and reopen retain the shared popup contract. - Empty current conditions and empty incident history have separate lowercase status rows. ## Constraints - The popup remains smooth on the target ThinkPad X230. - Metric history is strictly bounded. - Closing the popup stops expensive probes. - Missing or failed data is shown as `n/a` or omitted and never creates a false issue. - Failed probes retain last-good data and do not crash the service. - Battery checks are skipped when no laptop battery exists. - Existing system-state IPC remains available with functional state only. ## Issue thresholds - Failed services are errors. - Disk use at 90% is a warning and at 95% is an error. - Memory pressure requires at least 90% memory use and growing swap. - Battery health below 80% is a warning and below 60% is an error. - Thermal degradation is an error; crossing a sensor high threshold is at least a warning. - Available updates are informational. ## Acceptance - Pure tests cover parsers, rate and ring behavior, formatting, thresholds, unavailable inputs, and action shaping. - Runtime self-test covers issue wiring and IPC shaping. - Nested-niri rendered screenshots and pointer/keyboard tests confirm identity and live cards, issue grouping, agent dispatch, failure/retry, long metadata, empty/unavailable history, and polling lifecycle. - Synthetic crash/OOM records and action-planner tests cover filtering, malformed data, stable identities, limits, service scopes, literal argv transport, and prompt safety. ## Evidence [Agent actions and incident detection](evidence/arrrgh-2026-09-06/README.md). [Stable repair directory and issue reports](evidence/repair-log-2026-09-06/README.md). [Local repair prompt and eligible actions](evidence/repair-prompt-2026-09-06/README.md). [Repeated application crashes](evidence/crash-groups-2026-09-06/README.md). [Process line above issues](evidence/process-line-2026-09-06/README.md). [Neutral incident history](evidence/neutral-history-2026-09-06/README.md).