Luigit
repositories / termux-janitor

termux-janitor

Interactive cleanup assistant for Termux: transparent, safe, confirmed disk reclamation.

owned by admin

spec/PERFORMANCE.md

Raw
Rendered preview

Performance model

PRODUCT.md incorporates this document as normative version 1 behavior. It defines the performance resource model, explicit budgets, expected bottlenecks, measurement methods, benchmarks, profiling evidence, and regression criteria.

Every numeric bound is owned by LIMITS.md through model/limits.zon. This document references registered limit identities and never restates their values. A budget without a registered identity requires a limits registry change first; this document never introduces a number.

Resource model

TJ-PERF-01. All application storage is allocated once during initialization from the registered fixed capacities. Runtime never grows a bound, never frees and reallocates a capacity pool, and never persists a scan cache. Exhaustion follows the registered fail-closed behavior.

Computation uses one UI thread and the registered fixed worker count. The UI thread performs at most its registered per-turn work and communicates with workers only through the registered bounded queues. No worker blocks the UI thread, and no UI wait depends on worker progress.

Input and output processing is bounded per turn: registered input bytes, events, cells, and terminal bytes per turn. External processes run with registered concurrency, output capture, argument, and timeout bounds. Filesystem access is descriptor-relative and batched at registered cancellation checkpoints. Network access occurs only through admitted adapter scope.

Memory, CPU, I/O, and process costs are therefore functions of the registered capacities and the scanned tree, never of elapsed time, arrival order, or operator input rate.

Explicit budgets

TJ-PERF-02. A performance budget is a registered limit identity plus its enforcement site and verification suite. The version 1 budgets are:

Budget Registered identity Enforced at Verified by
Acknowledged input transition input_acknowledgement UI event loop PTY
UI turn work ui_work_per_turn UI event loop SM, PTY
Worker cancellation acknowledgment worker_cancellation_interval workers SIM
Discovery, planning, mutation completion discovery_timeout, planning_timeout, mutation_timeout process supervision A, SIM
Terminal output per turn terminal_bytes_per_turn, cells_per_turn, terminal_writes_per_turn renderer R, PTY
Process escalation interrupt_grace_period, termination_grace_period, kill_and_reap_period process supervision A

A budget violation is a specification and implementation defect, never a tuning hint. Deleting or weakening a budget follows the limit-change policy in LIMITS.md.

Expected bottlenecks

TJ-PERF-03. The version 1 bottlenecks, in expected impact order, are:

  1. directory enumeration and metadata batches, bounded by cancellation checkpoints and descriptor batching;
  2. filesystem identity deduplication for hard-link accounting and per-filesystem totals;
  3. external package-manager process latency, bounded by supervision timeouts;
  4. Unicode width-table lookup during wrap and track rendering;
  5. cell-grid differencing and pending-output draining;
  6. bounded log append and flush batches.

Each bottleneck has a registered bound or a cancellation checkpoint. An optimization must not introduce an unregistered capacity, a second allocation path, or work outside the turn and queue model.

Measurement methods

TJ-PERF-04. Deterministic measurement uses seeded simulation corpora, traced checkpoint counts, and fixed iteration counts. Deterministic product logic never reads a clock for control flow; the generation realtime reference exists only for age eligibility.

On-device measurement records the device model, Android and Termux versions, thermal and power state when available, the method, and the repetition count, and dates the record per EXPERIMENTAL_EVIDENCE.md. A measured result supports conclusions only for the recorded environment and never generalizes across devices or versions. Timing use follows CODING_STYLE.md: measurement code is separate from product logic and is never persisted as a scan cache.

Benchmarks

TJ-PERF-05. Benchmarks are a named build.zig step that runs deterministic seeded scenarios and writes a bounded text report. Version 1 benchmark scenarios are:

  1. cold scan of a seeded tree at the retained-finding capacity boundary;
  2. planning at the plan-action and manifest-entry capacity boundaries;
  3. rendering the maximum registered cell grid;
  4. input decoding at the registered per-turn input bound;
  5. log append at the registered audit-event capacity.

Benchmarks use no network, no external package managers, and no timing assertions inside deterministic test gates. A benchmark result never passes or fails a gate by itself; regression criteria below decide.

Profiling evidence

TJ-PERF-06. Profiling evidence records the tool, exact command, capture location, retention location, date, and recorded environment. Captures are repository evidence, not release artifacts, unless ARTIFACT.md registers them. A profile supports a bottleneck or budget conclusion only together with the deterministic measurement that reproduces it. Platform limitation records accompany every affected requirement, and expired evidence follows the evidence rules in EXPERIMENTAL_EVIDENCE.md.

Prior-art evaluation

Traversal, matching, aggregation, and bounded-concurrency techniques from ripgrep and ugrep were evaluated against the version 1 model. Adopted concepts are already structural: bounded parallel directory workers instead of unbounded parallelism, precomputed lookup tables instead of per-byte branching (the checked-in Unicode width table), and byte-oriented scanning delegated to standard-library routines that already use vectorized implementations. Rejected: memory-mapped I/O as a correctness dependency because Android shared-storage and FUSE semantics do not support it, and persisted indexes or scan caches because the specification forbids them. Deferred: SIMD beyond standard-library routines and data-oriented storage rewrites until profiling evidence under TJ-PERF-06 identifies a bottleneck; any such change then passes the TJ-PERF-07 regression criteria.

Regression criteria

TJ-PERF-07. A performance regression is any of:

  1. a deterministic budget violation: any measured run exceeding a registered budget identity on a capable platform;
  2. an allocation after initialization that grows a registered capacity or creates an unregistered one;
  3. work outside the registered turn, queue, and worker model that a traced checkpoint can observe.

A regression fails the gate that owns the evidence. It is never resolved by a warning, a retry budget, or an undocumented bound change. Restoring compliance requires either an implementation change or a limits-registry change with memory-cost evidence and boundary tests per LIMITS.md.

# Performance model

[`PRODUCT.md`](PRODUCT.md) incorporates this document as normative version 1 behavior.
It defines the performance resource model, explicit budgets, expected bottlenecks, measurement
methods, benchmarks, profiling evidence, and regression criteria.

Every numeric bound is owned by [`LIMITS.md`](LIMITS.md) through
[`model/limits.zon`](model/limits.zon). This document references registered limit identities and
never restates their values. A budget without a registered identity requires a limits registry
change first; this document never introduces a number.

## Resource model

**TJ-PERF-01.** All application storage is allocated once during initialization from the registered
fixed capacities. Runtime never grows a bound, never frees and reallocates a capacity pool, and
never persists a scan cache. Exhaustion follows the registered fail-closed behavior.

Computation uses one UI thread and the registered fixed worker count. The UI thread performs at most
its registered per-turn work and communicates with workers only through the registered bounded
queues. No worker blocks the UI thread, and no UI wait depends on worker progress.

Input and output processing is bounded per turn: registered input bytes, events, cells, and terminal
bytes per turn. External processes run with registered concurrency, output capture, argument, and
timeout bounds. Filesystem access is descriptor-relative and batched at registered cancellation
checkpoints. Network access occurs only through admitted adapter scope.

Memory, CPU, I/O, and process costs are therefore functions of the registered capacities and the
scanned tree, never of elapsed time, arrival order, or operator input rate.

## Explicit budgets

**TJ-PERF-02.** A performance budget is a registered limit identity plus its enforcement site and
verification suite. The version 1 budgets are:

| Budget | Registered identity | Enforced at | Verified by |
| --- | --- | --- | --- |
| Acknowledged input transition | `input_acknowledgement` | UI event loop | PTY |
| UI turn work | `ui_work_per_turn` | UI event loop | SM, PTY |
| Worker cancellation acknowledgment | `worker_cancellation_interval` | workers | SIM |
| Discovery, planning, mutation completion | `discovery_timeout`, `planning_timeout`, `mutation_timeout` | process supervision | A, SIM |
| Terminal output per turn | `terminal_bytes_per_turn`, `cells_per_turn`, `terminal_writes_per_turn` | renderer | R, PTY |
| Process escalation | `interrupt_grace_period`, `termination_grace_period`, `kill_and_reap_period` | process supervision | A |

A budget violation is a specification and implementation defect, never a tuning hint. Deleting or
weakening a budget follows the limit-change policy in [`LIMITS.md`](LIMITS.md#limit-changes).

## Expected bottlenecks

**TJ-PERF-03.** The version 1 bottlenecks, in expected impact order, are:

1. directory enumeration and metadata batches, bounded by cancellation checkpoints and descriptor
   batching;
2. filesystem identity deduplication for hard-link accounting and per-filesystem totals;
3. external package-manager process latency, bounded by supervision timeouts;
4. Unicode width-table lookup during wrap and track rendering;
5. cell-grid differencing and pending-output draining;
6. bounded log append and flush batches.

Each bottleneck has a registered bound or a cancellation checkpoint. An optimization must not
introduce an unregistered capacity, a second allocation path, or work outside the turn and queue
model.

## Measurement methods

**TJ-PERF-04.** Deterministic measurement uses seeded simulation corpora, traced checkpoint counts,
and fixed iteration counts. Deterministic product logic never reads a clock for control flow; the
generation realtime reference exists only for age eligibility.

On-device measurement records the device model, Android and Termux versions, thermal and power state
when available, the method, and the repetition count, and dates the record per
[`EXPERIMENTAL_EVIDENCE.md`](EXPERIMENTAL_EVIDENCE.md#evidence-rules). A measured result supports
conclusions only for the recorded environment and never generalizes across devices or versions.
Timing use follows [`CODING_STYLE.md`](CODING_STYLE.md#runtime-architecture): measurement code is
separate from product logic and is never persisted as a scan cache.

## Benchmarks

**TJ-PERF-05.** Benchmarks are a named `build.zig` step that runs deterministic seeded scenarios and
writes a bounded text report. Version 1 benchmark scenarios are:

1. cold scan of a seeded tree at the retained-finding capacity boundary;
2. planning at the plan-action and manifest-entry capacity boundaries;
3. rendering the maximum registered cell grid;
4. input decoding at the registered per-turn input bound;
5. log append at the registered audit-event capacity.

Benchmarks use no network, no external package managers, and no timing assertions inside
deterministic test gates. A benchmark result never passes or fails a gate by itself; regression
criteria below decide.

## Profiling evidence

**TJ-PERF-06.** Profiling evidence records the tool, exact command, capture location, retention
location, date, and recorded environment. Captures are repository evidence, not release artifacts,
unless [`ARTIFACT.md`](ARTIFACT.md) registers them. A profile supports a bottleneck or budget
conclusion only together with the deterministic measurement that reproduces it. Platform limitation
records accompany every affected requirement, and expired evidence follows the evidence rules in
[`EXPERIMENTAL_EVIDENCE.md`](EXPERIMENTAL_EVIDENCE.md#evidence-rules).

## Prior-art evaluation

Traversal, matching, aggregation, and bounded-concurrency techniques from ripgrep and ugrep were
evaluated against the version 1 model. Adopted concepts are already structural: bounded parallel
directory workers instead of unbounded parallelism, precomputed lookup tables instead of per-byte
branching (the checked-in Unicode width table), and byte-oriented scanning delegated to
standard-library routines that already use vectorized implementations. Rejected: memory-mapped I/O
as a correctness dependency because Android shared-storage and FUSE semantics do not support it,
and persisted indexes or scan caches because the specification forbids them. Deferred: SIMD beyond
standard-library routines and data-oriented storage rewrites until profiling evidence under
TJ-PERF-06 identifies a bottleneck; any such change then passes the TJ-PERF-07 regression criteria.

## Regression criteria

**TJ-PERF-07.** A performance regression is any of:

1. a deterministic budget violation: any measured run exceeding a registered budget identity on a
   capable platform;
2. an allocation after initialization that grows a registered capacity or creates an unregistered
   one;
3. work outside the registered turn, queue, and worker model that a traced checkpoint can observe.

A regression fails the gate that owns the evidence. It is never resolved by a warning, a retry
budget, or an undocumented bound change. Restoring compliance requires either an implementation
change or a limits-registry change with memory-cost evidence and boundary tests per
[`LIMITS.md`](LIMITS.md#limit-changes).