# Startup context overview Run `mise run //:context-startup`, optionally followed by `-- /path/to/project`. Requires installed Pi on PATH and valid existing credentials. Makes three live first-turn requests with `Reply only OK. Do not use tools.`. For a local-only split, run `mise run //:context-startup-sources`. It estimates `APPEND_SYSTEM.md`, `pi/agent` skill entries, extension/package skill entries, and shared catalog instructions/wrappers separately, using installed Pi's formatter and token heuristic. It uses the configured zero-message startup, including extension `resources_discover` hooks and the same read-only credential backend, but sends no measurement prompt. Skills are classified using Pi's loaded source metadata: package or `extension:` sources count as extensions, agent-directory skills count as `pi/agent`, and direct project-local skills are excluded from attribution without changing discovery. Unknown visible sources and skill diagnostics fail the estimate rather than silently dropping entries. The catalog includes names, descriptions, and locations, not skill bodies or prompt templates. Shared instructions are counted once; each group is rounded independently, so group sums may differ slightly from the combined total. These numbers measure the startup skill catalog, not arbitrary first-turn text injected by extensions. Run `mise run //:context-startup-skill-descriptions` to rank all visible `pi/agent` skills by description length and display the largest five descriptions for review. That ranking estimates raw description text only, excluding names, paths, XML escaping, and catalog markup; equal-length descriptions sort by skill name. ## Measurement contract | Case | Config | Mode | User messages | Measurement | | --- | --- | --- | --- | --- | | baseline+1msg | Stock Pi | Headless | 1 | Provider input tokens and estimated component tokens | | headless+0msg | Your config | Headless | 0 | Estimated startup tokens only | | headless+1msg | Your config | Headless | 1 | Provider input tokens and estimated component tokens | | headless+1msg+no_pi_config | Your extensions/settings, no agent-dir text resources | Headless | 1 | Provider input tokens and estimated component tokens | The configured zero-message run resolves the model and thinking level, which are then pinned for every request. All cases use the same cwd and installed Pi version in separate processes. The baseline uses empty in-memory settings, an empty agent directory, stock tools, and no discovered extensions, skills, prompts, themes, or context files. It retains Pi's built-in extensions and stock system prompt. A selected model absent from stock Pi fails the baseline rather than importing user model config. Configured runs honor existing project trust; the script never grants trust. `no_pi_config` removes the agent directory's context file (`AGENTS.md` or `CLAUDE.md`), `SYSTEM.md`, `APPEND_SYSTEM.md`, and resources under its `skills/` and `prompts/` directories using Pi's resource-loader overrides. It keeps the same real agent directory, settings, models, extension discovery, extension-specific settings, and extension-injected content. Project context, trusted project resources, themes, and package-provided skills/prompts remain unchanged. Explicitly configured resources outside those agent-directory paths are retained. This is text-resource isolation, not an empty settings directory; nothing is renamed, copied, or disabled on disk. Unresolved project trust blocks the comparison rather than silently dropping project config. All cases use Pi's SDK startup services, ephemeral sessions, offline startup catalogs, and Pi's read-only credential backend. Credentials are consumed internally by Pi, never copied or printed; expired OAuth credentials cannot be refreshed by this script. This bypasses CLI onboarding, migrations, and update checks, not the headless extension lifecycle. Extensions may still perform their normal state/log writes and background requests. ## Reading the output `input tok` is provider-reported input plus cache reads and cache writes, excluding output tokens. Cached input occupies context too. `~sys tok`, `~tools tok`, and `~msgs tok` use the installed Pi SDK's `estimateTokens()` heuristic. For text, Pi uses `Math.ceil(text.length / 4)`, based on JavaScript string length, not UTF-8 bytes. System text and serialized active tool definitions are each estimated as text. Tool definitions include names, descriptions, and parameter schemas, not implementation code. Messages use Pi's per-message estimator directly, including its image allowance and per-message rounding. Empty components count as zero. Zero-message estimates are captured after startup hooks finish. One-message estimates are captured at the agent's stream boundary after context transformations, before provider serialization. No raw prompts, provider payloads, credentials, or response transcripts are saved by the script. Configured headless minus baseline measures total config overhead, including changed tool definitions. `Pi text resources` is configured headless minus `no_pi_config`, using provider input tokens. It measures the effect of removing those resources with extensions/settings held fixed, not the residual extensions-only overhead: project context still remains. Extensions that react to missing resources can also contribute to this difference. Configured first-request component estimates minus configured zero-message estimates show first-turn growth, including the short user message. All displayed sizes are tokens; `~` distinguishes heuristic estimates from provider-reported counts. Estimates need not sum to provider input because tokenization, provider serialization, and payload-rewriting hooks can differ. A single run is an overview, not a stability benchmark or per-extension attribution. TUI measurement was removed after two live comparisons matched configured headless input at 24,579 tokens. Recheck manually if UI-dependent context behavior changes. Every worker exits after measurement; no terminal window is opened. Temporary measurement files contain counts and run options only and are removed after collection. Failures, extra agent requests, tool calls, startup errors, or mismatched model/version/cwd prevent misleading diffs. ## Per-extension measurement Run `mise run //:context-extensions`, optionally followed by `-- /path/to/project`. A zero-message startup discovers the current file-backed extensions from Pi's runtime, including registered tools, event-handler names, and queued provider ownership. There is no extension-name allowlist: newly installed extensions automatically enter the next measurement. Registration signals prioritize likely context contributors but cannot prove an extension is harmless; arbitrary startup side effects, unknown hooks, and lazy registration remain possible. Every discovered file extension is therefore measured, including uncertain cases. The script makes N+2 live requests: a full-set control, one fresh process omitting each of N extensions, and a repeated full-set control. All request runs explicitly load the discovered paths in the same order, minus the one omitted path. Omission happens before module evaluation, not by filtering registrations after an extension already ran. Pi's inline built-ins and the measurement guard remain loaded. Settings, project trust, model, thinking level, cwd, prompt, and static package resources stay fixed. An extension's dynamically discovered resources disappear with its discovery hook, but package-declared skills/prompts remain: this measures an extension module, not uninstalling its whole package. Positive savings mean removal reduces provider input; negative savings mean removal increases input. Heuristic system/tool/message savings are reported separately. Deltas are not additive because extensions can interact. A failing omission, dependency break, changed model, or incorrect extension set is a failed measurement, never zero savings. If the final full-set control fails or differs in provider input or local component estimates, all savings are withheld. Matching endpoint controls detect some drift, not every transient change; do not edit configuration during measurement. Each worker retains the existing two-minute timeout, count-only output, read-only credentials, and no-tool-execution guard. Temporary inventory/options/results are removed after collection; raw prompts and provider payloads are never saved. ## On-demand guidance checkpoint Current measurement: Pi 0.85.1, `openai-codex/gpt-6-astra`, high thinking, 39 file-backed extensions. Provider input is 8,934 tokens, 477 below the earlier 9,411-token checkpoint. The 41-request extension profile produced matching full-set controls at 8,934 tokens. The no-Pi-text case is 4,857 provider tokens. | Extension | Earlier removal saving | Current removal saving | | --- | ---: | ---: | | Ultra | 617 | 353 | | Web | 561 | 546 | | The System | 393 | 288 | | Terse | 240 | 240 | | Git-safe | 210 | 120 | The other 34 extensions had zero fixed first-turn impact for this probe. These are provider-token snapshot comparisons, not source-isolated before/after runs; removal savings remain non-additive. Current first-turn heuristic estimates are 7,596 system tokens, 2,250 tool tokens, and 8 message tokens, not a breakdown of the provider total. ## Terse retirement checkpoint With Pi 0.85.1 and `openai-codex/gpt-6-astra` at high thinking, replacing Terse with five response-style sentences in the global `APPEND_SYSTEM.md` reduced configured first-request provider input from 8,072 to 7,911 tokens, a saving of 161 tokens. The stock baseline remained 1,071 tokens; the no-Pi-text case fell from 4,716 to 4,476 tokens. The configured system estimate fell from 6,847 to 6,713 heuristic tokens; tools and messages were unchanged. These before/after startup runs measure the combined migration, not response quality or an additive per-extension profile. The extension set now contains 38 packages; the earlier Terse profile above is historical. ## Checks Run `mise run //:context-startup-test` for formatting, unit checks, and controlled subprocess fixtures. Run `mise run //:context-startup-fmt` to format only these measurement files. A real overview run additionally exercises installed Pi; fixture checks alone do not prove live behavior. Run `mise run //:context-guidance-check` to verify read-only live activation of tool-owned mockup guidance. It requires the configured model; it plans previews without writing files or opening a browser. Run `mise run //extensions/ultra:authoring-check` for the corresponding workflow authoring check.