# ultra Multi-agent workflow orchestration for Pi. A declarative JSON spec fans sub-agents out across phases under a concurrency cap, validates their structured output, and aggregates the results — with live progress on both an interactive overlay and a model-callable tool surface. ## Two surfaces - **`/ultra`**: enables the hidden `run_workflow` tool for later agent turns, confirms activation, and does not start a turn. - **`/ultra exec `**: enables the tool, sends the request as a visible user message, and starts an agent turn. - **`/ultra run [args]`**: runs a saved workflow behind a phase-first live overlay with keyboard phase toggling, agent focus, streaming transcripts, steering, targeted abort, and confirmed whole-run abort. - **`run_workflow` tool**: model-callable after `/ultra` enables it. Pass a saved workflow `name` (+ optional `args`), or an inline `spec` object for a one-off run. Returns the aggregated JSON; progress streams to a collapsible result surface. ## Workflows A workflow is a single JSON object — phases (`single` / `fanout`), per-step prompts/tools/schemas, and a closed interpolation grammar that pipes one phase's results into the next. The authoritative format reference is - [`prompts/references/spec-format.md`](prompts/references/spec-format.md) — the canonical authoring reference, kept in lockstep with the runtime validator by a test. The runtime schema is also exposed on the active tool. Don't reconstruct the format from memory; read it. ### Authoring Ultra registers the [Ultra authoring skill](prompts/ultra-authoring/SKILL.md) while the extension is enabled. Its description appears in Pi's skill catalog and its body remains on demand; it explains command routing, workflow mechanics, child context, dynamic extensions, results, and model fields. The active `run_workflow` tool adds one concise guideline describing when decomposition reduces context or verification risk. `/ultra exec` sends only the user's request. The [canonical spec reference](prompts/references/spec-format.md) provides detailed field semantics and examples. The provider-facing schema exposes the structural workflow fields and descriptions for inline `spec` objects. Both command and tool invocations still pass through the strict runtime validator, including conditional fanout and write-isolation requirements that JSON Schema cannot express. The effective model-tier data remains dynamically injected while the workflow tool is active. Each compact entry preserves the configured model, default thinking level, and selection guidance. `mise run //extensions/ultra:authoring-check` runs a read-only planning request using the configured model. It checks the live event stream for a successful skill read and a completed turn, rejecting any tool use other than `read`. ### Where workflows live Discovery merges three locations, lowest-to-highest precedence (a later source overrides a bundled name). Workflow files are rescanned on every completion or invocation, so newly saved JSON works immediately without `/reload`. 1. **bundled** — `extensions/ultra/workflows/*.json` (ships with ultra) 2. **global** — `~/.pi/agent/workflows/*.json` 3. **project** — `/.pi/workflows/*.json` The shipped `review` workflow reviews a target across dimensions, adversarially verifies each finding, then synthesizes a concise analysis with concrete next steps. Run it with `/ultra run review [target]` for a PR, branch/ref, recent changes, folder, or file; no target reviews the working tree and recent local changes. Verification preserves the original target, stable unique finding IDs, evidence, and explicit `confirmed`, `refuted`, or `unresolved` verdict statuses with reasons. The report keeps unresolved candidates in an explicit unresolved list; missing verification is not refutation. The shipped `feature` workflow drives a small feature end to end: research → plan → adversarial critique → strengthened plan → conditional implement → conditional security/quality review → final report. Run it with `/ultra run feature ` (the whole trailing text is the task). A dedicated read-only phase rewrites the plan to resolve critique blockers before implementation. The strengthened `Plan` readiness gate permits writes only for a settled assignment; unresolved decisions block implementation instead of being delegated as routine execution. Only its `implement` phase writes; every other phase is read-only. The shipped `bug` workflow drives investigate → conditional regression test → conditional fix → adversarial verification → final report. Run it with `/ultra run bug `. The investigator supplies a settled correction brief with relevant instructions, evidence, and acceptance checks before regression and fix work proceeds. The fix runs only after the regression phase proves the expected failure; verification runs only after a fix. Both bug and feature select a final report preserving `completed`, `blocked`, or `unresolved` status, actual changes and checks, blockers, and remaining work. If cancellation or a dropped reporting step prevents that report, the output envelope still exposes execution failure and the full evidence path. Empty, skipped, dropped, or inconclusive verification never establishes successful completion. The shipped `research` workflow drives question planning → parallel multi-source evidence gathering → gap and contradiction audit → targeted adversarial verification → cited synthesis. Run it with `/ultra run research `. Loaded research extensions may activate their own tools for its agents. ## Results While the model streams `run_workflow` arguments, the call preview shows received phase counts, the latest three phase summaries, and two lines from the latest received prompt. Text advances with incoming argument chunks, like Pi's write preview, without artificial delays or background animation timers. These are unvalidated inputs, not running agents or completion estimates; saved workflow calls show their name and received input count without loading files during rendering. Once execution starts, the draft collapses to a compact call header and the existing progress board takes over. Use Pi's tool expansion shortcut (`Ctrl+O` by default) to inspect every received phase, tool/model metadata, and full received prompts, including after execution or failure. Completed review runs show synthesized analysis, confirmed and refuted findings, unresolved candidates, drops, and concrete recommended next steps. Expand results to inspect complete workflow details as labeled fields and nested lists, not raw JSON. Agent-facing structured payloads remain JSON. For both tool and command results, the full aggregate is always stored once and every bounded model-facing envelope includes `fullOutputPath`. The envelope contains the selected `return` and configured `report` values when they fit, phase counts, bounded failure diagnostics, and usage, not duplicated `phaseResults`. When `report` is configured, it is persisted in the full aggregate, included in the bounded model-facing envelope, and rendered on human-facing surfaces. `executionStatus` is `finished`, `incomplete` for dropped steps, or `aborted`; `finished` describes execution only, not artifact quality or semantic success. Oversized values use a retrieval notice within the 50KB or 2000 lines limit; read the saved aggregate selectively. Complete contents, including intermediate results and failure inputs/transcript tails, remain in UI details. Bounding the envelope does not compress or discard the stored evidence. Workflow runs append incomplete-run checkpoints to the current Pi session as custom entries that never enter LLM context. Each launched step also gets a named persistent child session in the parent's session directory, linked to the parent session and resumable after interruption. The parent journal stores child identity before prompting; the next matching run continues unchanged incomplete steps and reruns changed work in clean children. Checkpoint identity includes the delivered agent envelope, immediate producer attribution, and resolved output schema. Changing assignment content, attribution, or a named schema invalidates stale child work. Reused outputs must satisfy the current schema, and a schema-valid submission alone never proves task completion. Completed runs start fresh but retain their named child sessions for `/resume` inspection. No-session parents and `ultra.subagentSessionRetention = "none"` keep children entirely in memory. Both progress surfaces use the same phase-first hierarchy, authored step summaries, status language, icons, elapsed time, counts, checkpoint state, token usage, and cost. All agent rows keep task and activity summaries above indented metadata. Summary interpolation uses short object labels and array counts, while prompts retain their complete JSON inputs. Ultra wraps every fresh child assignment in an agent envelope containing workflow, phase, and `phase#index` identity. Interpolated prior-agent outputs are quoted with their immediate producer IDs and resolved task summaries, while workflow arguments and literal items remain unquoted. The overlay shows this exact delivered sub-agent prompt as the first user message in its transcript. Sub-agent tool calls keep one native Pi tool row from invocation through result and follow Pi's global expand state. Collapsed workflow tool output shows counts, reported cost, and output tokens with Pi's standard expand hint. Interactive completion also notifies the user with the status, phase count, and usage; headless runs stay silent. Expanded tools and overlays add the uncached-input, cache-read, and cache-write breakdown. Expanded tool output summarizes every phase, automatically opens the active and failed phases, shows their agents, and ends with the configured human-readable report when present. The command overlay uses the same board, automatically collapses completed phases, and adds keyboard interaction. Use `↑`/`↓` to focus visible phase or agent rows, `Enter` to toggle a phase, and `Tab` to cycle visible agents. Use `Alt+↑`/`Alt+↓` to scroll the focused transcript, `s` to steer, and `x` to abort the focused agent. Aborting the whole run requires two `Esc` presses within 1.5 seconds. Known agents appear queued before launch; dynamic fan-outs show `agents not planned yet`; conditional phases show `conditional` until their selectors resolve. Rows transition through queued, running, done, dropped, or skipped without disappearing from their phase summary. Configured tiers are shown as `tier → provider/model`. Completion notifications report cost and output tokens from completed sub-agent turns. Configured human reports do not include usage; notifications and persisted JSON retain usage and input/cache details. The final JSON contains `tokenUsage` with `input`, `output`, `total`, `cacheRead`, `cacheWrite`, and `cost`. Totals count every assistant response from launched steps exactly once, including result-submission reminders and failed or aborted attempts; reused checkpoints add no new usage. Every dropped step produces a typed diagnostic in `phaseFailures` with its code, message, retryability, attempt count, model, thinking level, original item, accumulated usage, and a bounded final transcript tail when available. Expanded boards show compact failure and retry metadata; the final JSON retains every diagnostic field. A later phase may consume `{PHASE.failures}` through `when`, `over`, prompts, or `return` to perform one explicit linear recovery pass. Recovery briefs retain the original goal, scope, constraints, instruction paths, evidence, tools, and schema; diagnostics supplement rather than replace that assignment. Ultra never drops work because of elapsed time and never creates automatic recovery loops. ## Debug Opt-in metadata diagnostics: [debug contract](../DEBUG.md). Safe events: `session.start`, `session.shutdown`, `command.handle.start`, `command.handle.finish`, `command.handle.error`. ## Settings Ultra reads one `ultra` block from Pi settings. Precedence is trusted project settings, then user settings (`~/.pi/agent/settings.json`), then built-in defaults. Ultra declares no flags or environment variables. Project settings are read only when Pi trusts the project; untrusted project settings are ignored. When both project and user define `ultra`, the blocks deep-merge with project winning per key; arrays and scalars replace. Model tiers therefore merge per field, and a project may partially override a tier inherited from defaults or user settings. `ultra.fanoutToolAllowlist` from the winning source extends the built-in core tools; a project list replaces a user list. Invalid values, wrong types, and unknown keys make the whole block invalid and never fall back to a lower source. Invalid settings surface as extension errors on session start, each turn, `/ultra`, and `run_workflow`. `ultra.auto_enable` was renamed to `ultra.autoEnable` and is rejected as an unknown key. | Key | Default | Meaning | | ----------------------------- | ------- | ---------------------------------------------------------- | | `ultra.enabled` | `true` | When false, the tool + command refuse to run. | | `ultra.autoEnable` | `false` | Enable `run_workflow` when each session starts. | | `ultra.concurrency` | `5` | Max concurrent sub-agents in a `fanout` phase. | | `ultra.maxRetries` | `2` | Result-submission reminders after a missing output. | | `ultra.subagentExtensions` | `"all"` | Load all, none, or a static extension-name list. | | `ultra.fanoutToolAllowlist` | core tools | Additional initial tool names allowed in fanout without `writeIsolation`. | | `ultra.subagentSkills` | `"all"` | Load host skills into sub-agents, or not. | | `ultra.subagentSessionRetention` | `"all"` | Retain parent-linked child sessions; `"none"` is stateless. | | `ultra.modelTiers` | built-in tiers | Named model, thinking, and orchestration guidance tiers. | | `ultra.providerModelTiers` | `{}` | Per-active-provider tier overrides folded over `ultra.modelTiers`. | | `ultra.actionSummarizerModel` | `null` | Fast model for concise tool-action phrases; off when null. | `ultra.actionSummarizerModel` accepts a concrete `provider/model` ref. Ultra immediately renders a terse tool-aware action such as `Reading settings.ts`, then replaces it with the best-effort LLM phrase when available. Collapsed transcripts hide redundant tool previews; expand tools to inspect their complete native Pi arguments and results. Requests use bounded input, 40 output tokens, no retries, no prompt caching, and a two-second timeout. Rapid actions coalesce independently per agent, only one request runs at once, and an agent's final pending action flushes when it finishes. Workflow completion cancels leftover requests. Failures keep the deterministic fallback and never affect workflow execution. Unknown tool arguments and results use bounded human-readable summaries; expanding shows labeled fields and nested lists. Source files, exact prompts, prose, and incomplete JSON text remain verbatim. Tool arguments are sent to the configured model without secret redaction. `ultra.modelTiers` is a map from arbitrary tier names matching `/^[a-z][a-z0-9_-]*$/` to objects with all three fields: `model`, `thinkingLevel`, and `instructions`. Names that are Object.prototype property names are forbidden. `model` is either a concrete `provider/model` reference or `null`, which uses the session model. `thinkingLevel` supplies the tier's default Pi reasoning budget; `instructions` supply its orchestrator-only selection guidance. Tier values are objects only: string aliases and `:thinking` suffixes are not supported. A new tier must provide all three fields; user and project settings may partially override an existing tier field by field. Validate a settings file with `mise run //extensions/ultra:check-settings -- /path/to/settings.json`. Built-in model tier defaults live in [`settings.ts`](settings.ts). When `run_workflow` is active, Ultra injects one effective rubric generated from the merged tiers into the orchestrator system prompt. Tier instructions guide the orchestrator only and are never passed to child agents as model-specific instructions. A step's `model` may name any configured tier or give a concrete `provider/model` reference. An unknown bare tier name is an error; a concrete provider/model reference remains valid. Omitting `model`, or resolving a tier whose `model` is `null`, uses the session model. An explicit step `thinkingLevel` overrides the tier default. Saved workflows retain their explicit tier selections when settings change. Changing effective model or thinking routes starts a fresh checkpoint journal instead of resuming children with stale routes. Changes only to the tier's orchestrator-facing selection instructions do not invalidate execution checkpoints. `ultra.providerModelTiers` maps provider names to per-tier field overrides that Ultra folds over `ultra.modelTiers` whenever the session model's provider matches. Provider names are matched exactly against the provider part of a `provider/model` reference. Overrides may set any of the three tier fields; tiers without an entry keep the base values, and providers without a block keep the base tiers. Provider blocks may only override tiers that exist in the merged base `ultra.modelTiers`; an unknown tier name is a settings error. The orchestrator rubric and all tier resolution use the folded tiers, so switching the session model's provider re-routes affected tiers on the next turn. For example, keep tier-neutral defaults and give each provider its own models: ```json { "ultra": { "providerModelTiers": { "openai-codex": { "tiny": { "model": "openai-codex/gpt-5.6-luna" }, "medium": { "model": "openai-codex/gpt-5.6-terra" }, "large": { "model": "openai-codex/gpt-6-sol" }, "huge": { "model": "openai-codex/gpt-6-astra" } }, "zai": { "tiny": { "model": "zai/glm-4.7-flash", "thinkingLevel": "low" }, "medium": { "model": "zai/glm-4.7" }, "large": { "model": "zai/glm-4.7", "thinkingLevel": "max" } } } } } ``` For example, override one built-in tier and add a custom one: ```json { "ultra": { "modelTiers": { "large": { "model": "openai-codex/gpt-6-astra", "instructions": "Architecture and ambiguous debugging only." }, "extractor": { "model": "openai-codex/gpt-5.6-luna", "thinkingLevel": "low", "instructions": "Literal extraction from supplied context." } } } } ``` `ultra.subagentExtensions` accepts `"all"` (default), `"none"`, or a static array of sibling pi-ext extension names. An array loads every named extension into every child; Ultra performs no per-model or per-tool selection. An empty array is equivalent to `"none"`. Unknown names fail workflow startup. A step's `tools` array selects its initial active tools; loaded extensions may register and activate additional tools through their hooks. Extension selection is therefore the trust boundary, including for fanout mutation safety. Children inherit the parent's CLI extension flags, and a step's `flags` object sets more, such as `{ "nushell": true }` to enable tools an extension keeps off by default. A step fails when a listed tool is still inactive after its extensions start; the error names the owning extension's flags, which the orchestrator also sees in the `ultra sub-agent flags` system section. See [extension flags](prompts/references/spec-format.md#extension-flags). `ultra.fanoutToolAllowlist` adds names to the default `read`, `grep`, `find`, `ls`, and `bash` list across global and trusted project settings. For example, `"fanoutToolAllowlist": ["web"]` permits `web` as an initial tool without `writeIsolation`; it does not make `web` available to children or prove the tool cannot write. Ultra advertises the effective list in the orchestrator's system policy while `run_workflow` is active. Tools outside that list require a genuine disjoint-ownership declaration in `writeIsolation`; the declaration is not a sandbox. Ultra strips its own `run_workflow` tool from the initial set to prevent ordinary recursive orchestration. Steps may name published catalog identities through `dynamicExtensions`; those load after normal extensions from exact run-local hash snapshots. Builder steps select all four dynamic-extension tools and receive fast-feedback guidance, while consumers receive none of it. The validator checks structural safety and tries a bounded Pi load and authored tests, reporting failures without blocking eligible publication. Publication proves artifact identity, not functionality; agents must repair reported load/test failures before consumption. Dynamic children are agent-owned artifacts until a human deliberately promotes them into the repository extension contract. No debug file or debug log path is presumed for a child. Ultra refuses to start a workflow when a listed tool is unavailable to its selected extensions. Normal child context-file loading is enabled, including Pi's global and cwd/ancestor context files. Builder guidance appends to discovered user system instructions rather than replacing `APPEND_SYSTEM.md`. Children do not inherit the parent conversation or parent chat history. Ultra identifies each child and marks prior-agent output inserted through interpolation, but this metadata does not replace task context. Authored prompts still supply task-specific parent constraints and applicable nested instruction paths, and require children to read those instructions before acting. Skill and reference contents remain on demand; normal context loading is not wholesale parent history inheritance. Available tools and extensions provide host capability, not permission to exceed the assignment's authorized scope. `ultra.subagentSkills` (`"all"` | `"none"`, default `"all"`) controls skill discovery for workflow sub-agents. Skills use Pi's normal global, project, package, and settings discovery, plus resources from every loaded extension. `"all"` also grants the built-in `read` tool so the agent can load selected `SKILL.md` files; every other tool remains workflow-gated. `"none"` omits all skills. ## Install / load This extension is loaded through the root pi-ext package. See [../../README.md](../../README.md).