# ultra workflow spec — canonical reference - a workflow is a single JSON object. - `parseWorkflow` uses the runtime validator to check it structurally and reject malformed specs. - this file is the authoritative format reference; the extension `README.md` links here rather than restating it. ## machine-readable contract - [`spec.ts`](../../spec.ts) defines the `WorkflowSpecSchema` used by the active `run_workflow` tool and by `parseWorkflow()` for saved and inline specs. - that schema is the sole source for field names, requiredness, types, unions, and concise field hints; consult it for exact structure. - `parseWorkflow()` adds semantic checks; this reference documents execution semantics, authoring rules, and valid examples without maintaining a second schema. - `args` may declare `params`, an ordered list of input names. - `/ultra run tokA tokB` maps tokens positionally; the last declared parameter absorbs remaining tokens as free text. - values are exposed as `{args.NAME}`. - without `params`, pass inputs as a trailing JSON object: `/ultra run {"k": "v"}`. - a missing `{args.NAME}` renders to the empty string. ## orchestrator output - `return` selects the final value; without it, the last phase's results are used. - `report` optionally selects the human-facing value; without it, human rendering uses the existing `return`-based formatting. - tool and command results store the full aggregate once; every bounded model-facing envelope includes `fullOutputPath`. - the selected `return` value and configured `report` value appear in the bounded model-facing envelope when they fit, with phase counts, bounded failure diagnostics, and usage; `phaseResults` is not duplicated. - the selected `report` value is also rendered on human-facing surfaces through Pi's native Markdown/structured renderer, so the agent and maintainer receive the same selected report data. - `executionStatus` is `finished`, `incomplete` for dropped steps, or `aborted`. - `finished` describes execution only, not artifact quality or semantic success. - oversized values use a retrieval notice within the 50KB or 2000 lines limit. - complete contents, including intermediate results and failure inputs/transcript tails, remain in UI details. - read that file selectively when needed instead of loading all intermediate work into context. ## phases - a phase has an `id`, `kind`, `step`, optional `concurrency`, optional `when`, and optional fan-out `writeIsolation`. - `id` is referenced by interpolation, `return`, and `report`. - phase IDs must match `[A-Za-z_][A-Za-z0-9_]*` and be unique. - the reserved phase ID `args` is forbidden because it belongs to the workflow-input namespace. - **`single`** — runs the step once for synthesis, judging, or aggregation over prior results; `over` is accepted but ignored in v1. - **`fanout`** — runs the step once per `over` element, concurrently within `concurrency` and the engine's global pool; `over` is required. - `when` uses the existing single-selector grammar. - `false`, `null`, `undefined`, an empty string, or an empty array skips the phase and records `{ ran: 0, ok: 0, dropped: 0 }` with empty results. - any other value runs the phase. - a skipped phase remains addressable through empty `{PHASE.results}` and `{PHASE.failures}` arrays. - there are no loops or arbitrary expressions. - `over` is a literal array or one interpolation token string, such as `"{review.results[].findings[]}"`. - `return` and `report` use the same single-selector grammar and resolve after all phases. ### fan-out mutation ownership - without `writeIsolation`, a `fanout` step's initial tools must be in the effective `ultra fanout tool policy` system section (defaults: `read`, `grep`, `find`, `ls`, `bash`; `ultra.fanoutToolAllowlist` adds names). - the prompt must instruct the agent not to mutate anything. - other initial tools require `writeIsolation` with a genuine non-empty ownership rationale assigning disjoint work to every item; do not invent ownership just to pass validation. - the allowlist classifies tool names, not capabilities; Ultra trusts isolation declarations and does not sandbox paths or arbitrary custom mutators. - loaded extensions remain trusted to activate additional tools. - every listed tool must be available to sub-agents or Ultra refuses to start the workflow. ## step semantics - schema-declared step fields and types are available through the active tool contract. - `summary` is rendered on tool and overlay agent rows. - `prompt` is the child assignment body before Ultra adds runtime identity and provenance markup. - model, tools, thinking-level, dynamic-extension, and output-schema behavior below describes runtime semantics not captured by field shape alone. - use the same interpolation grammar as `prompt` so fan-out rows can identify their item. - summary interpolation renders objects as short labels and arrays as counts; prompts retain compact JSON. - Ultra wraps every fresh assignment with workflow, phase, and `phase#index` identity. - prior-agent outputs inserted through interpolation are quoted with immediate producer identity and the producer's resolved task summary. - workflow arguments, literal items, failure records, and operator steering are not attributed to sub-agents. - workflow authors must not add manual identity or provenance wrappers. - the overlay shows the exact delivered envelope as the first user message in that agent's transcript. - normal child context-file loading is enabled; children do not inherit the parent conversation. - prompts must supply the task context, constraints, instruction paths, evidence, and authority boundaries each child needs. - see [child context mechanics](../ultra-authoring/SKILL.md#give-children-context). - every step returns one JSON object through the terminating `structured_output` tool. - Ultra requests provider-side strict schema sampling when supported and always validates the submitted object locally. - when an agent stops without submitting, Ultra sends up to `ultra.maxRetries` focused reminders in the same session. - without `schema`, any JSON object is accepted. - `schema` validates that object and is required when downstream phases depend on a stable shape. - a string `schema` is a name reference and must resolve in the spec's `schemas` map. - configured model tiers live in Pi settings under `ultra.modelTiers`. - tier names are resolved dynamically, so a step may select any configured tier name. - an unknown bare tier name is invalid; a concrete `provider/model` reference is always valid. - omitting `model`, or selecting a tier whose `model` is `null`, uses the session model. - an explicit `thinkingLevel` overrides the tier default. ### extension flags - children inherit the parent Pi's CLI extension flags, applied the way Pi's CLI applies them: only flags a child extension registers. - `flags` sets Pi extension flags for one step, keyed by flag name without dashes, such as `{ "nushell": true }`; step values override inherited ones. - a step flag must be registered by a loaded child extension and match its boolean or string type, or the step fails before prompting. - some extensions register tools but keep them off unless a flag enables them at startup. - when `run_workflow` is active, the `ultra sub-agent flags` system section lists each child extension that owns tools and flags, with flag descriptions. - a step whose listed tool is still inactive after startup fails with a `session` error naming the owning extension and its flags; Ultra never force-activates tools. ### dynamic extensions and builders - `dynamicExtensions` names stable catalog identities, never hashes. - Ultra checks the published content hash and copies an exact run-local snapshot before child creation; historical validator metadata does not expire extensions. - attestation proves identity, not successful loading or test quality; inspect builder diagnostics before use. - normal configured extensions load first, then dynamic extensions in declared order. - a builder step selects all four tools or none: `search_dynamic_extensions`, `create_dynamic_extension`, `copy_dynamic_extension`, and `validate_dynamic_extension`. - builders receive every available child tool except `run_workflow` plus the [builder contract](../dynamic-extension-builder.md). - known canonical names may load directly; discovery, creation, improvement, stale validation, and repair require a builder. - a builder failure blocks every later phase with typed `dynamic_extension_builder_failure` outcomes. ### model mechanics - when `run_workflow` is active, Ultra injects the effective tier rubric generated from merged `ultra.modelTiers` settings into the orchestrator system prompt. - a step's `model` names a configured tier or a concrete `provider/model` reference. - an unknown bare tier name is invalid; omitting `model` uses the session model. - a tier supplies a default `thinkingLevel`; an explicit step value overrides it. - tier instructions guide the orchestrator and are not passed to child agents. ## interpolation grammar - interpolation is a closed mini-syntax, not an expression evaluator, and has no `eval`. - braces below are literal and shown without JSON quoting: ```text {args.NAME} workflow input value {item} current fanout item (string verbatim; else compact JSON) {PHASE.results} a prior phase's results array (nulls retained positionally) {PHASE.failures} a prior phase's typed failure array {PHASE.results[].FIELD} project each result's FIELD (null/missing skipped) {PHASE.results[].FIELD[]} flatten: concat each result's FIELD array (nulls/missing skipped) {EXPR | where FIELD} keep elements whose boolean FIELD is truthy {EXPR | where FIELD == "VALUE"} keep elements whose string FIELD exactly equals VALUE {EXPR | groupBy FIELD} group records by exact string FIELD value ``` - `{PHASE.results[].FIELD}` projects one field from each record into a new array, preserving order; null, missing, and non-record values are skipped, while false, zero, empty strings, objects, and arrays are retained. - `where FIELD` keeps records whose field is truthy. - `where FIELD == "VALUE"` keeps records whose field is a string exactly equal to VALUE; missing, null, and non-string fields do not match. - `groupBy FIELD` returns first-seen groups shaped as `{ "key": "...", "items": [...] }`; records with missing, null, or non-string fields are skipped. - group item order and group order follow the source array. - the left side of `where` and `groupBy` must resolve to an array. - anything outside this grammar raises a validation error at run time. - in a step `prompt`, literal braces are allowed only as these tokens. - describe inline JSON shapes in words and use a `schema` instead. ### explicit linear recovery - a dropped step appears in `{PHASE.failures}` with `agentId`, `index`, optional original `item`, `code`, `message`, `retryable`, `attempts`, optional `model`, optional `thinkingLevel`, optional accumulated `usage`, and an optional 2,000-character `transcriptTail`. - failure codes are `auth`, `transport`, `schema`, `missing-output`, `aborted`, `session`, or `dynamic_extension_builder_failure`. - use a non-empty failure array to condition and fan out one later recovery phase. - recovery must restate the original goal, scope, constraints, instruction paths, required evidence, tools, and output schema; failure diagnostics supplement the assignment, never replace it. - the example below reads only the root `package.json`, then recovers and selects an outcome that preserves blocked work. - supply `instructions` as a run input containing applicable parent constraints and instruction paths, or explicitly state that no additional task-specific instructions apply. - no child may infer a license from repository conventions or change files to make the task succeed. - complete example: [read-package-license](examples/safety-and-results.md#explicit-linear-recovery). - Ultra has no timeout, automatic rerun, loop, or back-edge. - `ultra.maxRetries` sends result-submission reminders only after an agent turn ends without output; it never reruns the work in a fresh session. - workflow authors explicitly place any recovery phase after the failed phase. ### schema join-keys - when one phase fans out over another's findings, require a stable unique `id` in both schemas and instruct the consumer to echo it verbatim. - use the same ID-based join as the bundled review workflow: a review dimension, a colon, and a locally unique identifier containing only letters, digits, underscores, or hyphens, such as `correctness:1`. - keep IDs stable through verification and distinguish multiple findings in one file. - `file` is a locator, not a unique join key; never merge or overwrite candidates merely because they share a file or title. - missing or duplicate verdict IDs remain unresolved work, not implicit refutations. ## example gallery Examples are split by authoring concern so a workflow can be copied without loading the whole gallery. Each linked file contains complete JSON workflow specs and is covered by the reference tests. - [core](examples/core.md): args, single phases, fanout, conditions, and fanout inputs. - [selectors](examples/selectors.md): projection, flattening, truthy filtering, exact equality, and grouping. - [schemas and routing](examples/schemas-and-routing.md): named/inline schemas, concurrency, model tiers, thinking, and provider/model references. - [safety and results](examples/safety-and-results.md): read-only tools, write isolation, recovery, return, and human reports. - [dynamic extensions](examples/dynamic-extensions.md): stable-name consumers and working-copy builders. - [pipelines](examples/pipelines.md): composition, joins, mixed fanout, and complete worked workflows. Use the smallest relevant file first. Read [dynamic-extension-builder.md](../dynamic-extension-builder.md) for generated-extension implementation requirements.