parseWorkflow uses the runtime validator to check it structurally and reject malformed specs.
this file is the authoritative format reference; the extension README.md links here rather than restating it.
machine-readable contract
spec.ts defines the WorkflowSpecSchema used by the active run_workflow tool and by parseWorkflow() for saved and inline specs.
that schema is the sole source for field names, requiredness, types, unions, and concise field hints; consult it for exact structure.
parseWorkflow() adds semantic checks; this reference documents execution semantics, authoring rules, and valid examples without maintaining a second schema.
args may declare params, an ordered list of input names.
/ultra run <name> tokA tokB maps tokens positionally; the last declared parameter absorbs remaining tokens as free text.
values are exposed as {args.NAME}.
without params, pass inputs as a trailing JSON object: /ultra run <name> {"k": "v"}.
a missing {args.NAME} renders to the empty string.
orchestrator output
return selects the final value; without it, the last phase's results are used.
report optionally selects the human-facing value; without it, human rendering uses the existing return-based formatting.
tool and command results store the full aggregate once; every bounded model-facing envelope includes fullOutputPath.
the selected return value and configured report value appear in the bounded model-facing envelope when they fit, with phase counts, bounded failure diagnostics, and usage; phaseResults is not duplicated.
the selected report value is also rendered on human-facing surfaces through Pi's native Markdown/structured renderer, so the agent and maintainer receive the same selected report data.
executionStatus is finished, incomplete for dropped steps, or aborted.
finished describes execution only, not artifact quality or semantic success.
oversized values use a retrieval notice within the 50KB or 2000 lines limit.
complete contents, including intermediate results and failure inputs/transcript tails, remain in UI details.
read that file selectively when needed instead of loading all intermediate work into context.
phases
a phase has an id, kind, step, optional concurrency, optional when, and optional fan-out writeIsolation.
id is referenced by interpolation, return, and report.
phase IDs must match [A-Za-z_][A-Za-z0-9_]* and be unique.
the reserved phase ID args is forbidden because it belongs to the workflow-input namespace.
single — runs the step once for synthesis, judging, or aggregation over prior results; over is accepted but ignored in v1.
fanout — runs the step once per over element, concurrently within concurrency and the engine's global pool; over is required.
when uses the existing single-selector grammar.
false, null, undefined, an empty string, or an empty array skips the phase and records { ran: 0, ok: 0, dropped: 0 } with empty results.
any other value runs the phase.
a skipped phase remains addressable through empty {PHASE.results} and {PHASE.failures} arrays.
there are no loops or arbitrary expressions.
over is a literal array or one interpolation token string, such as "{review.results[].findings[]}".
return and report use the same single-selector grammar and resolve after all phases.
fan-out mutation ownership
without writeIsolation, a fanout step's initial tools must be in the effective ultra fanout tool policy system section (defaults: read, grep, find, ls, bash; ultra.fanoutToolAllowlist adds names).
the prompt must instruct the agent not to mutate anything.
other initial tools require writeIsolation with a genuine non-empty ownership rationale assigning disjoint work to every item; do not invent ownership just to pass validation.
the allowlist classifies tool names, not capabilities; Ultra trusts isolation declarations and does not sandbox paths or arbitrary custom mutators.
loaded extensions remain trusted to activate additional tools.
every listed tool must be available to sub-agents or Ultra refuses to start the workflow.
step semantics
schema-declared step fields and types are available through the active tool contract.
summary is rendered on tool and overlay agent rows.
prompt is the child assignment body before Ultra adds runtime identity and provenance markup.
model, tools, thinking-level, dynamic-extension, and output-schema behavior below describes runtime semantics not captured by field shape alone.
use the same interpolation grammar as prompt so fan-out rows can identify their item.
summary interpolation renders objects as short labels and arrays as counts; prompts retain compact JSON.
Ultra wraps every fresh assignment with workflow, phase, and phase#index identity.
prior-agent outputs inserted through interpolation are quoted with immediate producer identity and the producer's resolved task summary.
workflow arguments, literal items, failure records, and operator steering are not attributed to sub-agents.
workflow authors must not add manual identity or provenance wrappers.
the overlay shows the exact delivered envelope as the first user message in that agent's transcript.
normal child context-file loading is enabled; children do not inherit the parent conversation.
prompts must supply the task context, constraints, instruction paths, evidence, and authority boundaries each child needs.
every step returns one JSON object through the terminating structured_output tool.
Ultra requests provider-side strict schema sampling when supported and always validates the submitted object locally.
when an agent stops without submitting, Ultra sends up to ultra.maxRetries focused reminders in the same session.
without schema, any JSON object is accepted.
schema validates that object and is required when downstream phases depend on a stable shape.
a string schema is a name reference and must resolve in the spec's schemas map.
configured model tiers live in Pi settings under ultra.modelTiers.
tier names are resolved dynamically, so a step may select any configured tier name.
an unknown bare tier name is invalid; a concrete provider/model reference is always valid.
omitting model, or selecting a tier whose model is null, uses the session model.
an explicit thinkingLevel overrides the tier default.
extension flags
children inherit the parent Pi's CLI extension flags, applied the way Pi's CLI applies them: only flags a child extension registers.
flags sets Pi extension flags for one step, keyed by flag name without dashes, such as { "nushell": true }; step values override inherited ones.
a step flag must be registered by a loaded child extension and match its boolean or string type, or the step fails before prompting.
some extensions register tools but keep them off unless a flag enables them at startup.
when run_workflow is active, the ultra sub-agent flags system section lists each child extension that owns tools and flags, with flag descriptions.
a step whose listed tool is still inactive after startup fails with a session error naming the owning extension and its flags; Ultra never force-activates tools.
dynamic extensions and builders
dynamicExtensions names stable catalog identities, never hashes.
Ultra checks the published content hash and copies an exact run-local snapshot before child creation; historical validator metadata does not expire extensions.
attestation proves identity, not successful loading or test quality; inspect builder diagnostics before use.
normal configured extensions load first, then dynamic extensions in declared order.
a builder step selects all four tools or none: search_dynamic_extensions, create_dynamic_extension, copy_dynamic_extension, and validate_dynamic_extension.
builders receive every available child tool except run_workflow plus the builder contract.
known canonical names may load directly; discovery, creation, improvement, stale validation, and repair require a builder.
a builder failure blocks every later phase with typed dynamic_extension_builder_failure outcomes.
model mechanics
when run_workflow is active, Ultra injects the effective tier rubric generated from merged ultra.modelTiers settings into the orchestrator system prompt.
a step's model names a configured tier or a concrete provider/model reference.
an unknown bare tier name is invalid; omitting model uses the session model.
a tier supplies a default thinkingLevel; an explicit step value overrides it.
tier instructions guide the orchestrator and are not passed to child agents.
interpolation grammar
interpolation is a closed mini-syntax, not an expression evaluator, and has no eval.
braces below are literal and shown without JSON quoting:
{args.NAME} workflow input value
{item} current fanout item (string verbatim; else compact JSON)
{PHASE.results} a prior phase's results array (nulls retained positionally)
{PHASE.failures} a prior phase's typed failure array
{PHASE.results[].FIELD} project each result's FIELD (null/missing skipped)
{PHASE.results[].FIELD[]} flatten: concat each result's FIELD array (nulls/missing skipped)
{EXPR | where FIELD} keep elements whose boolean FIELD is truthy
{EXPR | where FIELD == "VALUE"} keep elements whose string FIELD exactly equals VALUE
{EXPR | groupBy FIELD} group records by exact string FIELD value
{PHASE.results[].FIELD} projects one field from each record into a new array, preserving order; null, missing, and non-record values are skipped, while false, zero, empty strings, objects, and arrays are retained.
where FIELD keeps records whose field is truthy.
where FIELD == "VALUE" keeps records whose field is a string exactly equal to VALUE; missing, null, and non-string fields do not match.
groupBy FIELD returns first-seen groups shaped as { "key": "...", "items": [...] }; records with missing, null, or non-string fields are skipped.
group item order and group order follow the source array.
the left side of where and groupBy must resolve to an array.
anything outside this grammar raises a validation error at run time.
in a step prompt, literal braces are allowed only as these tokens.
describe inline JSON shapes in words and use a schema instead.
explicit linear recovery
a dropped step appears in {PHASE.failures} with agentId, index, optional original item, code, message, retryable, attempts, optional model, optional thinkingLevel, optional accumulated usage, and an optional 2,000-character transcriptTail.
failure codes are auth, transport, schema, missing-output, aborted, session, or dynamic_extension_builder_failure.
use a non-empty failure array to condition and fan out one later recovery phase.
recovery must restate the original goal, scope, constraints, instruction paths, required evidence, tools, and output schema; failure diagnostics supplement the assignment, never replace it.
the example below reads only the root package.json, then recovers and selects an outcome that preserves blocked work.
supply instructions as a run input containing applicable parent constraints and instruction paths, or explicitly state that no additional task-specific instructions apply.
no child may infer a license from repository conventions or change files to make the task succeed.
Ultra has no timeout, automatic rerun, loop, or back-edge.
ultra.maxRetries sends result-submission reminders only after an agent turn ends without output; it never reruns the work in a fresh session.
workflow authors explicitly place any recovery phase after the failed phase.
schema join-keys
when one phase fans out over another's findings, require a stable unique id in both schemas and instruct the consumer to echo it verbatim.
use the same ID-based join as the bundled review workflow: a review dimension, a colon, and a locally unique identifier containing only letters, digits, underscores, or hyphens, such as correctness:1.
keep IDs stable through verification and distinguish multiple findings in one file.
file is a locator, not a unique join key; never merge or overwrite candidates merely because they share a file or title.
missing or duplicate verdict IDs remain unresolved work, not implicit refutations.
example gallery
Examples are split by authoring concern so a workflow can be copied without loading the whole gallery.
Each linked file contains complete JSON workflow specs and is covered by the reference tests.
core: args, single phases, fanout, conditions, and fanout inputs.
selectors: projection, flattening, truthy filtering, exact equality, and grouping.
schemas and routing: named/inline schemas, concurrency, model tiers, thinking, and provider/model references.
safety and results: read-only tools, write isolation, recovery, return, and human reports.
pipelines: composition, joins, mixed fanout, and complete worked workflows.
Use the smallest relevant file first.
Read dynamic-extension-builder.md for generated-extension implementation requirements.
# ultra workflow spec — canonical reference
- a workflow is a single JSON object.
- `parseWorkflow` uses the runtime validator to check it structurally and reject malformed specs.
- this file is the authoritative format reference; the extension `README.md` links here rather than restating it.
## machine-readable contract
- [`spec.ts`](../../spec.ts) defines the `WorkflowSpecSchema` used by the active `run_workflow` tool and by `parseWorkflow()` for saved and inline specs.
- that schema is the sole source for field names, requiredness, types, unions, and concise field hints; consult it for exact structure.
- `parseWorkflow()` adds semantic checks; this reference documents execution semantics, authoring rules, and valid examples without maintaining a second schema.
- `args` may declare `params`, an ordered list of input names.
- `/ultra run <name> tokA tokB` maps tokens positionally; the last declared parameter absorbs remaining tokens as free text.
- values are exposed as `{args.NAME}`.
- without `params`, pass inputs as a trailing JSON object: `/ultra run <name> {"k": "v"}`.
- a missing `{args.NAME}` renders to the empty string.
## orchestrator output
- `return` selects the final value; without it, the last phase's results are used.
- `report` optionally selects the human-facing value; without it, human rendering uses the existing `return`-based formatting.
- tool and command results store the full aggregate once; every bounded model-facing envelope includes `fullOutputPath`.
- the selected `return` value and configured `report` value appear in the bounded model-facing envelope when they fit, with phase counts, bounded failure diagnostics, and usage; `phaseResults` is not duplicated.
- the selected `report` value is also rendered on human-facing surfaces through Pi's native Markdown/structured renderer, so the agent and maintainer receive the same selected report data.
- `executionStatus` is `finished`, `incomplete` for dropped steps, or `aborted`.
- `finished` describes execution only, not artifact quality or semantic success.
- oversized values use a retrieval notice within the 50KB or 2000 lines limit.
- complete contents, including intermediate results and failure inputs/transcript tails, remain in UI details.
- read that file selectively when needed instead of loading all intermediate work into context.
## phases
- a phase has an `id`, `kind`, `step`, optional `concurrency`, optional `when`, and optional fan-out `writeIsolation`.
- `id` is referenced by interpolation, `return`, and `report`.
- phase IDs must match `[A-Za-z_][A-Za-z0-9_]*` and be unique.
- the reserved phase ID `args` is forbidden because it belongs to the workflow-input namespace.
- **`single`** — runs the step once for synthesis, judging, or aggregation over prior results; `over` is accepted but ignored in v1.
- **`fanout`** — runs the step once per `over` element, concurrently within `concurrency` and the engine's global pool; `over` is required.
- `when` uses the existing single-selector grammar.
- `false`, `null`, `undefined`, an empty string, or an empty array skips the phase and records `{ ran: 0, ok: 0, dropped: 0 }` with empty results.
- any other value runs the phase.
- a skipped phase remains addressable through empty `{PHASE.results}` and `{PHASE.failures}` arrays.
- there are no loops or arbitrary expressions.
- `over` is a literal array or one interpolation token string, such as `"{review.results[].findings[]}"`.
- `return` and `report` use the same single-selector grammar and resolve after all phases.
### fan-out mutation ownership
- without `writeIsolation`, a `fanout` step's initial tools must be in the effective `ultra fanout tool policy` system section (defaults: `read`, `grep`, `find`, `ls`, `bash`; `ultra.fanoutToolAllowlist` adds names).
- the prompt must instruct the agent not to mutate anything.
- other initial tools require `writeIsolation` with a genuine non-empty ownership rationale assigning disjoint work to every item; do not invent ownership just to pass validation.
- the allowlist classifies tool names, not capabilities; Ultra trusts isolation declarations and does not sandbox paths or arbitrary custom mutators.
- loaded extensions remain trusted to activate additional tools.
- every listed tool must be available to sub-agents or Ultra refuses to start the workflow.
## step semantics
- schema-declared step fields and types are available through the active tool contract.
- `summary` is rendered on tool and overlay agent rows.
- `prompt` is the child assignment body before Ultra adds runtime identity and provenance markup.
- model, tools, thinking-level, dynamic-extension, and output-schema behavior below describes runtime semantics not captured by field shape alone.
- use the same interpolation grammar as `prompt` so fan-out rows can identify their item.
- summary interpolation renders objects as short labels and arrays as counts; prompts retain compact JSON.
- Ultra wraps every fresh assignment with workflow, phase, and `phase#index` identity.
- prior-agent outputs inserted through interpolation are quoted with immediate producer identity and the producer's resolved task summary.
- workflow arguments, literal items, failure records, and operator steering are not attributed to sub-agents.
- workflow authors must not add manual identity or provenance wrappers.
- the overlay shows the exact delivered envelope as the first user message in that agent's transcript.
- normal child context-file loading is enabled; children do not inherit the parent conversation.
- prompts must supply the task context, constraints, instruction paths, evidence, and authority boundaries each child needs.
- see [child context mechanics](../ultra-authoring/SKILL.md#give-children-context).
- every step returns one JSON object through the terminating `structured_output` tool.
- Ultra requests provider-side strict schema sampling when supported and always validates the submitted object locally.
- when an agent stops without submitting, Ultra sends up to `ultra.maxRetries` focused reminders in the same session.
- without `schema`, any JSON object is accepted.
- `schema` validates that object and is required when downstream phases depend on a stable shape.
- a string `schema` is a name reference and must resolve in the spec's `schemas` map.
- configured model tiers live in Pi settings under `ultra.modelTiers`.
- tier names are resolved dynamically, so a step may select any configured tier name.
- an unknown bare tier name is invalid; a concrete `provider/model` reference is always valid.
- omitting `model`, or selecting a tier whose `model` is `null`, uses the session model.
- an explicit `thinkingLevel` overrides the tier default.
### extension flags
- children inherit the parent Pi's CLI extension flags, applied the way Pi's CLI applies them: only flags a child extension registers.
- `flags` sets Pi extension flags for one step, keyed by flag name without dashes, such as `{ "nushell": true }`; step values override inherited ones.
- a step flag must be registered by a loaded child extension and match its boolean or string type, or the step fails before prompting.
- some extensions register tools but keep them off unless a flag enables them at startup.
- when `run_workflow` is active, the `ultra sub-agent flags` system section lists each child extension that owns tools and flags, with flag descriptions.
- a step whose listed tool is still inactive after startup fails with a `session` error naming the owning extension and its flags; Ultra never force-activates tools.
### dynamic extensions and builders
- `dynamicExtensions` names stable catalog identities, never hashes.
- Ultra checks the published content hash and copies an exact run-local snapshot before child creation; historical validator metadata does not expire extensions.
- attestation proves identity, not successful loading or test quality; inspect builder diagnostics before use.
- normal configured extensions load first, then dynamic extensions in declared order.
- a builder step selects all four tools or none: `search_dynamic_extensions`, `create_dynamic_extension`, `copy_dynamic_extension`, and `validate_dynamic_extension`.
- builders receive every available child tool except `run_workflow` plus the [builder contract](../dynamic-extension-builder.md).
- known canonical names may load directly; discovery, creation, improvement, stale validation, and repair require a builder.
- a builder failure blocks every later phase with typed `dynamic_extension_builder_failure` outcomes.
### model mechanics
- when `run_workflow` is active, Ultra injects the effective tier rubric generated from merged `ultra.modelTiers` settings into the orchestrator system prompt.
- a step's `model` names a configured tier or a concrete `provider/model` reference.
- an unknown bare tier name is invalid; omitting `model` uses the session model.
- a tier supplies a default `thinkingLevel`; an explicit step value overrides it.
- tier instructions guide the orchestrator and are not passed to child agents.
## interpolation grammar
- interpolation is a closed mini-syntax, not an expression evaluator, and has no `eval`.
- braces below are literal and shown without JSON quoting:
```text
{args.NAME} workflow input value
{item} current fanout item (string verbatim; else compact JSON)
{PHASE.results} a prior phase's results array (nulls retained positionally)
{PHASE.failures} a prior phase's typed failure array
{PHASE.results[].FIELD} project each result's FIELD (null/missing skipped)
{PHASE.results[].FIELD[]} flatten: concat each result's FIELD array (nulls/missing skipped)
{EXPR | where FIELD} keep elements whose boolean FIELD is truthy
{EXPR | where FIELD == "VALUE"} keep elements whose string FIELD exactly equals VALUE
{EXPR | groupBy FIELD} group records by exact string FIELD value
```
- `{PHASE.results[].FIELD}` projects one field from each record into a new array, preserving order; null, missing, and non-record values are skipped, while false, zero, empty strings, objects, and arrays are retained.
- `where FIELD` keeps records whose field is truthy.
- `where FIELD == "VALUE"` keeps records whose field is a string exactly equal to VALUE; missing, null, and non-string fields do not match.
- `groupBy FIELD` returns first-seen groups shaped as `{ "key": "...", "items": [...] }`; records with missing, null, or non-string fields are skipped.
- group item order and group order follow the source array.
- the left side of `where` and `groupBy` must resolve to an array.
- anything outside this grammar raises a validation error at run time.
- in a step `prompt`, literal braces are allowed only as these tokens.
- describe inline JSON shapes in words and use a `schema` instead.
### explicit linear recovery
- a dropped step appears in `{PHASE.failures}` with `agentId`, `index`, optional original `item`, `code`, `message`, `retryable`, `attempts`, optional `model`, optional `thinkingLevel`, optional accumulated `usage`, and an optional 2,000-character `transcriptTail`.
- failure codes are `auth`, `transport`, `schema`, `missing-output`, `aborted`, `session`, or `dynamic_extension_builder_failure`.
- use a non-empty failure array to condition and fan out one later recovery phase.
- recovery must restate the original goal, scope, constraints, instruction paths, required evidence, tools, and output schema; failure diagnostics supplement the assignment, never replace it.
- the example below reads only the root `package.json`, then recovers and selects an outcome that preserves blocked work.
- supply `instructions` as a run input containing applicable parent constraints and instruction paths, or explicitly state that no additional task-specific instructions apply.
- no child may infer a license from repository conventions or change files to make the task succeed.
- complete example: [read-package-license](examples/safety-and-results.md#explicit-linear-recovery).
- Ultra has no timeout, automatic rerun, loop, or back-edge.
- `ultra.maxRetries` sends result-submission reminders only after an agent turn ends without output; it never reruns the work in a fresh session.
- workflow authors explicitly place any recovery phase after the failed phase.
### schema join-keys
- when one phase fans out over another's findings, require a stable unique `id` in both schemas and instruct the consumer to echo it verbatim.
- use the same ID-based join as the bundled review workflow: a review dimension, a colon, and a locally unique identifier containing only letters, digits, underscores, or hyphens, such as `correctness:1`.
- keep IDs stable through verification and distinguish multiple findings in one file.
- `file` is a locator, not a unique join key; never merge or overwrite candidates merely because they share a file or title.
- missing or duplicate verdict IDs remain unresolved work, not implicit refutations.
## example gallery
Examples are split by authoring concern so a workflow can be copied without loading the whole gallery.
Each linked file contains complete JSON workflow specs and is covered by the reference tests.
- [core](examples/core.md): args, single phases, fanout, conditions, and fanout inputs.
- [selectors](examples/selectors.md): projection, flattening, truthy filtering, exact equality, and grouping.
- [schemas and routing](examples/schemas-and-routing.md): named/inline schemas, concurrency, model tiers, thinking, and provider/model references.
- [safety and results](examples/safety-and-results.md): read-only tools, write isolation, recovery, return, and human reports.
- [dynamic extensions](examples/dynamic-extensions.md): stable-name consumers and working-copy builders.
- [pipelines](examples/pipelines.md): composition, joins, mixed fanout, and complete worked workflows.
Use the smallest relevant file first.
Read [dynamic-extension-builder.md](../dynamic-extension-builder.md) for generated-extension implementation requirements.