--- id: PX-RESEARCH-50EAC989 type: research title: Ultra Workflow Evolution --- ## Scope This records durable Ultra runtime and authoring decisions from legacy material against present `extensions/ultra` behavior. Legacy sources are recoverable from Git commit `88660125c313dd5eff4f8c4dee431caa8e0ddd22`. ## Historical architecture Ultra was conceived as an in-process declarative JSON workflow runtime with ordered `single` and concurrency-capped `fanout` phases, structured child output, and aggregation. Inline one-off execution and named saved workflows were intentionally two invocations of one engine rather than separate orchestration systems. Closed interpolation replaced expression evaluation, making phase dataflow inspectable and rejecting unsupported syntax rather than executing arbitrary code. Phase ordering remains necessary because conditions and fan-out inputs may depend on earlier results. Early fan-out guidance identified shared-cwd concurrent mutation as a race, leading to a read-only default and explicit disjoint ownership when stronger tools are selected. ## Verified current runtime behavior Current Ultra exposes `/ultra`, `/ultra exec`, `/ultra run`, and the model-callable `run_workflow` surface described by `extensions/ultra/README.md`. Workflow discovery merges bundled, global, and project files with later sources overriding equal names, and rescans without reload. The runtime validates saved and inline specifications through one schema plus semantic checks, including interpolation-safe unique phase IDs and reserved `args` rejection. The active tool exposes the inline schema, while canonical authoring guidance, runtime contracts, and tests are kept semantically synchronized. Children receive explicit assignment envelopes and normal context-file discovery but not parent conversation history. Workflow output stores one complete aggregate, bounds the model-facing envelope, and preserves a path for selective evidence retrieval. Incomplete runs use checkpoints and named child sessions when retention permits, while changed assignment, provenance, schema, model, or thinking route invalidates stale work. ## Execution safety and recovery The historical extension-loading spike established that available extensions are capability supply, while each step's tool selection remains the activation gate and Ultra excludes its own workflow tool to prevent ordinary recursion. Current settings support all, none, or a static sibling-extension list, so extension selection remains a trust boundary rather than proof that a tool is read-only. Ultra-owned elapsed-time deadlines were removed, leaving external cancellation, provider failures, missing output, schema failures, and related typed diagnostics as drop causes. Dropped work carries bounded typed evidence, and later phases may explicitly consume prior failures for one linear recovery pass. No automatic recovery loop, back-edge, or timeout-driven retry is implied by that mechanism. ## Visibility and interaction Historical UI work separated the idle command overlay from mid-turn tool rendering because focus-grabbing interaction belongs only on the command path. Current progress is phase-first and shows known queued work, unresolved dynamic or conditional future work, active agents, terminal states, elapsed time, usage, and cost. Planning visibility must not speculate about selectors, alter scheduling, or fabricate agents whose existence depends on earlier results. The compact tool header combines brand, workflow identity, elapsed time, phase summary, and counts while expanded views retain evidence and agent detail. Human steering and abort controls belong to the interactive command surface, while model-tool progress remains read-only. ## Lessons A workflow DSL needs one runtime truth, one exposed schema, and concise semantic guidance because structural validity alone does not explain dataflow, ownership, or recovery. Concurrency policy, child-context boundaries, tool activation, and failure propagation are core correctness contracts rather than presentation details. Progress state should distinguish planned, queued, conditional, running, completed, dropped, and skipped work without claiming completion quality. ## Legacy sources `docs/super/specs/2026-06-27-ultra-workflow-engine-design.md` records original engine, isolation, two-surface, cancellation, and observability rationale. `docs/super/specs/2026-06-29-ultra-subagent-extensions-design.md` and `docs/super/plans/2026-06-29-ultra-subagent-extensions.md` record extension loading, allowlist gating, and recursion defense. `docs/super/specs/2026-07-13-ultra-infinite-step-runtime-design.md` and `docs/super/plans/2026-07-13-ultra-infinite-step-runtime.md` record removal of Ultra-owned timeouts. `docs/super/specs/2026-07-13-ultra-actionable-failures-design.md` and `docs/super/plans/2026-07-13-ultra-actionable-failures.md` record typed failures and explicit linear recovery. `docs/super/specs/2026-07-13-ultra-planned-work-visibility-design.md` and `docs/super/plans/2026-07-13-ultra-planned-work-visibility.md` record planned-work states and ordered ledger behavior. `docs/super/specs/2026-07-16-ultra-tool-header-design.md` and `docs/super/plans/2026-07-16-ultra-tool-header.md` record compact tool-header rationale. `docs/super/specs/2026-07-17-ultra-tool-schema-design.md`, `docs/super/plans/2026-07-17-ultra-tool-schema.md`, `docs/super/specs/2026-07-17-ultra-tool-guidance-design.md`, and `docs/super/plans/2026-07-17-ultra-tool-guidance.md` record schema exposure, semantic guidance, and synchronization requirements.