Luigit
repositories / pi-ext

pi-ext

bugabingas pi extensions

owned by admin

extensions/angel/SPEC.md

Raw
Rendered preview

Angel: investigative advisor specification

Status and purpose

This specification records the agreed redesign of Angel. It defines the required result, not an implementation plan or a description of the current code. It is the binding contract for the authorized implementation.

Angel provides deep reasoning and independent investigation to improve the executor's next decision. Advice quality and verified recovery take priority over cost or consultation frequency. Angel is an advisor, not a permission system, command firewall, or implementation worker.

Consultation lifecycle

Each consultation runs a capable, tool-using Pi agent in a new persistent child session linked to the originating parent session. The advisor investigates until it can provide useful advice or identify what prevents a supported conclusion. The parent receives the final advice, not the advisor's intermediate investigation.

A consultation has three possible origins:

  • Executor: the model invokes the angel tool with a question or decision to investigate.
  • Human: the user invokes /angel <question> directly.
  • Error: Angel automatically investigates a completed executor tool batch containing a failed tool result.

All origins share the same investigation engine and advice presentation. Their instructions and delivery behavior differ as specified below.

Error origin

Automatic consultation is a deterministic stalled-recovery mechanism, not a response to every failed tool result. The first eligible operation failure is remembered without consulting. If the same opaque operation fails again in a later completed executor turn before the task boundary changes, Angel starts one consultation before the executor's next model call. A matching successful operation clears the remembered failure and makes a later failure a new incident. After one consultation, the same operation does not trigger again until it succeeds or the task boundary changes.

Operation identity is the tool name plus a cryptographic hash of canonically ordered structured arguments. Arguments are matched exactly after canonical ordering; Angel does not normalize tool-specific fields, inspect command text, classify error prose, or persist argument values. A result whose originating structured call cannot be mapped reliably is not eligible for automatic consultation. Duplicate failures for one operation in the same completed turn count as one attempt, not stalled recovery.

Tool provenance comes from Pi's active tool metadata. Known low-signal Pi built-ins read, edit, write, grep, find, and ls are excluded from automatic consultation because their ordinary recovery is local discovery or retry. This exclusion applies only when Pi identifies the actual registration as built-in; an extension override with the same name remains an opaque non-built-in tool. Built-in process tools such as bash and powershell, extension tools, SDK tools, future unknown tools, and tools with unavailable provenance use the same generic repeated-operation rule. Angel must not contain names or domain assumptions for unrelated extensions. The angel tool itself is always excluded to prevent recursion.

A triggering batch produces one consultation with all completed sibling results available, including successes and first-time failures. There is no every-Nth-error counter or session-wide error-review budget. Explicit cancellation is not a failure candidate. Remembered state is session-local and is cleared on new human input, session replacement, tree navigation, reload, or /angel on|off. It is intentionally not restored from transcripts because stale recovery state is worse than a missed automatic consultation.

Error-origin instructions require Angel to diagnose the repeated failure, examine the intervening evidence, and recommend a concrete recovery step with validation. Errors produced inside Angel's child session must never recursively launch another Angel consultation. A failed explicit angel consultation must not trigger automatic Angel recovery either. Difficult one-off failures remain available to executor-origin or human-origin consultation.

Executor origin

The angel tool is available at the executor's discretion when Angel is enabled and its advisor configuration is usable. Calling it does not require separate human permission. Its purpose includes difficult diagnosis, architectural trade-offs, competing explanations, contradictory evidence, and consequential decisions. It is not restricted to errors or explicit human requests.

Executor-origin instructions require Angel to answer the supplied question, investigate alternatives, and challenge unsupported assumptions. The executor waits for the tool result and then continues normally.

Human origin

/angel <question> starts the advisor directly, without asking the executor to formulate or relay the request. Human-origin instructions prioritize the user's explicit focus while retaining the shared advisor role and applicable instructions. The final advice is displayed and made available in the parent's context. It does not automatically start or restart an executor turn. Invocation while the executor is busy must respect a safe consultation boundary rather than racing shared workspace activity.

/angel on|off remains a session-local control over the tool and automatic consultations. /angel cancel cancels an active consultation without changing that control state; this is the non-TUI cancellation route when the host does not propagate idle-command aborts. Its effective state must be reported accurately.

Advisor instructions and tool contract

Instructions have three distinct homes:

  1. The angel tool's description explains its capabilities, useful invocation situations, inputs, and output.
  2. Its promptGuidelines instruct the executor and participate in Pi's system-prompt Guidelines while the tool is active.
  3. The advisor's own instructions define investigative behavior, workspace boundaries, and origin-specific focus.

The tool description must communicate this contract clearly:

Consult Angel for deep reasoning and independent investigation: difficult diagnosis, competing explanations, architectural trade-offs, contradictory evidence, or a second opinion before committing to an approach. Angel receives the focused assignment, loads project instructions, and can inspect the pinned parent session, project, and research evidence on demand. Provide the question or decision you need resolved, including any competing hypotheses. Returns evidence-backed advice, not implementation or permission.

Exact wording may improve without changing that contract. The tool's schema and descriptions must make supplying a focused question straightforward. Any optional extra context sharpens the assignment rather than requiring the executor to rebuild a context packet.

Executor guidance must establish that:

  • Angel should be called when independent investigation can materially improve the next decision, not merely for generic reassurance.
  • A concrete question, uncertainty, or competing hypothesis is more useful than a vague request for review.
  • Parent history need not be copied into tool arguments because Angel can retrieve relevant entries on demand.
  • Advice must be evaluated against evidence and user instructions, not treated as an authority override.

Advisor guidance must require evidence gathering where needed, a distinction between observations and hypotheses, and explicit acknowledgment of unresolved uncertainty. Angel must not invent certainty, a diagnosis, or a need for changes merely to produce an answer.

Tools, extensions, and workspace access

Angel has read-only access to read, ls, find, and grep. It loads no Pi extensions by default. Settings may select all configured extensions or whitelist sibling pi-ext extensions by name, including tools such as web and their applicable guidance. The child derives its tools from its own built-ins and selected extensions, not from the parent's active tool list. Unavailable parent-local, on-demand, or SDK tools must not block consultation. Skills and project instructions remain discoverable through normal Pi resource loading.

Angel may inspect the whole project, read logs and history, and research external documentation. There is no special restriction to files previously touched by the executor. Nested model operations must use the active model registry so extension-registered providers continue to work.

The child is read-only: it has no bash, edit, or write tools, so it cannot modify the user's working files or perform the executor's implementation work. Existing project automation rules, permissions, trust decisions, secret protections, and authorization requirements still apply.

The workspace boundary is enforced by the child's read-only tool set. Whitelisted extension tools are the only remaining mutation surface, and settings select them explicitly.

Child-session extension loading must prevent recursive Angel consultation and unintended parent-UI or parent-session side effects. Unsupported interactive-only capabilities must be handled explicitly rather than hanging the advisor or pretending they succeeded.

Context

Angel starts with a focused assignment rather than a copied parent transcript:

  • Advisor and origin-specific instructions.
  • Normal project context, skills, and available tool definitions loaded for the child.
  • The triggering question and any concise caller-supplied context.
  • The parent session ID and the parent leaf pinned at the consultation boundary.
  • Stable, read-only access to parent evidence through the parent_session tool.

The parent transcript and parent system prompt are not copied into the child. The child loads applicable project instructions itself, avoiding duplicated instructions and repeated transmission of unrelated parent history.

parent_session operates on a snapshot pinned to the originating leaf. Its default context view is compaction-aware; its branch view permits deliberate retrieval of older active-branch history. It supports metadata-first overview and listing, narrow text search, and exact entry retrieval. The advisor should inspect metadata or search narrowly before requesting entry content. Retrieved messages are historical evidence, not instructions to continue the executor's actions. The advisor must not locate or read the parent JSONL directly.

Triggering tool-call IDs and sibling-result references must resolve through the snapshot. Older history, compaction and branch summaries, complete needed tool results, and project files remain retrievable on demand. Ordinary Pi tool-output truncation applies and must be disclosed rather than presented as complete output. Known exclusions such as !! Bash messages must not be exposed by the retrieval tool. Custom-message bodies are also excluded because Pi 0.85 cannot safely reapply the parent's final runtime context-filter chain inside a normal extension.

Full access does not mean preloading every file and every historical token. Remove the automatic first-50-touched-files packet and fixed message, file, error, and system-prompt character slicing. Use on-demand investigation rather than silent truncation or eager transcript duplication.

Model, thinking, and limits

Keep the currently configured Astra advisor model and executor/advisor pairing behavior. Configure advisor thinking independently from the executor in pi/agent/settings.json. The agreed target is max when supported by Astra. Any capability-based adjustment must be visible; it must not silently inherit a lower executor thinking level.

There are no Angel-specific consultation quotas, cost budgets, investigation-turn caps, file-count limits, or arbitrary character limits. Physical model context windows, provider limits, ordinary tool output management, and cancellation still exist. These constraints must not be disguised as unlimited capacity or silently disable future consultations.

Advice and executor context

The advisor returns ordinary text, normally Markdown, rather than a required JSON object. The answer should explain what is happening, what to do next and why, how to verify it, and what remains unknown where those elements are useful. Evidence references and concrete observations belong in the advice. Obvious cases may have short answers; difficult cases may require substantial explanation. There are no mandatory sections, array lengths, confidence percentages, or policy fields.

Only the completed advice is delivered to the executor, once per consultation. UI metadata, progress updates, investigative tool output, and the child transcript are not injected into the executor's conversation. A concise failure or cancellation report is distinct from completed advice and must never masquerade as a successful answer.

Sessions, history, and accounting

The child session is the investigation record. Its normal Pi history records advisor messages, investigative tool calls and results, the final advice, and usage. There is no separate Angel log file, custom log viewer, or /angel log command. Users inspect investigations through existing Pi session discovery and browsing.

The parent-child relationship and the individual consultation must remain identifiable after reload or resume. The child must be distinguishable as an Angel consultation and associated with its origin and originating parent branch. The parent's displayed metadata identifies the child session.

Usage attributable to Angel includes the work performed during that consultation, including nested model work where reported. Parent history is charged only when retrieved into the child and subsequently processed; unrelated parent messages are not copied or charged. Parent attribution and child-session usage must be traceable without counting one consultation twice in aggregate accounting. Parent linkage alone must not be assumed to provide this accounting automatically. Unavailable pricing or usage must be identified as unavailable rather than fabricated as zero.

Human interface

The presentation hierarchy is:

  1. Final advice, prominently displayed as readable Markdown.
  2. Subdued metadata: advisor model, thinking level, origin, runtime, tokens, cost where available, and child-session identity.
  3. Investigation history, available through normal session browsing rather than displayed in the parent transcript.

Tool, command, and error origins share the advice component, not necessarily the same Pi delivery API. The final advice appears once in the normal parent view. No permanent availability footer, policy badges, numeric confidence badges, or duplicate advice messages are required.

Native tool behavior

The angel tool follows Pi's default tool standards. Its collapsed widget contains only compact invocation/progress/status information and bounded metadata or summaries. Full bodies and investigation output must not be placed in collapsed renderResult output. Expansion respects app.tools.expand, using keyHint("app.tools.expand", "to expand") where a hint is needed.

To keep complete advice front and center without violating compact tool behavior, a tool-origin consultation displays its advice in a separate persistent UI-only transcript entry. The executor receives the advice through the actual tool result, not a duplicate custom message. Command and error origins may use rendered custom messages with advice in content and UI metadata in details. UI-only records must not leak into the executor's model context.

Streaming and completion

Consultations provide live progress and streamed answer text using supported Pi interfaces. Tool invocations use native tool updates; command and error origins may use a temporary widget. Investigative output remains in the child session, with at most compact activity information in the parent's live presentation.

Streamed text is provisional until the advisor finishes, because an assistant text stream may precede further tool calls rather than constitute the final answer. Temporary presentation is removed when persistent final advice or an explicit failure/cancellation state is displayed. Intermediate text must not be persisted or injected as completed advice.

Cancellation stops the child investigation and its abort-aware operations. Session switching, reload, and shutdown must prevent stale results or UI updates from being delivered into a different parent context. Child history already recorded remains available for inspection. The extension must remain usable across Windows, macOS, and Linux, with supported non-TUI behavior rather than a dependency on terminal-only components.

Removed behavior and non-goals

The redesign removes:

  • Dangerous-Bash regex detection and one-time command blocking.
  • Exact-command retry exemptions and policy-gate state.
  • enforceablePolicy, confidence scoring, and policy enforcement claims.
  • Mandatory JSON advice parsing and fallback schemas.
  • maxCallsPerTask and silent exhaustion behavior.
  • Error cadence counters and fixed context-packet slicing.
  • Separate investigation logs, custom log navigation, and /angel log.

New heuristic danger filters, universal pre-action approval, multi-advisor debate, and semantic error classifiers are outside this specification. More sophisticated automatic triggers require separate evidence and agreement.

Acceptance criteria

Implementation is complete only when focused tests establish the following behavior:

  • A first eligible operation failure is remembered without consulting, and the same canonical operation failing in a later turn causes one investigation with all current sibling results before the next executor model call.
  • Canonical argument matching is key-order independent, while changed arguments identify a different operation.
  • Matching success, new human input, disable/enable control, session replacement, reload, and tree navigation clear remembered failure state.
  • Known low-signal built-ins are excluded by verified built-in provenance, while non-built-in and unknown tools receive generic opaque handling without extension-specific names.
  • Duplicate same-operation failures in one turn, successful batches, cancellations, unmappable calls, and errors inside Angel do not trigger consultations.
  • Further distinct stalled operations remain eligible after more than four consultations in one session.
  • Executor tool invocation and direct human invocation use their respective instructions and delivery semantics.
  • Human invocation does not start an otherwise idle executor.
  • The configured Astra advisor uses the independently selected supported thinking level.
  • Core tools, skill discovery, and project instructions are available to the child; extension tools such as web are available only when selected by angel.subagentExtensions, and unavailable parent-local tools do not block consultation.
  • The initial child context does not contain the parent transcript or a duplicate parent system prompt.
  • parent_session provides metadata-first, pinned, read-only retrieval of compaction and branch summaries, recent results, exact entries, and older active-branch evidence.
  • Ordinary Markdown advice works without JSON, policy fields, or confidence scores.
  • The executor receives completed advice exactly once, without UI metadata or investigation transcripts.
  • Human presentation prioritizes advice, supports live updates, follows compact tool standards, and does not duplicate the final answer.
  • Each consultation remains discoverable as a parent-linked child session after reload, with attributable usage and no parent-history double counting.
  • Cancellation, provider failures, reload, and session replacement leave explicit and accurate state without stale delivery.
  • The removed command gate no longer intercepts Bash commands.

Quality validation uses representative failure episodes to compare the existing Angel, the redesigned Angel, and an executor without Angel. Judge verified recovery, repeated failed approaches, unsupported recommendations, and incorrect advice followed by the executor. Longer answers, higher confidence, and more tool calls are not success metrics. Live evaluations are separate from deterministic regression tests; their results must not be claimed without running them.

Evidence informing the design

These sources motivate the design but do not replace the binding requirements above:

The installed Pi docs/extensions.md, docs/sdk.md, docs/tui.md, and docs/session-format.md, together with the message and entry renderer examples, establish the relevant SDK surfaces. Exact runtime wiring, cost aggregation, cancellation, streaming, and restoration must be verified against the installed SDK during implementation rather than assumed from documentation alone.

# Angel: investigative advisor specification

## Status and purpose

This specification records the agreed redesign of Angel.
It defines the required result, not an implementation plan or a description of the current code.
It is the binding contract for the authorized implementation.

**Angel provides deep reasoning and independent investigation to improve the executor's next decision.**
Advice quality and verified recovery take priority over cost or consultation frequency.
Angel is an advisor, not a permission system, command firewall, or implementation worker.

## Consultation lifecycle

Each consultation runs a capable, tool-using Pi agent in a new persistent child session linked to the originating parent session.
The advisor investigates until it can provide useful advice or identify what prevents a supported conclusion.
The parent receives the final advice, not the advisor's intermediate investigation.

A consultation has three possible origins:

- Executor: the model invokes the `angel` tool with a question or decision to investigate.
- Human: the user invokes `/angel <question>` directly.
- Error: Angel automatically investigates a completed executor tool batch containing a failed tool result.

All origins share the same investigation engine and advice presentation.
Their instructions and delivery behavior differ as specified below.

### Error origin

Automatic consultation is a deterministic stalled-recovery mechanism, not a response to every failed tool result.
The first eligible operation failure is remembered without consulting.
If the same opaque operation fails again in a later completed executor turn before the task boundary changes, Angel starts one consultation before the executor's next model call.
A matching successful operation clears the remembered failure and makes a later failure a new incident.
After one consultation, the same operation does not trigger again until it succeeds or the task boundary changes.

Operation identity is the tool name plus a cryptographic hash of canonically ordered structured arguments.
Arguments are matched exactly after canonical ordering; Angel does not normalize tool-specific fields, inspect command text, classify error prose, or persist argument values.
A result whose originating structured call cannot be mapped reliably is not eligible for automatic consultation.
Duplicate failures for one operation in the same completed turn count as one attempt, not stalled recovery.

Tool provenance comes from Pi's active tool metadata.
Known low-signal Pi built-ins `read`, `edit`, `write`, `grep`, `find`, and `ls` are excluded from automatic consultation because their ordinary recovery is local discovery or retry.
This exclusion applies only when Pi identifies the actual registration as built-in; an extension override with the same name remains an opaque non-built-in tool.
Built-in process tools such as `bash` and `powershell`, extension tools, SDK tools, future unknown tools, and tools with unavailable provenance use the same generic repeated-operation rule.
Angel must not contain names or domain assumptions for unrelated extensions.
The `angel` tool itself is always excluded to prevent recursion.

A triggering batch produces one consultation with all completed sibling results available, including successes and first-time failures.
There is no every-Nth-error counter or session-wide error-review budget.
Explicit cancellation is not a failure candidate.
Remembered state is session-local and is cleared on new human input, session replacement, tree navigation, reload, or `/angel on|off`.
It is intentionally not restored from transcripts because stale recovery state is worse than a missed automatic consultation.

Error-origin instructions require Angel to diagnose the repeated failure, examine the intervening evidence, and recommend a concrete recovery step with validation.
Errors produced inside Angel's child session must never recursively launch another Angel consultation.
A failed explicit `angel` consultation must not trigger automatic Angel recovery either.
Difficult one-off failures remain available to executor-origin or human-origin consultation.

### Executor origin

The `angel` tool is available at the executor's discretion when Angel is enabled and its advisor configuration is usable.
Calling it does not require separate human permission.
Its purpose includes difficult diagnosis, architectural trade-offs, competing explanations, contradictory evidence, and consequential decisions.
It is not restricted to errors or explicit human requests.

Executor-origin instructions require Angel to answer the supplied question, investigate alternatives, and challenge unsupported assumptions.
The executor waits for the tool result and then continues normally.

### Human origin

`/angel <question>` starts the advisor directly, without asking the executor to formulate or relay the request.
Human-origin instructions prioritize the user's explicit focus while retaining the shared advisor role and applicable instructions.
The final advice is displayed and made available in the parent's context.
It does not automatically start or restart an executor turn.
Invocation while the executor is busy must respect a safe consultation boundary rather than racing shared workspace activity.

`/angel on|off` remains a session-local control over the tool and automatic consultations.
`/angel cancel` cancels an active consultation without changing that control state; this is the non-TUI cancellation route when the host does not propagate idle-command aborts.
Its effective state must be reported accurately.

## Advisor instructions and tool contract

Instructions have three distinct homes:

1. The `angel` tool's `description` explains its capabilities, useful invocation situations, inputs, and output.
2. Its `promptGuidelines` instruct the executor and participate in Pi's system-prompt Guidelines while the tool is active.
3. The advisor's own instructions define investigative behavior, workspace boundaries, and origin-specific focus.

The tool description must communicate this contract clearly:

> Consult Angel for deep reasoning and independent investigation: difficult diagnosis, competing explanations, architectural trade-offs, contradictory evidence, or a second opinion before committing to an approach.
> Angel receives the focused assignment, loads project instructions, and can inspect the pinned parent session, project, and research evidence on demand.
> Provide the question or decision you need resolved, including any competing hypotheses.
> Returns evidence-backed advice, not implementation or permission.

Exact wording may improve without changing that contract.
The tool's schema and descriptions must make supplying a focused question straightforward.
Any optional extra context sharpens the assignment rather than requiring the executor to rebuild a context packet.

Executor guidance must establish that:

- Angel should be called when independent investigation can materially improve the next decision, not merely for generic reassurance.
- A concrete question, uncertainty, or competing hypothesis is more useful than a vague request for review.
- Parent history need not be copied into tool arguments because Angel can retrieve relevant entries on demand.
- Advice must be evaluated against evidence and user instructions, not treated as an authority override.

Advisor guidance must require evidence gathering where needed, a distinction between observations and hypotheses, and explicit acknowledgment of unresolved uncertainty.
Angel must not invent certainty, a diagnosis, or a need for changes merely to produce an answer.

## Tools, extensions, and workspace access

Angel has read-only access to `read`, `ls`, `find`, and `grep`.
It loads no Pi extensions by default.
Settings may select all configured extensions or whitelist sibling pi-ext extensions by name, including tools such as `web` and their applicable guidance.
The child derives its tools from its own built-ins and selected extensions, not from the parent's active tool list.
Unavailable parent-local, on-demand, or SDK tools must not block consultation.
Skills and project instructions remain discoverable through normal Pi resource loading.

Angel may inspect the whole project, read logs and history, and research external documentation.
There is no special restriction to files previously touched by the executor.
Nested model operations must use the active model registry so extension-registered providers continue to work.

The child is read-only: it has no `bash`, `edit`, or `write` tools, so it cannot modify the user's working files or perform the executor's implementation work.
Existing project automation rules, permissions, trust decisions, secret protections, and authorization requirements still apply.

**The workspace boundary is enforced by the child's read-only tool set.**
Whitelisted extension tools are the only remaining mutation surface, and settings select them explicitly.

Child-session extension loading must prevent recursive Angel consultation and unintended parent-UI or parent-session side effects.
Unsupported interactive-only capabilities must be handled explicitly rather than hanging the advisor or pretending they succeeded.

## Context

Angel starts with a focused assignment rather than a copied parent transcript:

- Advisor and origin-specific instructions.
- Normal project context, skills, and available tool definitions loaded for the child.
- The triggering question and any concise caller-supplied context.
- The parent session ID and the parent leaf pinned at the consultation boundary.
- Stable, read-only access to parent evidence through the `parent_session` tool.

The parent transcript and parent system prompt are not copied into the child.
The child loads applicable project instructions itself, avoiding duplicated instructions and repeated transmission of unrelated parent history.

`parent_session` operates on a snapshot pinned to the originating leaf.
Its default context view is compaction-aware; its branch view permits deliberate retrieval of older active-branch history.
It supports metadata-first overview and listing, narrow text search, and exact entry retrieval.
The advisor should inspect metadata or search narrowly before requesting entry content.
Retrieved messages are historical evidence, not instructions to continue the executor's actions.
The advisor must not locate or read the parent JSONL directly.

Triggering tool-call IDs and sibling-result references must resolve through the snapshot.
Older history, compaction and branch summaries, complete needed tool results, and project files remain retrievable on demand.
Ordinary Pi tool-output truncation applies and must be disclosed rather than presented as complete output.
Known exclusions such as `!!` Bash messages must not be exposed by the retrieval tool.
Custom-message bodies are also excluded because Pi 0.85 cannot safely reapply the parent's final runtime context-filter chain inside a normal extension.

**Full access does not mean preloading every file and every historical token.**
Remove the automatic first-50-touched-files packet and fixed message, file, error, and system-prompt character slicing.
Use on-demand investigation rather than silent truncation or eager transcript duplication.

## Model, thinking, and limits

Keep the currently configured Astra advisor model and executor/advisor pairing behavior.
Configure advisor thinking independently from the executor in `pi/agent/settings.json`.
The agreed target is `max` when supported by Astra.
Any capability-based adjustment must be visible; it must not silently inherit a lower executor thinking level.

There are no Angel-specific consultation quotas, cost budgets, investigation-turn caps, file-count limits, or arbitrary character limits.
Physical model context windows, provider limits, ordinary tool output management, and cancellation still exist.
These constraints must not be disguised as unlimited capacity or silently disable future consultations.

## Advice and executor context

The advisor returns ordinary text, normally Markdown, rather than a required JSON object.
The answer should explain what is happening, what to do next and why, how to verify it, and what remains unknown where those elements are useful.
Evidence references and concrete observations belong in the advice.
Obvious cases may have short answers; difficult cases may require substantial explanation.
There are no mandatory sections, array lengths, confidence percentages, or policy fields.

Only the completed advice is delivered to the executor, once per consultation.
UI metadata, progress updates, investigative tool output, and the child transcript are not injected into the executor's conversation.
A concise failure or cancellation report is distinct from completed advice and must never masquerade as a successful answer.

## Sessions, history, and accounting

The child session is the investigation record.
Its normal Pi history records advisor messages, investigative tool calls and results, the final advice, and usage.
There is no separate Angel log file, custom log viewer, or `/angel log` command.
Users inspect investigations through existing Pi session discovery and browsing.

The parent-child relationship and the individual consultation must remain identifiable after reload or resume.
The child must be distinguishable as an Angel consultation and associated with its origin and originating parent branch.
The parent's displayed metadata identifies the child session.

Usage attributable to Angel includes the work performed during that consultation, including nested model work where reported.
Parent history is charged only when retrieved into the child and subsequently processed; unrelated parent messages are not copied or charged.
Parent attribution and child-session usage must be traceable without counting one consultation twice in aggregate accounting.
Parent linkage alone must not be assumed to provide this accounting automatically.
Unavailable pricing or usage must be identified as unavailable rather than fabricated as zero.

## Human interface

The presentation hierarchy is:

1. Final advice, prominently displayed as readable Markdown.
2. Subdued metadata: advisor model, thinking level, origin, runtime, tokens, cost where available, and child-session identity.
3. Investigation history, available through normal session browsing rather than displayed in the parent transcript.

Tool, command, and error origins share the advice component, not necessarily the same Pi delivery API.
The final advice appears once in the normal parent view.
No permanent availability footer, policy badges, numeric confidence badges, or duplicate advice messages are required.

### Native tool behavior

The `angel` tool follows Pi's default tool standards.
Its collapsed widget contains only compact invocation/progress/status information and bounded metadata or summaries.
Full bodies and investigation output must not be placed in collapsed `renderResult` output.
Expansion respects `app.tools.expand`, using `keyHint("app.tools.expand", "to expand")` where a hint is needed.

To keep complete advice front and center without violating compact tool behavior, a tool-origin consultation displays its advice in a separate persistent UI-only transcript entry.
The executor receives the advice through the actual tool result, not a duplicate custom message.
Command and error origins may use rendered custom messages with advice in `content` and UI metadata in `details`.
UI-only records must not leak into the executor's model context.

### Streaming and completion

Consultations provide live progress and streamed answer text using supported Pi interfaces.
Tool invocations use native tool updates; command and error origins may use a temporary widget.
Investigative output remains in the child session, with at most compact activity information in the parent's live presentation.

Streamed text is provisional until the advisor finishes, because an assistant text stream may precede further tool calls rather than constitute the final answer.
Temporary presentation is removed when persistent final advice or an explicit failure/cancellation state is displayed.
Intermediate text must not be persisted or injected as completed advice.

Cancellation stops the child investigation and its abort-aware operations.
Session switching, reload, and shutdown must prevent stale results or UI updates from being delivered into a different parent context.
Child history already recorded remains available for inspection.
The extension must remain usable across Windows, macOS, and Linux, with supported non-TUI behavior rather than a dependency on terminal-only components.

## Removed behavior and non-goals

The redesign removes:

- Dangerous-Bash regex detection and one-time command blocking.
- Exact-command retry exemptions and policy-gate state.
- `enforceablePolicy`, confidence scoring, and policy enforcement claims.
- Mandatory JSON advice parsing and fallback schemas.
- `maxCallsPerTask` and silent exhaustion behavior.
- Error cadence counters and fixed context-packet slicing.
- Separate investigation logs, custom log navigation, and `/angel log`.

New heuristic danger filters, universal pre-action approval, multi-advisor debate, and semantic error classifiers are outside this specification.
More sophisticated automatic triggers require separate evidence and agreement.

## Acceptance criteria

Implementation is complete only when focused tests establish the following behavior:

- A first eligible operation failure is remembered without consulting, and the same canonical operation failing in a later turn causes one investigation with all current sibling results before the next executor model call.
- Canonical argument matching is key-order independent, while changed arguments identify a different operation.
- Matching success, new human input, disable/enable control, session replacement, reload, and tree navigation clear remembered failure state.
- Known low-signal built-ins are excluded by verified built-in provenance, while non-built-in and unknown tools receive generic opaque handling without extension-specific names.
- Duplicate same-operation failures in one turn, successful batches, cancellations, unmappable calls, and errors inside Angel do not trigger consultations.
- Further distinct stalled operations remain eligible after more than four consultations in one session.
- Executor tool invocation and direct human invocation use their respective instructions and delivery semantics.
- Human invocation does not start an otherwise idle executor.
- The configured Astra advisor uses the independently selected supported thinking level.
- Core tools, skill discovery, and project instructions are available to the child; extension tools such as `web` are available only when selected by `angel.subagentExtensions`, and unavailable parent-local tools do not block consultation.
- The initial child context does not contain the parent transcript or a duplicate parent system prompt.
- `parent_session` provides metadata-first, pinned, read-only retrieval of compaction and branch summaries, recent results, exact entries, and older active-branch evidence.
- Ordinary Markdown advice works without JSON, policy fields, or confidence scores.
- The executor receives completed advice exactly once, without UI metadata or investigation transcripts.
- Human presentation prioritizes advice, supports live updates, follows compact tool standards, and does not duplicate the final answer.
- Each consultation remains discoverable as a parent-linked child session after reload, with attributable usage and no parent-history double counting.
- Cancellation, provider failures, reload, and session replacement leave explicit and accurate state without stale delivery.
- The removed command gate no longer intercepts Bash commands.

Quality validation uses representative failure episodes to compare the existing Angel, the redesigned Angel, and an executor without Angel.
Judge verified recovery, repeated failed approaches, unsupported recommendations, and incorrect advice followed by the executor.
Longer answers, higher confidence, and more tool calls are not success metrics.
Live evaluations are separate from deterministic regression tests; their results must not be claimed without running them.

## Evidence informing the design

These sources motivate the design but do not replace the binding requirements above:

- [Anthropic advisor tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/advisor-tool): transcript-informed advice returned to an executor.
- [Claude Code subagents](https://code.claude.com/docs/en/sub-agents): separate agent contexts with tools and focused assignments.
- [Codex subagents](https://developers.openai.com/codex/multi-agent): independent investigative work with results returned to the parent.
- [Effective context engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents): just-in-time retrieval and context management rather than indiscriminate preloading.
- [CRITIC](https://arxiv.org/abs/2305.11738) and [the critical survey of self-correction](https://arxiv.org/abs/2406.01297): external evidence matters more than ungrounded self-critique; these studies do not establish current-model performance for Angel.

The installed Pi `docs/extensions.md`, `docs/sdk.md`, `docs/tui.md`, and `docs/session-format.md`, together with the message and entry renderer examples, establish the relevant SDK surfaces.
Exact runtime wiring, cost aggregation, cancellation, streaming, and restoration must be verified against the installed SDK during implementation rather than assumed from documentation alone.