id: PX-RESEARCH-TVTIE1_F
type: research
title: Claude Code and Codex Compaction Compared with Pi and Decay
Scope
Comparison concerns Claude Code, open-source Codex CLI, installed Pi default compaction, and this repository's Decay extension as inspected on 2026-09-27.
Claude Messages API compaction and OpenAI Responses API compaction are separate products, not evidence of every CLI implementation detail.
Strategies at a glance
System
Trigger and transform
Context kept verbatim
Recovery and inspectability
Claude Code
Model/config-dependent auto threshold or focused /compact; clears older tool outputs before summarizing if needed.
Re-injected persistent instructions and memory, bounded recent files and skills; exact retained conversation boundary undocumented.
Summary prompt and failure policy are not public; documented auto-compaction thrash guard.
Codex CLI
Provider capability selects remote Responses compaction V2 or local model-written checkpoint; manual and auto paths exist.
Remote: eligible messages within 64k-token retention budget; local: recent user messages within 20k-token budget.
Remote checkpoint is opaque; local checkpoint is readable; both use lifecycle hooks.
Pi default
Context estimate exceeds window minus 16,384-token reserve, manual /compact, or overflow recovery; one structured summary request.
Pi-selected whole-entry suffix targeting 20k tokens, plus summary and file-operation appendix.
Readable checkpoint and raw JSONL history; transient request retry and one overflow compact-and-retry.
Decay
Uses Pi's trigger, cut, and retained suffix; classifies discarded chunks into atoms, then allocates and renders a checkpoint.
Same Pi-selected suffix; prior atom identities and file lists carried in compaction details.
Validated chunk coverage; bounded gap fills; classifier failure delegates to Pi, summarizer failure salvages atoms.
The table summarizes documented/source-visible behavior, not a measured quality or latency ranking.
Claude Code
Claude Code first clears older tool outputs, then summarizes when necessary; the auto threshold depends on model and configuration rather than one universal number.
Manual /compact accepts a focus, /autocompact changes the trigger, and a Compact Instructions section in CLAUDE.md can guide preservation.
Early conversational instructions can still disappear, so persistent rules belong in CLAUDE.md rather than in old turns. Claude workflow · Context-window controls.
After compaction, Claude Code reapplies system prompt/output style, reloads project-root CLAUDE.md and auto memory from disk, refreshes git state, and re-injects the plan.
It re-reads at most five recently modified files; files above 5,000 tokens return as path references, not file bodies.
Invoked skills are re-injected under 5,000 tokens per skill and 25,000 tokens total; matching SessionStart hooks run again.
Since Claude Code v2.1.198 its compaction summary request also inherits the session's extended-thinking setting without changing that setting afterward. What survives compaction.
The public documentation does not disclose Claude Code's exact checkpoint prompt, retained-turn selection, model-call retry policy, or measured compaction cost.
Claude's separately documented Messages API compaction cannot establish those CLI internals.
Codex CLI
Codex selects the implementation by provider capability: RemoteCompactionSupport::V2 uses Responses-backed remote compaction, while Unsupported uses the local summarization path; a separate token-budget feature can select another path.
Automatic compaction can occur before a turn, during a continuing turn, or after a turn, including context-limit and model-downshift cases. Codex turn selection.
Remote V2 sends history and a compaction trigger to the provider, expects exactly one compaction output item, and installs it after eligible messages.
Its client retention filter keeps user/hook messages and selected small agent messages within a 64k-token budget; it excludes most tool outputs and does not simply preserve a contiguous recent-turn suffix.
It allows up to two stream retries and has a model-fallback path for selected failures. Remote V2 request · Remote retention and retries · Retention filter.
The Responses API guide describes the returned state as encrypted and opaque; it does not disclose what semantic details the server preserved.
The local path asks a model for a concise handoff of progress, decisions, constraints, remaining work, and critical references.
It installs a prefixed readable summary alongside up to 20k tokens of most-recent user messages, then restores initial context at the appropriate subsequent boundary.
On compaction-request overflow it removes oldest history items; stream failures use bounded retries.
The CLI warns that long threads and repeated compactions may reduce accuracy. Local prompt · Local implementation · Retained user-message limit.
Pi default and Decay
Pi compacts when estimated context exceeds contextWindow - reserveTokens, normally with a 16,384-token reserve and a 20,000-token recent-suffix target.
It chooses a valid whole-entry cut, never separates a tool result from its call, and summarizes discarded serialized messages into fixed handoff sections.
A prior summary is updated iteratively; read/modified file paths are appended separately; raw session entries remain available outside rebuilt model context.
Tool-result text is truncated to 2,000 characters during serialization.
Split user-message spans receive a separate prefix summary; transient summary requests can retry, and overflow recovery allows one compact-and-retry. Pi compaction implementation · Cut and generation · Installed Pi reference.
Decay intercepts Pi's session_before_compact without replacing Pi's trigger or cut point.
It groups discarded messages into roughly 6k-estimated-token chunks, requests numbered record_chunk tool calls, validates coverage, and permits up to two gap-fill rounds per classifier model before failure or model fallback.
Local code merges atoms with a bounded prior identity catalog, weighs recurrence and recency, protects critical kinds, and allocates summary tokens before the second model formats Markdown.
It stores atoms and file lists in the compaction entry, uses separate classifier/summarizer model settings, and preserves the same Pi-selected suffix.
Abort cancels; classifier failure delegates to Pi's default compactor; summarizer failure renders the allocated atoms without prose synthesis. Decay implementation · Atom policy · Model prompts · Decay behavior.
Conclusions
Claude Code's durable advantage is explicit re-injection of disk-backed instructions and bounded source files; a summary alone is not asked to preserve that entire layer.
Codex remote compaction optimizes provider-managed continuity through opaque state; Codex local, Pi default, and Decay instead expose readable handoffs with different retention rules.
Pi default minimizes application-specific machinery, while Decay spends an extra classification stage for inspectable recurrence, priority, and safe fallback.
Decay's complete chunk coverage prevents silent structural omissions but cannot prove every important fact was extracted; no controlled four-way retention, latency, or cost evaluation is established here.
Claude Code's tool-output eviction and Codex remote retention filter make direct token or speed comparisons against Pi's serialized-history compaction invalid without identical task traces and model budgets.
Unresolved questions
Which Claude Code compaction details are stable beyond documented behavior, especially summary prompt, retry policy, exact recent-turn boundary, and cost accounting?
Which Codex provider/model combinations currently select remote V2, local summarization, or token-budget reset in practice?
Across repeated long coding sessions, which strategy best retains user constraints, exact literals, tool provenance, and pending work at equal total token and monetary budgets?
Does Decay's classifier-plus-summarizer improve real continuation outcomes enough to justify its added latency and failure modes over Pi's single-request default?
---
id: PX-RESEARCH-TVTIE1_F
type: research
title: Claude Code and Codex Compaction Compared with Pi and Decay
---
## Scope
Comparison concerns Claude Code, open-source Codex CLI, installed Pi default compaction, and this repository's Decay extension as inspected on 2026-09-27.
Claude Messages API compaction and OpenAI Responses API compaction are separate products, not evidence of every CLI implementation detail.
## Strategies at a glance
| System | Trigger and transform | Context kept verbatim | Recovery and inspectability |
| --- | --- | --- | --- |
| Claude Code | Model/config-dependent auto threshold or focused `/compact`; clears older tool outputs before summarizing if needed. | Re-injected persistent instructions and memory, bounded recent files and skills; exact retained conversation boundary undocumented. | Summary prompt and failure policy are not public; documented auto-compaction thrash guard. |
| Codex CLI | Provider capability selects remote Responses compaction V2 or local model-written checkpoint; manual and auto paths exist. | Remote: eligible messages within 64k-token retention budget; local: recent user messages within 20k-token budget. | Remote checkpoint is opaque; local checkpoint is readable; both use lifecycle hooks. |
| Pi default | Context estimate exceeds window minus 16,384-token reserve, manual `/compact`, or overflow recovery; one structured summary request. | Pi-selected whole-entry suffix targeting 20k tokens, plus summary and file-operation appendix. | Readable checkpoint and raw JSONL history; transient request retry and one overflow compact-and-retry. |
| Decay | Uses Pi's trigger, cut, and retained suffix; classifies discarded chunks into atoms, then allocates and renders a checkpoint. | Same Pi-selected suffix; prior atom identities and file lists carried in compaction details. | Validated chunk coverage; bounded gap fills; classifier failure delegates to Pi, summarizer failure salvages atoms. |
The table summarizes documented/source-visible behavior, not a measured quality or latency ranking.
## Claude Code
Claude Code first clears older tool outputs, then summarizes when necessary; the auto threshold depends on model and configuration rather than one universal number.
Manual `/compact` accepts a focus, `/autocompact` changes the trigger, and a `Compact Instructions` section in CLAUDE.md can guide preservation.
Early conversational instructions can still disappear, so persistent rules belong in CLAUDE.md rather than in old turns. [Claude workflow](https://code.claude.com/docs/en/how-claude-code-works#when-context-fills-up) · [Context-window controls](https://code.claude.com/docs/en/context-window#when-your-context-fills-up).
After compaction, Claude Code reapplies system prompt/output style, reloads project-root CLAUDE.md and auto memory from disk, refreshes git state, and re-injects the plan.
It re-reads at most five recently modified files; files above 5,000 tokens return as path references, not file bodies.
Invoked skills are re-injected under 5,000 tokens per skill and 25,000 tokens total; matching `SessionStart` hooks run again.
Since Claude Code v2.1.198 its compaction summary request also inherits the session's extended-thinking setting without changing that setting afterward. [What survives compaction](https://code.claude.com/docs/en/context-window#what-survives-compaction).
The public documentation does not disclose Claude Code's exact checkpoint prompt, retained-turn selection, model-call retry policy, or measured compaction cost.
Claude's separately documented [Messages API compaction](https://platform.claude.com/docs/en/build-with-claude/compaction) cannot establish those CLI internals.
## Codex CLI
Codex selects the implementation by provider capability: `RemoteCompactionSupport::V2` uses Responses-backed remote compaction, while `Unsupported` uses the local summarization path; a separate token-budget feature can select another path.
Automatic compaction can occur before a turn, during a continuing turn, or after a turn, including context-limit and model-downshift cases. [Codex turn selection](https://github.com/openai/codex/blob/819cdb726dab1c01f49099ce7d960a5f1383c2fb/codex-rs/core/src/session/turn.rs#L1450-L1507).
```mermaid
flowchart TB
session["Codex CLI compaction"] --> budget{"Token-budget mode?"}
budget -->|yes| reset["Separate reset path"]
budget -->|no| capability{"Provider supports remote V2?"}
capability -->|yes| remote["Responses compaction item"]
capability -->|no| local["Model-written checkpoint"]
remote --> remote_context["Eligible recent messages plus opaque item"]
local --> local_context["Recent user messages plus readable summary"]
```
Remote V2 sends history and a compaction trigger to the provider, expects exactly one compaction output item, and installs it after eligible messages.
Its client retention filter keeps user/hook messages and selected small agent messages within a 64k-token budget; it excludes most tool outputs and does not simply preserve a contiguous recent-turn suffix.
It allows up to two stream retries and has a model-fallback path for selected failures. [Remote V2 request](https://github.com/openai/codex/blob/819cdb726dab1c01f49099ce7d960a5f1383c2fb/codex-rs/core/src/compact_remote_v2_attempt.rs#L30-L132) · [Remote retention and retries](https://github.com/openai/codex/blob/819cdb726dab1c01f49099ce7d960a5f1383c2fb/codex-rs/core/src/compact_remote_v2.rs#L386-L530) · [Retention filter](https://github.com/openai/codex/blob/819cdb726dab1c01f49099ce7d960a5f1383c2fb/codex-rs/core/src/compact_remote_v2.rs#L555-L660).
The [Responses API guide](https://developers.openai.com/api/docs/guides/compaction) describes the returned state as encrypted and opaque; it does not disclose what semantic details the server preserved.
The local path asks a model for a concise handoff of progress, decisions, constraints, remaining work, and critical references.
It installs a prefixed readable summary alongside up to 20k tokens of most-recent user messages, then restores initial context at the appropriate subsequent boundary.
On compaction-request overflow it removes oldest history items; stream failures use bounded retries.
The CLI warns that long threads and repeated compactions may reduce accuracy. [Local prompt](https://github.com/openai/codex/blob/819cdb726dab1c01f49099ce7d960a5f1383c2fb/codex-rs/prompts/templates/compact/prompt.md) · [Local implementation](https://github.com/openai/codex/blob/819cdb726dab1c01f49099ce7d960a5f1383c2fb/codex-rs/core/src/compact.rs#L255-L400) · [Retained user-message limit](https://github.com/openai/codex/blob/819cdb726dab1c01f49099ce7d960a5f1383c2fb/codex-rs/core/src/compact.rs#L649-L739).
## Pi default and Decay
Pi compacts when estimated context exceeds `contextWindow - reserveTokens`, normally with a 16,384-token reserve and a 20,000-token recent-suffix target.
It chooses a valid whole-entry cut, never separates a tool result from its call, and summarizes discarded serialized messages into fixed handoff sections.
A prior summary is updated iteratively; read/modified file paths are appended separately; raw session entries remain available outside rebuilt model context.
Tool-result text is truncated to 2,000 characters during serialization.
Split user-message spans receive a separate prefix summary; transient summary requests can retry, and overflow recovery allows one compact-and-retry. [Pi compaction implementation](https://github.com/earendil-works/pi/blob/2b0a123de98318c2ff8069661721ce0c3794c34e/packages/coding-agent/src/core/compaction/compaction.ts#L148-L152) · [Cut and generation](https://github.com/earendil-works/pi/blob/2b0a123de98318c2ff8069661721ce0c3794c34e/packages/coding-agent/src/core/compaction/compaction.ts#L453-L560) · [Installed Pi reference](https://github.com/earendil-works/pi/blob/2b0a123de98318c2ff8069661721ce0c3794c34e/packages/coding-agent/docs/compaction.md).
Decay intercepts Pi's `session_before_compact` without replacing Pi's trigger or cut point.
It groups discarded messages into roughly 6k-estimated-token chunks, requests numbered `record_chunk` tool calls, validates coverage, and permits up to two gap-fill rounds per classifier model before failure or model fallback.
Local code merges atoms with a bounded prior identity catalog, weighs recurrence and recency, protects critical kinds, and allocates summary tokens before the second model formats Markdown.
It stores atoms and file lists in the compaction entry, uses separate classifier/summarizer model settings, and preserves the same Pi-selected suffix.
Abort cancels; classifier failure delegates to Pi's default compactor; summarizer failure renders the allocated atoms without prose synthesis. [Decay implementation](../../../extensions/decay/implementation.ts) · [Atom policy](../../../extensions/decay/core.ts) · [Model prompts](../../../extensions/decay/prompts.ts) · [Decay behavior](../../../extensions/decay/README.md).
## Conclusions
- Claude Code's durable advantage is explicit re-injection of disk-backed instructions and bounded source files; a summary alone is not asked to preserve that entire layer.
- Codex remote compaction optimizes provider-managed continuity through opaque state; Codex local, Pi default, and Decay instead expose readable handoffs with different retention rules.
- Pi default minimizes application-specific machinery, while Decay spends an extra classification stage for inspectable recurrence, priority, and safe fallback.
- Decay's complete chunk coverage prevents silent structural omissions but cannot prove every important fact was extracted; no controlled four-way retention, latency, or cost evaluation is established here.
- Claude Code's tool-output eviction and Codex remote retention filter make direct token or speed comparisons against Pi's serialized-history compaction invalid without identical task traces and model budgets.
## Unresolved questions
- Which Claude Code compaction details are stable beyond documented behavior, especially summary prompt, retry policy, exact recent-turn boundary, and cost accounting?
- Which Codex provider/model combinations currently select remote V2, local summarization, or token-budget reset in practice?
- Across repeated long coding sessions, which strategy best retains user constraints, exact literals, tool provenance, and pending work at equal total token and monetary budgets?
- Does Decay's classifier-plus-summarizer improve real continuation outcomes enough to justify its added latency and failure modes over Pi's single-request default?