--- id: PX-RESEARCH-LXU8HWNJ type: research title: Free Extension Post-Mortem --- ## Free Extension Post-Mortem Prior attempt at the same idea as [Gratis Synthetic Free Provider](../../specs/PX-SPEC-CNDFHNO2-gratis-synthetic-free-provider/index.md): a `free` extension in the dotfiles Pi setup, April-May 2026. Sources: dotfiles git history commits `5fbf178c` (added), `b87a61ad` (packaged), `5ca1296c` (reworked), `a6c52214` (deleted). ### What it was Single extension registering three synthetic providers over keyless free endpoints: | Synthetic provider | Upstream | API | Catalog | | --- | --- | --- | --- | | `free` | Kilo gateway (`api.kilo.ai/api/gateway/v1`) | openai-completions | hardcoded, manual `/free-refresh` re-fetch | | `zen` | opencode.ai/zen (big-pickle, minimax-m2.5-free) | anthropic-messages | startup fetch of `zen/v1/models` | | `zen-free` | opencode.ai/zen (hy3, ling, nemotron `-free`) | openai-completions | same startup fetch | Model IDs leaked upstream paths (`kilo-auto/free`, `openrouter/free`, `nvidia/nemotron-3-super-120b-a12b:free`). Auth was `apiKey: "-"` with `authHeader: false` — both endpoints served free models without accounts in April 2026. ### Timeline ```mermaid flowchart LR add["2026-05-04 added
5fbf178c"] --> pkg["2026-05-04 packaged
b87a61ad"] --> rw["2026-05-09 reworked
5ca1296c"] --> del["2026-05-21 deleted
a6c52214"] ``` Deleted in a commit otherwise titled "config: add kimi coding model" (same change migrated all extensions to a new package scope and rewrote `models.json`) — no failure note, dropped as dead weight after ~2.5 weeks. ### Failure analysis 1. **Keyless assumption expired.** Free access without any account was the load-bearing design. Zen now requires login + API key; Kilo free models moved behind its own gateway account. When endpoints started rejecting unauthenticated calls, every registered model broke at request time. 2. **Visible but broken, the worst failure mode.** Fake key `-` meant providers always registered. Stale catalogs and dead endpoints surfaced as per-request errors, never as absence. 3. **Provider sprawl.** Three synthetic providers named after API shapes (`zen` vs `zen-free`) forced users to know transport details to pick a model. 4. **Catalog churn.** Hardcoded Kilo list rotted between manual `/free-refresh` runs; a session-start notify reminded every session, which is noise. 5. **No resilience.** No retry, no failover, no cooldown; one 429 ended the request. 6. **Static model params blocked the auto router.** Pi model attributes (context window, max tokens) were fixed per registered model entry at the time; there was no way to give one synthetic router model per-backend parameters. (Human-reported; consistent with the single-entry-per-model design of the old code.) Current Pi resolves 6: `ProviderModelConfig` supports per-model `api` and `baseUrl` overrides, `registerProvider` calls after load apply immediately (model sets replaceable live), custom `streamSimple` gives full per-request control, and `message_end` error normalization makes context-overflow recovery work for custom providers ([custom-provider docs](https://github.com/earendil-works/pi/blob/main/packages/ai/test)). Viable auto strategy: advertise the minimum context window across keyed backends, normalize upstream overflow errors to `context_length_exceeded`, and treat overflow like exhaustion in the cascade. Points 1-4 are evidenced by the code and current provider docs; point 5's removal trigger (commit pairing) is inferred, not documented. ### Lessons vs current spec | Old failure | Gratis spec countermeasure | | --- | --- | | Keyless endpoints died | GRATIS_*/native key chain; unkeyed backend invisible | | Visible but broken models | Register only keyed backends; no key → no provider | | Three API-shaped providers | One `grts` provider, backend-nested model IDs | | Stale hardcoded catalogs | Static curated + Zen live discovery of `-free` models | | No failover | `grts/auto` cascade with per-backend cooldown | ### New constraint discovered Zen free models are split across two API types: `big-pickle`/`minimax`-class via anthropic-messages, others via openai-completions — and the split can change between catalog refreshes. The spec's backends table pins opencode to "OpenAI completions (Zen v1)"; correct model is per-model API type carried by each catalog entry, not per-backend API choice. Same applies in principle to every backend: catalog entries, not backends, own the API type. ### Unresolved questions - Why exactly deletion happened on 2026-05-21: no note in commits; if session logs hold the actual breakage story, it would confirm inference 1 vs an unrelated cleanup motive. - Kilo re-evaluated separately: [Kilo Gateway Free Access](../PX-RESEARCH-KZ2M6WQ4-kilo-gateway-free-access/index.md) — verdict: include as the keyless v1 tier; the old Kilo failure was catalog rot, not auth.