Luigit
repositories / pi-ext

pi-ext

bugabingas pi extensions

owned by admin

.system/research/PX-RESEARCH-LXU8HWNJ-free-extension-post-mortem/index.md

Raw
Rendered preview

id: PX-RESEARCH-LXU8HWNJ type: research title: Free Extension Post-Mortem

Free Extension Post-Mortem

Prior attempt at the same idea as Gratis Synthetic Free Provider: a free extension in the dotfiles Pi setup, April-May 2026. Sources: dotfiles git history commits 5fbf178c (added), b87a61ad (packaged), 5ca1296c (reworked), a6c52214 (deleted).

What it was

Single extension registering three synthetic providers over keyless free endpoints:

Synthetic provider Upstream API Catalog
free Kilo gateway (api.kilo.ai/api/gateway/v1) openai-completions hardcoded, manual /free-refresh re-fetch
zen opencode.ai/zen (big-pickle, minimax-m2.5-free) anthropic-messages startup fetch of zen/v1/models
zen-free opencode.ai/zen (hy3, ling, nemotron -free) openai-completions same startup fetch

Model IDs leaked upstream paths (kilo-auto/free, openrouter/free, nvidia/nemotron-3-super-120b-a12b:free). Auth was apiKey: "-" with authHeader: false — both endpoints served free models without accounts in April 2026.

Timeline

flowchart LR
  add["2026-05-04 added<br/>5fbf178c"] --> pkg["2026-05-04 packaged<br/>b87a61ad"] --> rw["2026-05-09 reworked<br/>5ca1296c"] --> del["2026-05-21 deleted<br/>a6c52214"]

Deleted in a commit otherwise titled "config: add kimi coding model" (same change migrated all extensions to a new package scope and rewrote models.json) — no failure note, dropped as dead weight after ~2.5 weeks.

Failure analysis

  1. Keyless assumption expired. Free access without any account was the load-bearing design. Zen now requires login + API key; Kilo free models moved behind its own gateway account. When endpoints started rejecting unauthenticated calls, every registered model broke at request time.
  2. Visible but broken, the worst failure mode. Fake key - meant providers always registered. Stale catalogs and dead endpoints surfaced as per-request errors, never as absence.
  3. Provider sprawl. Three synthetic providers named after API shapes (zen vs zen-free) forced users to know transport details to pick a model.
  4. Catalog churn. Hardcoded Kilo list rotted between manual /free-refresh runs; a session-start notify reminded every session, which is noise.
  5. No resilience. No retry, no failover, no cooldown; one 429 ended the request.
  6. Static model params blocked the auto router. Pi model attributes (context window, max tokens) were fixed per registered model entry at the time; there was no way to give one synthetic router model per-backend parameters. (Human-reported; consistent with the single-entry-per-model design of the old code.)

Current Pi resolves 6: ProviderModelConfig supports per-model api and baseUrl overrides, registerProvider calls after load apply immediately (model sets replaceable live), custom streamSimple gives full per-request control, and message_end error normalization makes context-overflow recovery work for custom providers (custom-provider docs). Viable auto strategy: advertise the minimum context window across keyed backends, normalize upstream overflow errors to context_length_exceeded, and treat overflow like exhaustion in the cascade.

Points 1-4 are evidenced by the code and current provider docs; point 5's removal trigger (commit pairing) is inferred, not documented.

Lessons vs current spec

Old failure Gratis spec countermeasure
Keyless endpoints died GRATIS_*/native key chain; unkeyed backend invisible
Visible but broken models Register only keyed backends; no key → no provider
Three API-shaped providers One grts provider, backend-nested model IDs
Stale hardcoded catalogs Static curated + Zen live discovery of -free models
No failover grts/auto cascade with per-backend cooldown

New constraint discovered

Zen free models are split across two API types: big-pickle/minimax-class via anthropic-messages, others via openai-completions — and the split can change between catalog refreshes. The spec's backends table pins opencode to "OpenAI completions (Zen v1)"; correct model is per-model API type carried by each catalog entry, not per-backend API choice. Same applies in principle to every backend: catalog entries, not backends, own the API type.

Unresolved questions

  • Why exactly deletion happened on 2026-05-21: no note in commits; if session logs hold the actual breakage story, it would confirm inference 1 vs an unrelated cleanup motive.
  • Kilo re-evaluated separately: Kilo Gateway Free Access — verdict: include as the keyless v1 tier; the old Kilo failure was catalog rot, not auth.
---
id: PX-RESEARCH-LXU8HWNJ
type: research
title: Free Extension Post-Mortem
---

## Free Extension Post-Mortem

Prior attempt at the same idea as [Gratis Synthetic Free Provider](../../specs/PX-SPEC-CNDFHNO2-gratis-synthetic-free-provider/index.md): a `free` extension in the dotfiles Pi setup, April-May 2026.
Sources: dotfiles git history commits `5fbf178c` (added), `b87a61ad` (packaged), `5ca1296c` (reworked), `a6c52214` (deleted).

### What it was

Single extension registering three synthetic providers over keyless free endpoints:

| Synthetic provider | Upstream | API | Catalog |
| --- | --- | --- | --- |
| `free` | Kilo gateway (`api.kilo.ai/api/gateway/v1`) | openai-completions | hardcoded, manual `/free-refresh` re-fetch |
| `zen` | opencode.ai/zen (big-pickle, minimax-m2.5-free) | anthropic-messages | startup fetch of `zen/v1/models` |
| `zen-free` | opencode.ai/zen (hy3, ling, nemotron `-free`) | openai-completions | same startup fetch |

Model IDs leaked upstream paths (`kilo-auto/free`, `openrouter/free`, `nvidia/nemotron-3-super-120b-a12b:free`).
Auth was `apiKey: "-"` with `authHeader: false` — both endpoints served free models without accounts in April 2026.

### Timeline

```mermaid
flowchart LR
  add["2026-05-04 added<br/>5fbf178c"] --> pkg["2026-05-04 packaged<br/>b87a61ad"] --> rw["2026-05-09 reworked<br/>5ca1296c"] --> del["2026-05-21 deleted<br/>a6c52214"]
```

Deleted in a commit otherwise titled "config: add kimi coding model" (same change migrated all extensions to a new package scope and rewrote `models.json`) — no failure note, dropped as dead weight after ~2.5 weeks.

### Failure analysis

1. **Keyless assumption expired.** Free access without any account was the load-bearing design. Zen now requires login + API key; Kilo free models moved behind its own gateway account. When endpoints started rejecting unauthenticated calls, every registered model broke at request time.
2. **Visible but broken, the worst failure mode.** Fake key `-` meant providers always registered. Stale catalogs and dead endpoints surfaced as per-request errors, never as absence.
3. **Provider sprawl.** Three synthetic providers named after API shapes (`zen` vs `zen-free`) forced users to know transport details to pick a model.
4. **Catalog churn.** Hardcoded Kilo list rotted between manual `/free-refresh` runs; a session-start notify reminded every session, which is noise.
5. **No resilience.** No retry, no failover, no cooldown; one 429 ended the request.
6. **Static model params blocked the auto router.** Pi model attributes (context window, max tokens) were fixed per registered model entry at the time; there was no way to give one synthetic router model per-backend parameters. (Human-reported; consistent with the single-entry-per-model design of the old code.)

Current Pi resolves 6: `ProviderModelConfig` supports per-model `api` and `baseUrl` overrides, `registerProvider` calls after load apply immediately (model sets replaceable live), custom `streamSimple` gives full per-request control, and `message_end` error normalization makes context-overflow recovery work for custom providers ([custom-provider docs](https://github.com/earendil-works/pi/blob/main/packages/ai/test)). Viable auto strategy: advertise the minimum context window across keyed backends, normalize upstream overflow errors to `context_length_exceeded`, and treat overflow like exhaustion in the cascade.

Points 1-4 are evidenced by the code and current provider docs; point 5's removal trigger (commit pairing) is inferred, not documented.

### Lessons vs current spec

| Old failure | Gratis spec countermeasure |
| --- | --- |
| Keyless endpoints died | GRATIS_*/native key chain; unkeyed backend invisible |
| Visible but broken models | Register only keyed backends; no key → no provider |
| Three API-shaped providers | One `grts` provider, backend-nested model IDs |
| Stale hardcoded catalogs | Static curated + Zen live discovery of `-free` models |
| No failover | `grts/auto` cascade with per-backend cooldown |

### New constraint discovered

Zen free models are split across two API types: `big-pickle`/`minimax`-class via anthropic-messages, others via openai-completions — and the split can change between catalog refreshes.
The spec's backends table pins opencode to "OpenAI completions (Zen v1)"; correct model is per-model API type carried by each catalog entry, not per-backend API choice.
Same applies in principle to every backend: catalog entries, not backends, own the API type.

### Unresolved questions

- Why exactly deletion happened on 2026-05-21: no note in commits; if session logs hold the actual breakage story, it would confirm inference 1 vs an unrelated cleanup motive.
- Kilo re-evaluated separately: [Kilo Gateway Free Access](../PX-RESEARCH-KZ2M6WQ4-kilo-gateway-free-access/index.md) — verdict: include as the keyless v1 tier; the old Kilo failure was catalog rot, not auth.