grts/<backend>/<model> pins one backend: google, zai, groq, mistral, nvidia, openrouter, hetzner, or kilo.
grts/auto cascades google → zai → groq → mistral → nvidia → openrouter → hetzner → kilo, answer quality first, trying each backend's top candidates before moving on.
grts/auto never waits twice on a congested backend: when a candidate streams nothing within 10 s, it skips that backend for 10 minutes.
A candidate that hit a rate limit is skipped for a minute; one the account cannot use (404) for an hour; a backend whose candidates all failed cools for its rate or quota window.
Cooling windows are estimates: before failing, grts/auto gives each skipped or slow backend's best candidate one more try, without a deadline.
Routing learns from real requests only, never probes: a model that usually starts fast gets a shorter deadline (3× its typical first token, 3–10 s), and models or backends that keep failing move behind their peers until the failures age out.
/gratis cool also votes against the model; two recent votes keep it behind its siblings for days after the cooldown. A backend with at least 20 clean requests moves up one place past an unreliable neighbor; speed never promotes.
What routing learned persists across restarts in $XDG_STATE_HOME/pi/gratis/routing.json, else ~/.local/state/pi/gratis/routing.json (macOS ~/Library/Application Support/pi/gratis/, Windows %LOCALAPPDATA%\pi\gratis\), owner-only; delete it to start fresh.
A backend appears only when usable; with no usable backend grts lists no models.
Every model costs 0 and its name carries the free-tier data caveat: free tiers may log or train on prompts.
Each reply records the model that actually answered as Pi's responseModel, <backend>/<model>, including what virtual routers such as kilo-auto/free resolved to; /session groups usage by it.
Z.ai: gratis uses only the models Z.ai prices Free (glm-4.7-flash, glm-4.5-flash) on the general endpoint, never the Coding Plan endpoints, so a Coding Plan key never spends plan quota through gratis.
Hetzner is free only while its Inference API stays experimental.
Per backend, first match wins:
The GRATIS_* variable.
The native variable.
Pi's credential store for the native provider (google, nvidia, openrouter, groq, mistral, zai, or a models.jsonhetzner entry), resolved after session start.
gratis never prompts for, stores, or logs keys.
Catalog refresh
Kilo, OpenRouter, NVIDIA, and Hetzner model lists come from their model endpoints through Pi's model-catalog refresh; only Hetzner's needs the key.
Google, Groq, Mistral, and NVIDIA models come from Pi's own catalog, which Pi refreshes at runtime, filtered by each backend's free rule: concrete Gemini Flash and Flash-Lite ids, Groq chat models, Mistral -latest text models, and NVIDIA models Pi prices at 0 that NVIDIA's public list serves.
Z.ai lists the models its pricing page marks Free for input and output; an id gratis does not know also needs models.dev at cost 0 on the general endpoint.
A failed list fetch keeps the last confirmed list; the curated lists in src/catalog.ts apply until a session exists or a source never answered.
Interactive mode and /model refresh from the network; print and RPC startup reuse Pi's stored catalog.
Unreachable Kilo means no Kilo models; other unreachable lists fall back to their curated entries.
Commands / tools / settings
Commands: /gratis shows visible backends with model counts, backends missing a key, cooling backends with time left, the last route (requested → served model, outcome, hops, what grts/auto skipped), and measured models (first token, tokens per second, failure rate, demotion).
/gratis cool [model] [minutes] makes grts/auto skip a model, by default the one that answered last, for 60 minutes (1–1440), including its last-resort pass; /gratis uncool [model] ends one or all such cooldowns. Behind a virtual router such as kilo-auto/free, gratis can only skip the router itself. Pinned models stay usable, and manual cooldowns end when Pi restarts.
Tools: none.
Settings: none; GRATIS_* env variables are the only configuration.
Hooks/events: session_start, message_end (normalizes unrecognized overflow errors so Pi compacts), session_shutdown.
Live checks
mise run //extensions/gratis:live sends one tiny pinned request to each backend's real endpoint, spending free quota.
Keyed backends run only when one of their env keys is set and skip otherwise; keyless Kilo always runs.
The suite lives in tests/backends.live.ts, outside //:test.
Benchmarks
mise run //extensions/gratis:bench prints Go benchmark lines; compare saved runs with //extensions/gratis:benchstat, and use //extensions/gratis:profile for a CPU profile.
bench/gratis.bench.ts covers getModels, live-catalog parsing, and 512-chunk streams (direct transport, pinned, auto) against a local SSE server.
Debug
Opt-in metadata diagnostics: debug contract.
Safe events: session.start, session.shutdown, keys.resolve, spans catalog.refresh and request with route, outcome, backend, and counts.
# gratis
Synthetic `grts` provider exposing curated free-tier LLM backends as normal Pi models.
Spec: `PX-SPEC-CNDFHNO2` (Gratis Synthetic Free Provider).
## Install / load
Loaded through root [pi-ext](../../README.md) package.
## Models
- `grts/<backend>/<model>` pins one backend: google, zai, groq, mistral, nvidia, openrouter, hetzner, or kilo.
- `grts/auto` cascades google → zai → groq → mistral → nvidia → openrouter → hetzner → kilo, answer quality first, trying each backend's top candidates before moving on.
- `grts/auto` never waits twice on a congested backend: when a candidate streams nothing within 10 s, it skips that backend for 10 minutes.
- A candidate that hit a rate limit is skipped for a minute; one the account cannot use (404) for an hour; a backend whose candidates all failed cools for its rate or quota window.
- Cooling windows are estimates: before failing, `grts/auto` gives each skipped or slow backend's best candidate one more try, without a deadline.
- Routing learns from real requests only, never probes: a model that usually starts fast gets a shorter deadline (3× its typical first token, 3–10 s), and models or backends that keep failing move behind their peers until the failures age out.
- `/gratis cool` also votes against the model; two recent votes keep it behind its siblings for days after the cooldown. A backend with at least 20 clean requests moves up one place past an unreliable neighbor; speed never promotes.
- What routing learned persists across restarts in `$XDG_STATE_HOME/pi/gratis/routing.json`, else `~/.local/state/pi/gratis/routing.json` (macOS `~/Library/Application Support/pi/gratis/`, Windows `%LOCALAPPDATA%\pi\gratis\`), owner-only; delete it to start fresh.
- A backend appears only when usable; with no usable backend `grts` lists no models.
- Every model costs 0 and its name carries the free-tier data caveat: free tiers may log or train on prompts.
- Each reply records the model that actually answered as Pi's `responseModel`, `<backend>/<model>`, including what virtual routers such as `kilo-auto/free` resolved to; `/session` groups usage by it.
- Catalogs: [src/catalog.ts](src/catalog.ts); cascade: [src/router.ts](src/router.ts).
## Keys
- **Need no key:** kilo works anonymously, about 200 requests per hour per IP.
- **Benefit from a key:** kilo; a key lifts it above the anonymous limit and off shared office or VPN IPs.
- **Need a key:** google, nvidia, openrouter, groq, mistral, zai, hetzner; each stays hidden until its key is set.
| Variable | Native fallback | Backend | How to get it |
| --- | --- | --- | --- |
| `GRATIS_KILO_TOKEN` | `KILO_API_KEY` | kilo (optional) | Sign in at [app.kilo.ai](https://app.kilo.ai), open **Your Profile** on your personal account, copy the API key. |
| `GRATIS_GOOGLE_TOKEN` | `GEMINI_API_KEY` | google | Sign in to [Google AI Studio](https://aistudio.google.com/apikey) with a Google account and create an API key; no card needed. |
| `GRATIS_NVIDIA_TOKEN` | `NVIDIA_API_KEY` | nvidia | Sign in at [build.nvidia.com/settings](https://build.nvidia.com/settings) and generate an API key (`nvapi-…`); free trial tier, no card needed. |
| `GRATIS_OPENROUTER_TOKEN` | `OPENROUTER_API_KEY` | openrouter | Create an account and a key at [openrouter.ai/keys](https://openrouter.ai/keys); no card needed. |
| `GRATIS_GROQ_TOKEN` | `GROQ_API_KEY` | groq | Create an account and a key at [console.groq.com/keys](https://console.groq.com/keys); no card needed. |
| `GRATIS_MISTRAL_TOKEN` | `MISTRAL_API_KEY` | mistral | Activate Mistral Studio in Free mode (phone verification), then create a key at [admin.mistral.ai](https://admin.mistral.ai/organization/api-keys). |
| `GRATIS_ZAI_TOKEN` | `ZAI_API_KEY` | zai | Create a key at [z.ai/manage-apikey](https://z.ai/manage-apikey/apikey-list). |
| `GRATIS_HETZNER_TOKEN` | `HETZNER_INFERENCE_API_KEY` | hetzner | Sign in at [experiments.hetzner.com](https://experiments.hetzner.com), open **Apps → Inference**, click **Create API Token**. |
Z.ai: gratis uses only the models Z.ai prices Free (`glm-4.7-flash`, `glm-4.5-flash`) on the general endpoint, never the Coding Plan endpoints, so a Coding Plan key never spends plan quota through gratis.
Hetzner is free only while its Inference API stays experimental.
Per backend, first match wins:
1. The `GRATIS_*` variable.
2. The native variable.
3. Pi's credential store for the native provider (`google`, `nvidia`, `openrouter`, `groq`, `mistral`, `zai`, or a `models.json` `hetzner` entry), resolved after session start.
gratis never prompts for, stores, or logs keys.
## Catalog refresh
Kilo, OpenRouter, NVIDIA, and Hetzner model lists come from their model endpoints through Pi's model-catalog refresh; only Hetzner's needs the key.
Google, Groq, Mistral, and NVIDIA models come from Pi's own catalog, which Pi refreshes at runtime, filtered by each backend's free rule: concrete Gemini Flash and Flash-Lite ids, Groq chat models, Mistral `-latest` text models, and NVIDIA models Pi prices at 0 that NVIDIA's public list serves.
Z.ai lists the models its pricing page marks Free for input and output; an id gratis does not know also needs models.dev at cost 0 on the general endpoint.
A failed list fetch keeps the last confirmed list; the curated lists in [src/catalog.ts](src/catalog.ts) apply until a session exists or a source never answered.
Interactive mode and `/model` refresh from the network; print and RPC startup reuse Pi's stored catalog.
Unreachable Kilo means no Kilo models; other unreachable lists fall back to their curated entries.
## Commands / tools / settings
- Commands: `/gratis` shows visible backends with model counts, backends missing a key, cooling backends with time left, the last route (requested → served model, outcome, hops, what `grts/auto` skipped), and measured models (first token, tokens per second, failure rate, demotion).
`/gratis cool [model] [minutes]` makes `grts/auto` skip a model, by default the one that answered last, for 60 minutes (1–1440), including its last-resort pass; `/gratis uncool [model]` ends one or all such cooldowns. Behind a virtual router such as `kilo-auto/free`, gratis can only skip the router itself. Pinned models stay usable, and manual cooldowns end when Pi restarts.
- Tools: none.
- Settings: none; `GRATIS_*` env variables are the only configuration.
- Hooks/events: `session_start`, `message_end` (normalizes unrecognized overflow errors so Pi compacts), `session_shutdown`.
## Live checks
`mise run //extensions/gratis:live` sends one tiny pinned request to each backend's real endpoint, spending free quota.
Keyed backends run only when one of their env keys is set and skip otherwise; keyless Kilo always runs.
The suite lives in [__tests__/backends.live.ts](__tests__/backends.live.ts), outside `//:test`.
## Benchmarks
`mise run //extensions/gratis:bench` prints Go benchmark lines; compare saved runs with `//extensions/gratis:benchstat`, and use `//extensions/gratis:profile` for a CPU profile.
[__bench__/gratis.bench.ts](__bench__/gratis.bench.ts) covers `getModels`, live-catalog parsing, and 512-chunk streams (direct transport, pinned, auto) against a local SSE server.
## Debug
Opt-in metadata diagnostics: [debug contract](../DEBUG.md).
Safe events: `session.start`, `session.shutdown`, `keys.resolve`, spans `catalog.refresh` and `request` with route, outcome, backend, and counts.