# gratis Synthetic `grts` provider exposing curated free-tier LLM backends as normal Pi models. Spec: `PX-SPEC-CNDFHNO2` (Gratis Synthetic Free Provider). ## Install / load Loaded through root [pi-ext](../../README.md) package. ## Models - `grts//` pins one backend: google, zai, groq, mistral, nvidia, openrouter, hetzner, or kilo. - `grts/auto` cascades google → zai → groq → mistral → nvidia → openrouter → hetzner → kilo, answer quality first, trying each backend's top candidates before moving on. - `grts/auto` never waits twice on a congested backend: when a candidate streams nothing within 10 s, it skips that backend for 10 minutes. - A candidate that hit a rate limit is skipped for a minute; one the account cannot use (404) for an hour; a backend whose candidates all failed cools for its rate or quota window. - Cooling windows are estimates: before failing, `grts/auto` gives each skipped or slow backend's best candidate one more try, without a deadline. - Routing learns from real requests only, never probes: a model that usually starts fast gets a shorter deadline (3× its typical first token, 3–10 s), and models or backends that keep failing move behind their peers until the failures age out. - `/gratis cool` also votes against the model; two recent votes keep it behind its siblings for days after the cooldown. A backend with at least 20 clean requests moves up one place past an unreliable neighbor; speed never promotes. - What routing learned persists across restarts in `$XDG_STATE_HOME/pi/gratis/routing.json`, else `~/.local/state/pi/gratis/routing.json` (macOS `~/Library/Application Support/pi/gratis/`, Windows `%LOCALAPPDATA%\pi\gratis\`), owner-only; delete it to start fresh. - A backend appears only when usable; with no usable backend `grts` lists no models. - Every model costs 0 and its name carries the free-tier data caveat: free tiers may log or train on prompts. - Each reply records the model that actually answered as Pi's `responseModel`, `/`, including what virtual routers such as `kilo-auto/free` resolved to; `/session` groups usage by it. - Catalogs: [src/catalog.ts](src/catalog.ts); cascade: [src/router.ts](src/router.ts). ## Keys - **Need no key:** kilo works anonymously, about 200 requests per hour per IP. - **Benefit from a key:** kilo; a key lifts it above the anonymous limit and off shared office or VPN IPs. - **Need a key:** google, nvidia, openrouter, groq, mistral, zai, hetzner; each stays hidden until its key is set. | Variable | Native fallback | Backend | How to get it | | --- | --- | --- | --- | | `GRATIS_KILO_TOKEN` | `KILO_API_KEY` | kilo (optional) | Sign in at [app.kilo.ai](https://app.kilo.ai), open **Your Profile** on your personal account, copy the API key. | | `GRATIS_GOOGLE_TOKEN` | `GEMINI_API_KEY` | google | Sign in to [Google AI Studio](https://aistudio.google.com/apikey) with a Google account and create an API key; no card needed. | | `GRATIS_NVIDIA_TOKEN` | `NVIDIA_API_KEY` | nvidia | Sign in at [build.nvidia.com/settings](https://build.nvidia.com/settings) and generate an API key (`nvapi-…`); free trial tier, no card needed. | | `GRATIS_OPENROUTER_TOKEN` | `OPENROUTER_API_KEY` | openrouter | Create an account and a key at [openrouter.ai/keys](https://openrouter.ai/keys); no card needed. | | `GRATIS_GROQ_TOKEN` | `GROQ_API_KEY` | groq | Create an account and a key at [console.groq.com/keys](https://console.groq.com/keys); no card needed. | | `GRATIS_MISTRAL_TOKEN` | `MISTRAL_API_KEY` | mistral | Activate Mistral Studio in Free mode (phone verification), then create a key at [admin.mistral.ai](https://admin.mistral.ai/organization/api-keys). | | `GRATIS_ZAI_TOKEN` | `ZAI_API_KEY` | zai | Create a key at [z.ai/manage-apikey](https://z.ai/manage-apikey/apikey-list). | | `GRATIS_HETZNER_TOKEN` | `HETZNER_INFERENCE_API_KEY` | hetzner | Sign in at [experiments.hetzner.com](https://experiments.hetzner.com), open **Apps → Inference**, click **Create API Token**. | Z.ai: gratis uses only the models Z.ai prices Free (`glm-4.7-flash`, `glm-4.5-flash`) on the general endpoint, never the Coding Plan endpoints, so a Coding Plan key never spends plan quota through gratis. Hetzner is free only while its Inference API stays experimental. Per backend, first match wins: 1. The `GRATIS_*` variable. 2. The native variable. 3. Pi's credential store for the native provider (`google`, `nvidia`, `openrouter`, `groq`, `mistral`, `zai`, or a `models.json` `hetzner` entry), resolved after session start. gratis never prompts for, stores, or logs keys. ## Catalog refresh Kilo, OpenRouter, NVIDIA, and Hetzner model lists come from their model endpoints through Pi's model-catalog refresh; only Hetzner's needs the key. Google, Groq, Mistral, and NVIDIA models come from Pi's own catalog, which Pi refreshes at runtime, filtered by each backend's free rule: concrete Gemini Flash and Flash-Lite ids, Groq chat models, Mistral `-latest` text models, and NVIDIA models Pi prices at 0 that NVIDIA's public list serves. Z.ai lists the models its pricing page marks Free for input and output; an id gratis does not know also needs models.dev at cost 0 on the general endpoint. A failed list fetch keeps the last confirmed list; the curated lists in [src/catalog.ts](src/catalog.ts) apply until a session exists or a source never answered. Interactive mode and `/model` refresh from the network; print and RPC startup reuse Pi's stored catalog. Unreachable Kilo means no Kilo models; other unreachable lists fall back to their curated entries. ## Commands / tools / settings - Commands: `/gratis` shows visible backends with model counts, backends missing a key, cooling backends with time left, the last route (requested → served model, outcome, hops, what `grts/auto` skipped), and measured models (first token, tokens per second, failure rate, demotion). `/gratis cool [model] [minutes]` makes `grts/auto` skip a model, by default the one that answered last, for 60 minutes (1–1440), including its last-resort pass; `/gratis uncool [model]` ends one or all such cooldowns. Behind a virtual router such as `kilo-auto/free`, gratis can only skip the router itself. Pinned models stay usable, and manual cooldowns end when Pi restarts. - Tools: none. - Settings: none; `GRATIS_*` env variables are the only configuration. - Hooks/events: `session_start`, `message_end` (normalizes unrecognized overflow errors so Pi compacts), `session_shutdown`. ## Live checks `mise run //extensions/gratis:live` sends one tiny pinned request to each backend's real endpoint, spending free quota. Keyed backends run only when one of their env keys is set and skip otherwise; keyless Kilo always runs. The suite lives in [__tests__/backends.live.ts](__tests__/backends.live.ts), outside `//:test`. ## Benchmarks `mise run //extensions/gratis:bench` prints Go benchmark lines; compare saved runs with `//extensions/gratis:benchstat`, and use `//extensions/gratis:profile` for a CPU profile. [__bench__/gratis.bench.ts](__bench__/gratis.bench.ts) covers `getModels`, live-catalog parsing, and 512-chunk streams (direct transport, pinned, auto) against a local SSE server. ## Debug Opt-in metadata diagnostics: [debug contract](../DEBUG.md). Safe events: `session.start`, `session.shutdown`, `keys.resolve`, spans `catalog.refresh` and `request` with route, outcome, backend, and counts.