Extension gratis registers one synthetic Pi provider grts that exposes curated free-tier LLM backends as normal models.
Intent: free inference with zero gratis-specific configuration; invisible when unusable.
grts/auto sells on ease of use and a comfortable, fast experience; waiting, disruption, or annoyance in it is a defect.
Nested per-backend model IDs plus one synthetic router model.
<backend> is one of kilo, google, nvidia, openrouter, groq, mistral, zai, hetzner.
grts/<backend>/<model> pinned: request routes to that backend only
grts/auto router: request cascades across available backends
grts/auto follows backend priority: google → zai → groq → mistral → nvidia → openrouter → hetzner → kilo.
Order is by answer quality first, because a weak answer is the worst annoyance: Gemini Flash, GLM Flash, gpt-oss-120b, Devstral.
NVIDIA serves the largest free models (Kimi, GLM 5.3) but trails Mistral: measured twice, those models queued past 30–60 s or returned 404 for the account while only gpt-oss-20b answered fast, so leading with NVIDIA would cost a deadline wait.
OpenRouter's free pool is mixed and capped near 50 requests per day; Hetzner is experimental at 10 requests per minute.
Kilo trails: it needs no key, so it is the floor that always exists, but its anonymous router often lands on weak models.
Within the kilo backend, virtual router models (openrouter/free, kilo-auto/free) rank first: they absorb upstream per-model saturation server-side (verified: pinned models 429 while virtual models answered in the same minute).
Auto model parameters
Pi holds model attributes per registered model entry, not per request (post-mortem failure 6).
grts/auto therefore advertises the minimum context window and max output tokens across its candidate set, and the transport layer may tighten parameters per hop.
An upstream context-overflow error on grts/auto normalizes to a recognized overflow error and advances the cascade exactly like exhaustion; pinned models surface it honestly.
Key resolution
Per backend, evaluated independently at extension load:
GRATIS_<BACKEND>_TOKEN (for example GRATIS_OPENROUTER_TOKEN)
Native env of the backend (OPENROUTER_API_KEY, GEMINI_API_KEY, NVIDIA_API_KEY, GROQ_API_KEY, MISTRAL_API_KEY, ZAI_API_KEY, HETZNER_INFERENCE_API_KEY)
Pi's credential resolution for the backend's native provider id (modelRegistry auth lookup), which covers keys stored in the auth store and OAuth — env-only access is not required
none
A backend that resolves a key by any step → its grts/<backend>/* models register.
No backend resolves a key → grts never registers → provider invisible in /model.
Partial keys → only resolved backends' models appear.
Kilo exception: kilo is keyless — its free models register whenever the catalog is reachable, no key required.
GRATIS_KILO_TOKEN or KILO_API_KEY optionally raise it above anonymous 200 req/h/IP and decouple it from shared egress IPs.
Keys are read from process env only; gratis never prompts, stores, or logs key material.
Transport and backends
API type is a per-catalog-entry property, never a backend property; a single backend can serve several transports and the split can shift between catalog refreshes (post-mortem).
All API types Pi implements are in scope from the start — the google backend alone needs its own API shape — and gratis does not assume openai-completions suffices.
A catalog entry carries: model id, API type, endpoint base, context window, max output tokens, input modalities, reasoning flag.
live discovery via unauthenticated models endpoint
google
Google account key
concrete Flash and Flash-Lite ids (gemini-<version>-flash[-lite])
Pi's runtime catalog; static list as fallback
openrouter
OpenRouter key
:free variants
live discovery via public models endpoint; static list as fallback
groq
Groq key
every chat model on the free tier
Pi's runtime catalog; static list as fallback
mistral
Mistral key
every -latest text model on Studio Free mode
Pi's runtime catalog; static list as fallback
nvidia
NVIDIA API key
models Pi's catalog prices at 0
Pi's runtime catalog, intersected with the public models endpoint; static list as fallback
zai
Z.ai API key
models Z.ai's pricing page lists as Free for input and output
live: first-party pricing page, confirmed by models.dev for ids gratis does not know; static list as fallback, general endpoint only
hetzner
Hetzner Inference token
every model while experimental
live discovery via keyed models endpoint; static list as fallback
Free-set identification never relies on a suffix alone: discovery intersects live lists with a gratis-maintained free-set of names and patterns, updated with the catalog.
Virtual router models advertise their published catalog parameters when present (kilo-auto/free: 256K ctx, 10K out); conservative fallback values apply only when nothing is published.
Static catalogs are extension data; stale entries are corrected by catalog update, never by user config.
Keeping catalogs fresh
Free pools churn, so no backend relies on gratis releases alone when a runtime source exists.
Pi refreshes its built-in provider catalogs at runtime (pi.dev, every 4 hours); gratis reads them in memory through Pi's registry for google, groq, mistral, and nvidia, and never through a call that lists its own models.
Those catalogs supply API type, limits, modalities, and thinking levels; gratis applies its free rule per backend, zeroes cost, and ranks its known preferences first, then the rest by version and context.
Until a session exists, or when Pi's catalog has no match, the static list applies.
Z.ai publishes no free flag in any API, and its models endpoint omits the free models; its first-party pricing page is the free authority.
Because a wrong free claim bills the user, an unknown Z.ai id is added only when models.dev also lists it at cost 0 under the general endpoint; a known id stays only while the page lists it as Free.
A reachable page without Free rows yields no Z.ai models; an unreachable or unparseable page keeps the last confirmed list.
Live lists still expire at runtime: a listed model that the account cannot use is skipped as unavailable.
Failover
Cascade applies to grts/auto only; pinned models surface errors honestly.
Cascade operates at two granularities: models within a backend, then backends.
Within an available backend, the router tries candidate models in catalog rank — virtual router models first — before cooling the backend; a 429 or quota response on one candidate advances to the next candidate.
A candidate that fails with a rate limit, quota, or its first-response deadline cools on its own for a short window and is skipped meanwhile, so a saturated model costs no round trip on the next request.
Only when every candidate fails does the backend cool until its window resets.
Each hop gets a first-response deadline; a hop that streams nothing before it marks its backend congested: the cascade skips that backend's remaining candidates and cools it, since free queues stall whole backends (measured: NVIDIA's large free models queued past 60 s while gpt-oss-20b answered in 0.6 s). Transport timeouts and network errors advance like rate limits.
A listed candidate the account cannot use (404, not found) is skipped for a long window and the cascade moves to the next candidate; pinned models surface it honestly.
An upstream context-overflow error on grts/auto advances the cascade without cooling; the router retries the same prompt on the next candidate or backend.
Cooling windows are estimates: when the first pass finds no answer, a last-resort pass retries, once per backend and in priority order, the best candidate that was skipped as cooling or timed out, without a deadline.
The request fails only when the last-resort pass also fails; the error then lists exhausted backends.
Performance-aware routing
Quality order stays primary; measured performance only adjusts it, because the fastest free models are usually the weakest.
Measurement is passive: every real hop, pinned or auto, records time to first token, output tokens per second, and whether it failed; gratis never sends probe requests, which would spend free budgets.
Measurements persist across restarts in a small owner-only state file (<user state dir>/pi/gratis/routing.json, XDG_STATE_HOME honored), loaded in the background so startup never waits on it; a missing or corrupt file starts fresh.
A model's first-response deadline adapts to its own recent time to first token (three times typical, within 3–10 s) once it has enough samples.
A model or backend that keeps failing (rate, quota, silence, unavailable) moves behind its peers in the order and recovers as failures age or it succeeds again.
A manual cool is also a quality vote against that model; votes fade over days rather than minutes, so a model the user rejected repeatedly stays behind its peers after its cooldown ends.
A backend with a long clean record moves up one place past an unreliable neighbor, at most one place; speed never promotes, so a fast but weak backend cannot climb.
The status command shows per-model time to first token, tokens per second, and failure rate.
Manual cooling
The user can make grts/auto skip a model on demand: by default the candidate that answered the last request, for 60 minutes (1–1440), or a named current grts model.
A manually cooled model is also excluded from the last-resort pass, because the user rejected it; pinned use of the same model is unaffected.
Behind a virtual router gratis cools the router candidate itself and says so, since it cannot steer the upstream's choice.
Manual cooldowns can be ended individually or all at once, show in the status command, and live in memory like the rest of routing state.
stateDiagram-v2
[*] --> Available: backend available at load
Available --> Cooling: 429 or quota exhausted
Cooling --> Available: rate window resets
Available --> [*]: unavailable at next load
Cooling --> [*]: unavailable at next load
flowchart LR
req["grts/auto request"] --> pick{"backend available<br/>and not cooling?"}
pick -->|yes| model["pick candidate model:<br/>virtual first"]
pick -->|no| next["next backend"]
next --> pick
model --> send["send"]
send -->|success| done["stream response"]
send -->|429 or quota,<br/>candidates remain| model
send -->|429 or quota,<br/>candidates exhausted| cool["cool backend"] --> next
send -->|overflow| next
send -->|no output before deadline| cool
send -->|other error| err["surface error"]
next -->|no backend left| last["last-resort pass:<br/>skipped and timed-out,<br/>no deadline"]
last -->|success| done
last -->|all fail| fail["fail, list backends"]
Constraints
Every routed model costs 0; gratis never emits paid-model fallbacks.
Free tiers commonly log or train on prompts; grts models are labeled in their description so users can avoid them for proprietary code.
Kilo's free pool rotates hard (post-mortem: every April 2026 ID is gone); discovery must tolerate arbitrary churn and empty pools.
Rate-limit and quota responses arrive status-wrapped with provider text in the body; failure classification inspects status and body text, never status alone.
Kilo exposes no rate-limit response headers; budget awareness is observation-based (cool on observed 429), never header-based.
opencode and Zen are rejected backends: the free tier is locked to the OpenCode client (403 FreeTierError even on a valid key); a Zen key grants paid models only. Neither returns for gratis unless OpenCode changes the lock (research: OpenCode Free Tier Lock).
Rate-limit metadata in catalogs is advisory, never user-facing config.
Z.ai: gratis calls only the general endpoint https://api.z.ai/api/paas/v4 and only models Z.ai prices Free; it never calls the Coding Plan endpoints (/api/coding/paas/v4, /api/anthropic), so a Coding Plan key never spends plan quota through gratis. "Flash" in a name is not proof of free: Z.ai glm-5.3-flash is paid.
NVIDIA: only models Pi's catalog prices at 0 are curated; priced NVIDIA models never register.
Hetzner: free only while the Inference API is experimental; the catalog follows /v1/models, which Hetzner declares definitive.
No new Pi config surface: GRATIS_* env vars are the only configuration gratis adds; native env and Pi's credential store are read, never written.
Acceptance criteria
With no GRATIS-style and native keys set and network up, grts/kilo/* models appear and grts/auto includes kilo; all other grts backends stay invisible.
With no keys set and the Kilo catalog unreachable, grts does not appear in provider or model lists.
With exactly one native key set, only that backend's and kilo's models appear under grts/.
GRATIS_OPENROUTER_TOKEN overrides a simultaneously set native OpenRouter key.
A key stored in Pi's credential store for a native provider, with no env var set, registers the matching grts backend.
grts/auto on an exhausted first-priority backend completes on the next available backend without user action.
A grts/auto hop that streams nothing before its first-response deadline advances to the next candidate instead of waiting.
A candidate that just hit a rate limit is skipped on the next grts/auto request while its short cooldown runs.
With every backend cooling, grts/auto still completes when a cooled candidate answers in the last-resort pass.
A model with a short measured time to first token gets a proportionally shorter first-response deadline, never below 3 s or above 10 s.
A model that failed repeatedly is tried after its healthy siblings once its cooldown ends, and returns to its rank after its failures age out.
After two manual cools, a model is still tried after its siblings once both cooldowns ended.
A demotion learned before a restart still applies after it.
A backend with a long clean record moves up exactly one place past an unreliable neighbor and never further.
Routing sends no request that the user's traffic did not cause.
A model that appears in Pi's runtime catalog for google, groq, mistral, or nvidia and matches that backend's free rule appears under grts without a gratis update.
A Z.ai model newly listed as Free on the pricing page and at cost 0 on models.dev appears; one that stops being listed as Free disappears; an unreachable page keeps the last confirmed list.
After the user cools the last-answering model, the next grts/auto request skips it, including in the last-resort pass, until the cooldown ends or the user lifts it; a pinned request to it still answers.
A context-overflow response during an grts/auto request advances the cascade; it never triggers cooling or fails while a next backend remains.
A pinned model that 429s returns the upstream error, not a cascade.
A pinned Kilo model returning upstream-saturation 429 while a virtual router model is available completes on that model without cooling the kilo backend.
Virtual models advertise published catalog parameters when present.
Catalog entries with non-openai-completions API types (anthropic-messages, google-shaped) stream correctly without conversion assumptions.
OpenRouter discovery with zero :free variants and Kilo discovery with zero :free models each yield no models for that backend and no error.
All registered models report input/output cost 0 and a free-tier data caveat in their description.
With only NVIDIA_API_KEY set, NVIDIA models appear only for curated ids present in the live models list; with the list unreachable, the curated list appears.
With only ZAI_API_KEY set, exactly the free Flash models appear, and requests go to the general endpoint, never a Coding Plan endpoint.
With only HETZNER_INFERENCE_API_KEY set, Hetzner models follow the keyed live models list; with the list unreachable, the static list appears.
Key-gated live tests run each backend's real endpoint only when that backend's key is set, skip otherwise, and stay outside //:test.
---
id: PX-SPEC-CNDFHNO2
type: spec
title: Gratis Synthetic Free Provider
research:
- PX-RESEARCH-KMBSJUZR
- PX-RESEARCH-LXU8HWNJ
- PX-RESEARCH-KZ2M6WQ4
- PX-RESEARCH-8WQ4KZ2M
- PX-RESEARCH-Q7N4M2K9
---
## Gratis Synthetic Free Provider
Extension `gratis` registers one synthetic Pi provider `grts` that exposes curated free-tier LLM backends as normal models.
Intent: free inference with zero gratis-specific configuration; invisible when unusable.
`grts/auto` sells on ease of use and a comfortable, fast experience; waiting, disruption, or annoyance in it is a defect.
### Scope
In: synthetic provider, model addressing, key resolution, backend visibility, failover cascade, free-model catalog for eight backends.
Out: paid models, key acquisition or signup flows, non-LLM free tiers (image, audio), usage accounting dashboards.
### Addressing
Nested per-backend model IDs plus one synthetic router model.
`<backend>` is one of `kilo`, `google`, `nvidia`, `openrouter`, `groq`, `mistral`, `zai`, `hetzner`.
```text
grts/<backend>/<model> pinned: request routes to that backend only
grts/auto router: request cascades across available backends
```
`grts/auto` follows backend priority: google → zai → groq → mistral → nvidia → openrouter → hetzner → kilo.
Order is by answer quality first, because a weak answer is the worst annoyance: Gemini Flash, GLM Flash, gpt-oss-120b, Devstral.
NVIDIA serves the largest free models (Kimi, GLM 5.3) but trails Mistral: measured twice, those models queued past 30–60 s or returned 404 for the account while only gpt-oss-20b answered fast, so leading with NVIDIA would cost a deadline wait.
OpenRouter's free pool is mixed and capped near 50 requests per day; Hetzner is experimental at 10 requests per minute.
Kilo trails: it needs no key, so it is the floor that always exists, but its anonymous router often lands on weak models.
Within the kilo backend, virtual router models (`openrouter/free`, `kilo-auto/free`) rank first: they absorb upstream per-model saturation server-side (verified: pinned models 429 while virtual models answered in the same minute).
#### Auto model parameters
Pi holds model attributes per registered model entry, not per request (post-mortem failure 6).
`grts/auto` therefore advertises the minimum context window and max output tokens across its candidate set, and the transport layer may tighten parameters per hop.
An upstream context-overflow error on `grts/auto` normalizes to a recognized overflow error and advances the cascade exactly like exhaustion; pinned models surface it honestly.
### Key resolution
Per backend, evaluated independently at extension load:
1. `GRATIS_<BACKEND>_TOKEN` (for example `GRATIS_OPENROUTER_TOKEN`)
2. Native env of the backend (`OPENROUTER_API_KEY`, `GEMINI_API_KEY`, `NVIDIA_API_KEY`, `GROQ_API_KEY`, `MISTRAL_API_KEY`, `ZAI_API_KEY`, `HETZNER_INFERENCE_API_KEY`)
3. Pi's credential resolution for the backend's native provider id (`modelRegistry` auth lookup), which covers keys stored in the auth store and OAuth — env-only access is not required
4. none
A backend that resolves a key by any step → its `grts/<backend>/*` models register.
No backend resolves a key → `grts` never registers → provider invisible in `/model`.
Partial keys → only resolved backends' models appear.
**Kilo exception:** kilo is keyless — its free models register whenever the catalog is reachable, no key required.
`GRATIS_KILO_TOKEN` or `KILO_API_KEY` optionally raise it above anonymous 200 req/h/IP and decouple it from shared egress IPs.
Keys are read from process env only; gratis never prompts, stores, or logs key material.
### Transport and backends
API type is a per-catalog-entry property, never a backend property; a single backend can serve several transports and the split can shift between catalog refreshes (post-mortem).
All API types Pi implements are in scope from the start — the google backend alone needs its own API shape — and gratis does not assume openai-completions suffices.
A catalog entry carries: model id, API type, endpoint base, context window, max output tokens, input modalities, reasoning flag.
| Backend | Auth | Free models | Catalog source |
| --- | --- | --- | --- |
| kilo | none (key optional) | `:free` suffix + virtual `kilo-auto/free`, `openrouter/free` | live discovery via unauthenticated models endpoint |
| google | Google account key | concrete Flash and Flash-Lite ids (`gemini-<version>-flash[-lite]`) | Pi's runtime catalog; static list as fallback |
| openrouter | OpenRouter key | `:free` variants | live discovery via public models endpoint; static list as fallback |
| groq | Groq key | every chat model on the free tier | Pi's runtime catalog; static list as fallback |
| mistral | Mistral key | every `-latest` text model on Studio Free mode | Pi's runtime catalog; static list as fallback |
| nvidia | NVIDIA API key | models Pi's catalog prices at 0 | Pi's runtime catalog, intersected with the public models endpoint; static list as fallback |
| zai | Z.ai API key | models Z.ai's pricing page lists as Free for input and output | live: first-party pricing page, confirmed by models.dev for ids gratis does not know; static list as fallback, general endpoint only |
| hetzner | Hetzner Inference token | every model while experimental | live discovery via keyed models endpoint; static list as fallback |
Free-set identification never relies on a suffix alone: discovery intersects live lists with a gratis-maintained free-set of names and patterns, updated with the catalog.
Virtual router models advertise their published catalog parameters when present (`kilo-auto/free`: 256K ctx, 10K out); conservative fallback values apply only when nothing is published.
Static catalogs are extension data; stale entries are corrected by catalog update, never by user config.
#### Keeping catalogs fresh
Free pools churn, so no backend relies on gratis releases alone when a runtime source exists.
Pi refreshes its built-in provider catalogs at runtime (pi.dev, every 4 hours); gratis reads them in memory through Pi's registry for google, groq, mistral, and nvidia, and never through a call that lists its own models.
Those catalogs supply API type, limits, modalities, and thinking levels; gratis applies its free rule per backend, zeroes cost, and ranks its known preferences first, then the rest by version and context.
Until a session exists, or when Pi's catalog has no match, the static list applies.
Z.ai publishes no free flag in any API, and its models endpoint omits the free models; its first-party pricing page is the free authority.
Because a wrong free claim bills the user, an unknown Z.ai id is added only when models.dev also lists it at cost 0 under the general endpoint; a known id stays only while the page lists it as Free.
A reachable page without Free rows yields no Z.ai models; an unreachable or unparseable page keeps the last confirmed list.
Live lists still expire at runtime: a listed model that the account cannot use is skipped as unavailable.
### Failover
Cascade applies to `grts/auto` only; pinned models surface errors honestly.
Cascade operates at two granularities: models within a backend, then backends.
Within an available backend, the router tries candidate models in catalog rank — virtual router models first — before cooling the backend; a 429 or quota response on one candidate advances to the next candidate.
A candidate that fails with a rate limit, quota, or its first-response deadline cools on its own for a short window and is skipped meanwhile, so a saturated model costs no round trip on the next request.
Only when every candidate fails does the backend cool until its window resets.
Each hop gets a first-response deadline; a hop that streams nothing before it marks its backend congested: the cascade skips that backend's remaining candidates and cools it, since free queues stall whole backends (measured: NVIDIA's large free models queued past 60 s while gpt-oss-20b answered in 0.6 s). Transport timeouts and network errors advance like rate limits.
A listed candidate the account cannot use (404, not found) is skipped for a long window and the cascade moves to the next candidate; pinned models surface it honestly.
An upstream context-overflow error on `grts/auto` advances the cascade without cooling; the router retries the same prompt on the next candidate or backend.
Cooling windows are estimates: when the first pass finds no answer, a last-resort pass retries, once per backend and in priority order, the best candidate that was skipped as cooling or timed out, without a deadline.
The request fails only when the last-resort pass also fails; the error then lists exhausted backends.
#### Performance-aware routing
Quality order stays primary; measured performance only adjusts it, because the fastest free models are usually the weakest.
Measurement is passive: every real hop, pinned or auto, records time to first token, output tokens per second, and whether it failed; gratis never sends probe requests, which would spend free budgets.
Measurements persist across restarts in a small owner-only state file (`<user state dir>/pi/gratis/routing.json`, `XDG_STATE_HOME` honored), loaded in the background so startup never waits on it; a missing or corrupt file starts fresh.
A model's first-response deadline adapts to its own recent time to first token (three times typical, within 3–10 s) once it has enough samples.
A model or backend that keeps failing (rate, quota, silence, unavailable) moves behind its peers in the order and recovers as failures age or it succeeds again.
A manual cool is also a quality vote against that model; votes fade over days rather than minutes, so a model the user rejected repeatedly stays behind its peers after its cooldown ends.
A backend with a long clean record moves up one place past an unreliable neighbor, at most one place; speed never promotes, so a fast but weak backend cannot climb.
The status command shows per-model time to first token, tokens per second, and failure rate.
#### Manual cooling
The user can make `grts/auto` skip a model on demand: by default the candidate that answered the last request, for 60 minutes (1–1440), or a named current `grts` model.
A manually cooled model is also excluded from the last-resort pass, because the user rejected it; pinned use of the same model is unaffected.
Behind a virtual router gratis cools the router candidate itself and says so, since it cannot steer the upstream's choice.
Manual cooldowns can be ended individually or all at once, show in the status command, and live in memory like the rest of routing state.
```mermaid
stateDiagram-v2
[*] --> Available: backend available at load
Available --> Cooling: 429 or quota exhausted
Cooling --> Available: rate window resets
Available --> [*]: unavailable at next load
Cooling --> [*]: unavailable at next load
```
```mermaid
flowchart LR
req["grts/auto request"] --> pick{"backend available<br/>and not cooling?"}
pick -->|yes| model["pick candidate model:<br/>virtual first"]
pick -->|no| next["next backend"]
next --> pick
model --> send["send"]
send -->|success| done["stream response"]
send -->|429 or quota,<br/>candidates remain| model
send -->|429 or quota,<br/>candidates exhausted| cool["cool backend"] --> next
send -->|overflow| next
send -->|no output before deadline| cool
send -->|other error| err["surface error"]
next -->|no backend left| last["last-resort pass:<br/>skipped and timed-out,<br/>no deadline"]
last -->|success| done
last -->|all fail| fail["fail, list backends"]
```
### Constraints
- Every routed model costs 0; gratis never emits paid-model fallbacks.
- Free tiers commonly log or train on prompts; `grts` models are labeled in their description so users can avoid them for proprietary code.
- Kilo's free pool rotates hard (post-mortem: every April 2026 ID is gone); discovery must tolerate arbitrary churn and empty pools.
- Rate-limit and quota responses arrive status-wrapped with provider text in the body; failure classification inspects status and body text, never status alone.
- Kilo exposes no rate-limit response headers; budget awareness is observation-based (cool on observed 429), never header-based.
- opencode and Zen are rejected backends: the free tier is locked to the OpenCode client (403 `FreeTierError` even on a valid key); a Zen key grants paid models only. Neither returns for gratis unless OpenCode changes the lock (research: OpenCode Free Tier Lock).
- Rate-limit metadata in catalogs is advisory, never user-facing config.
- Z.ai: gratis calls only the general endpoint `https://api.z.ai/api/paas/v4` and only models Z.ai prices Free; it never calls the Coding Plan endpoints (`/api/coding/paas/v4`, `/api/anthropic`), so a Coding Plan key never spends plan quota through gratis. "Flash" in a name is not proof of free: Z.ai `glm-5.3-flash` is paid.
- NVIDIA: only models Pi's catalog prices at 0 are curated; priced NVIDIA models never register.
- Hetzner: free only while the Inference API is experimental; the catalog follows `/v1/models`, which Hetzner declares definitive.
- No new Pi config surface: `GRATIS_*` env vars are the only configuration gratis adds; native env and Pi's credential store are read, never written.
### Acceptance criteria
- [ ] With no GRATIS-style and native keys set and network up, `grts/kilo/*` models appear and `grts/auto` includes kilo; all other `grts` backends stay invisible.
- [ ] With no keys set and the Kilo catalog unreachable, `grts` does not appear in provider or model lists.
- [ ] With exactly one native key set, only that backend's and kilo's models appear under `grts/`.
- [ ] `GRATIS_OPENROUTER_TOKEN` overrides a simultaneously set native OpenRouter key.
- [ ] A key stored in Pi's credential store for a native provider, with no env var set, registers the matching `grts` backend.
- [ ] `grts/auto` on an exhausted first-priority backend completes on the next available backend without user action.
- [ ] A `grts/auto` hop that streams nothing before its first-response deadline advances to the next candidate instead of waiting.
- [ ] A candidate that just hit a rate limit is skipped on the next `grts/auto` request while its short cooldown runs.
- [ ] With every backend cooling, `grts/auto` still completes when a cooled candidate answers in the last-resort pass.
- [ ] A model with a short measured time to first token gets a proportionally shorter first-response deadline, never below 3 s or above 10 s.
- [ ] A model that failed repeatedly is tried after its healthy siblings once its cooldown ends, and returns to its rank after its failures age out.
- [ ] After two manual cools, a model is still tried after its siblings once both cooldowns ended.
- [ ] A demotion learned before a restart still applies after it.
- [ ] A backend with a long clean record moves up exactly one place past an unreliable neighbor and never further.
- [ ] Routing sends no request that the user's traffic did not cause.
- [ ] A model that appears in Pi's runtime catalog for google, groq, mistral, or nvidia and matches that backend's free rule appears under `grts` without a gratis update.
- [ ] A Z.ai model newly listed as Free on the pricing page and at cost 0 on models.dev appears; one that stops being listed as Free disappears; an unreachable page keeps the last confirmed list.
- [ ] After the user cools the last-answering model, the next `grts/auto` request skips it, including in the last-resort pass, until the cooldown ends or the user lifts it; a pinned request to it still answers.
- [ ] A context-overflow response during an `grts/auto` request advances the cascade; it never triggers cooling or fails while a next backend remains.
- [ ] A pinned model that 429s returns the upstream error, not a cascade.
- [ ] A pinned Kilo model returning upstream-saturation 429 while a virtual router model is available completes on that model without cooling the kilo backend.
- [ ] Virtual models advertise published catalog parameters when present.
- [ ] Catalog entries with non-openai-completions API types (anthropic-messages, google-shaped) stream correctly without conversion assumptions.
- [ ] OpenRouter discovery with zero `:free` variants and Kilo discovery with zero `:free` models each yield no models for that backend and no error.
- [ ] All registered models report input/output cost 0 and a free-tier data caveat in their description.
- [ ] With only `NVIDIA_API_KEY` set, NVIDIA models appear only for curated ids present in the live models list; with the list unreachable, the curated list appears.
- [ ] With only `ZAI_API_KEY` set, exactly the free Flash models appear, and requests go to the general endpoint, never a Coding Plan endpoint.
- [ ] With only `HETZNER_INFERENCE_API_KEY` set, Hetzner models follow the keyed live models list; with the list unreachable, the static list appears.
- [ ] Key-gated live tests run each backend's real endpoint only when that backend's key is set, skip otherwise, and stay outside `//:test`.