pi-ext bugabingas pi extensions
owned by admin
README Source History Refs Compare Notes Search .system/research/PX-RESEARCH-KMBSJUZR-free-llm-inference-providers/index.md Raw Rendered preview
id: PX-RESEARCH-KMBSJUZR
type: research
title: Free LLM Inference Providers
Free LLM Inference Providers
Free-tier LLM inference providers that are established (multi-year track record), their access requirements, and estimated free usage.
Numbers change frequently; each row cites its source and all sources were checked at research time.
Landscape
Provider
Since
Access requirements
Free usage estimate
Caveats
OpenRouter
2023
Account + API key, no card
:free models: 50 req/day, 20 req/min; with ≥$10 lifetime credit purchase: 1000 req/day
Free endpoints may log prompts; 429s common; lowest tier caps prompt tokens per request
OpenCode Zen
2025
OpenCode account + API key, no card
Rotating free stealth/trial models (Big Pickle, MiMo-V2.5 Free, Nemotron Free, Ling Free); no published rate limit
Free = "limited time"; data may be collected to improve model; NVIDIA-served ones are trial-only
Google AI Studio / Gemini API
2023
Google account, no card
Per model, viewed in AI Studio; Flash class ≈ 10-15 RPM, 250K-1M TPM, 250-1500 RPD; RPD resets midnight PT
No published universal table anymore; free tier data may be used for product improvement
Groq
API free tier 2024
Account, no card
Per model; e.g. gpt-oss / qwen3-27b: 30 RPM, 1K RPD, 8K TPM, 200K TPD; audio + guard models higher RPD
Curated open models only; limits are org-level; exact values on limits page
Cerebras
2024
Account, no card
Free trial: 5 RPM, 30K TPM, 1M tokens/day (gpt-oss-120b, GLM-4.7)
5 RPM breaks parallel agent tool calls; eval-only
Mistral La Plateforme
Free tier 2024
Account + phone (SMS) verification, no card
"Experiment" plan: ≈1 RPS, 500K TPM, ≈1B tokens/month, all models incl. Codestral
1 RPS slow for concurrency; free data may train unless opted out
GitHub Models
2024
GitHub account + PAT
≈10 RPM, ≈50 RPD on free Copilot tier; higher with paid Copilot plans
Hard 8K input / 4K output token cap per request; prototyping only
Cloudflare Workers AI
GA 2024/25
Cloudflare account
10,000 Neurons/day ≈ 100-300 LLM requests/day; text generation ≈300 RPM
Frontier models (kimi, glm, deepseek) need paid plan or prepaid AI Gateway credits
NVIDIA NIM
2024
NVIDIA account
Free prototype endpoints, ≈40 RPM most models, no per-token billing
Prototype/trial use; older credit system (1000 credits signup) replaced by RPM limits
Cohere
Trial keys since ~2023
Account, no card
Trial key: 20 req/min chat, 1000 API calls/month total
Embed 2000 inputs/min; production key requires billing
Z.ai / Zhipu
GLM-4.5 free tier 2025
Account, no card
GLM-4.5-Flash and GLM-4.7-Flash: $0/token for registered users, 128K-200K ctx
Rate limits unpublished; concurrency caps by account tier
Hugging Face Inference Providers
2024/25
HF account, no card
$0.10 credits/month free users; $2/month with PRO ($9/mo)
After credits: must purchase credits; smallest allowance in this list
Vercel AI Gateway
2025
Vercel team account
$5 credits per 30 days for non-paying teams; free-tier model subset, lower per-model rate limits
Buying credits ends monthly free credit; BYOK needs paid tier
Discontinued free access worth knowing: Chutes removed its $5-deposit 200 req/day free tier in Aug 2025 → paid subscriptions only.
NanoGPT offers one anonymous free model only; paid access from $0.10 crypto / $1 card deposit.
Reading the estimates
Free request budgets per day are the headline number, but token budgets decide usability for coding agents:
OpenRouter 50 req/day ≈ 1-3M tokens/day at agent-typical 20-50K tokens/request → one short coding session.
Gemini free Flash ≈ 250-1500 RPD with 250K-1M TPM → the largest generous allowance; viable for real daily agent use.
Groq 200K-500K TPD on main models → small contexts only.
Mistral 1B tokens/month ≈ 33M/day → largest total volume, throttled to 1 RPS.
Cerebras 1M tokens/day but 5 RPM → single-threaded experiments.
GitHub Models 50 RPD × 8K in/4K out → smoke tests, not agents.
Conclusions
Best sustainable free coding-agent budget: Google AI Studio (Flash class), then Mistral (volume, 1 RPS) and OpenRouter with the one-time $10 top-up (1000 free req/day, permanent).
Cheapest "free forever" unlock: OpenRouter $10 one-time purchase multiplies free quota 20x; no subscription.
All free tiers can change or vanish (Chutes precedent); none is contractual capacity.
Privacy: free tiers routinely log or train on prompts (Google, Mistral, OpenRouter free endpoints, Zen free models); avoid proprietary code on them.
Unresolved questions
OpenCode Zen free-model rate limits: undocumented; only observable by hitting 429s.
Google per-model free limits: no longer published as a static table; must be read per project in AI Studio.
Z.ai free-model concurrency caps: unpublished, vary by account tier.
---
id: PX-RESEARCH-KMBSJUZR
type: research
title: Free LLM Inference Providers
---
## Free LLM Inference Providers
Free-tier LLM inference providers that are established (multi-year track record), their access requirements, and estimated free usage.
Numbers change frequently; each row cites its source and all sources were checked at research time.
### Landscape
| Provider | Since | Access requirements | Free usage estimate | Caveats |
| --- | --- | --- | --- | --- |
| [OpenRouter](https://openrouter.ai/docs/faq) | 2023 | Account + API key, no card | `:free` models: 50 req/day, 20 req/min; with ≥\$10 lifetime credit purchase: 1000 req/day | Free endpoints may log prompts; 429s common; lowest tier caps prompt tokens per request |
| [OpenCode Zen](https://opencode.ai/docs/zen) | 2025 | OpenCode account + API key, no card | Rotating free stealth/trial models (Big Pickle, MiMo-V2.5 Free, Nemotron Free, Ling Free); no published rate limit | Free = "limited time"; data may be collected to improve model; NVIDIA-served ones are trial-only |
| [Google AI Studio / Gemini API](https://ai.google.dev/gemini-api/docs/rate-limits) | 2023 | Google account, no card | Per model, viewed in AI Studio; Flash class ≈ 10-15 RPM, 250K-1M TPM, 250-1500 RPD; RPD resets midnight PT | No published universal table anymore; free tier data may be used for product improvement |
| [Groq](https://console.groq.com/docs/rate-limits) | API free tier 2024 | Account, no card | Per model; e.g. gpt-oss / qwen3-27b: 30 RPM, 1K RPD, 8K TPM, 200K TPD; audio + guard models higher RPD | Curated open models only; limits are org-level; exact values on limits page |
| [Cerebras](https://inference-docs.cerebras.ai/) | 2024 | Account, no card | Free trial: 5 RPM, 30K TPM, 1M tokens/day (gpt-oss-120b, GLM-4.7) | 5 RPM breaks parallel agent tool calls; eval-only |
| [Mistral La Plateforme](https://help.mistral.ai/en/articles/698531-why-am-i-hitting-api-rate-limits-and-how-do-i-increase-them) | Free tier 2024 | Account + phone (SMS) verification, no card | "Experiment" plan: ≈1 RPS, 500K TPM, ≈1B tokens/month, all models incl. Codestral | 1 RPS slow for concurrency; free data may train unless opted out |
| [GitHub Models](https://docs.github.com/en/copilot/concepts/billing/copilot-requests) | 2024 | GitHub account + PAT | ≈10 RPM, ≈50 RPD on free Copilot tier; higher with paid Copilot plans | Hard 8K input / 4K output token cap per request; prototyping only |
| [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/platform/pricing/) | GA 2024/25 | Cloudflare account | 10,000 Neurons/day ≈ 100-300 LLM requests/day; text generation ≈300 RPM | Frontier models (kimi, glm, deepseek) need paid plan or prepaid AI Gateway credits |
| [NVIDIA NIM](https://build.nvidia.com/models) | 2024 | NVIDIA account | Free prototype endpoints, ≈40 RPM most models, no per-token billing | Prototype/trial use; older credit system (1000 credits signup) replaced by RPM limits |
| [Cohere](https://docs.cohere.com/docs/rate-limits) | Trial keys since ~2023 | Account, no card | Trial key: 20 req/min chat, 1000 API calls/month total | Embed 2000 inputs/min; production key requires billing |
| [Z.ai / Zhipu](https://docs.z.ai/guides/llm/glm-4.5) | GLM-4.5 free tier 2025 | Account, no card | GLM-4.5-Flash and GLM-4.7-Flash: \$0/token for registered users, 128K-200K ctx | Rate limits unpublished; concurrency caps by account tier |
| [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers/pricing) | 2024/25 | HF account, no card | \$0.10 credits/month free users; \$2/month with PRO (\$9/mo) | After credits: must purchase credits; smallest allowance in this list |
| [Vercel AI Gateway](https://vercel.com/docs/ai-gateway/pricing) | 2025 | Vercel team account | \$5 credits per 30 days for non-paying teams; free-tier model subset, lower per-model rate limits | Buying credits ends monthly free credit; BYOK needs paid tier |
Discontinued free access worth knowing: [Chutes](https://chutes.ai/pricing) removed its \$5-deposit 200 req/day free tier in Aug 2025 → paid subscriptions only.
[NanoGPT](https://nano-gpt.com/get-started) offers one anonymous free model only; paid access from \$0.10 crypto / \$1 card deposit.
### Reading the estimates
Free request budgets per day are the headline number, but token budgets decide usability for coding agents:
- OpenRouter 50 req/day ≈ 1-3M tokens/day at agent-typical 20-50K tokens/request → one short coding session.
- Gemini free Flash ≈ 250-1500 RPD with 250K-1M TPM → the largest generous allowance; viable for real daily agent use.
- Groq 200K-500K TPD on main models → small contexts only.
- Mistral 1B tokens/month ≈ 33M/day → largest total volume, throttled to 1 RPS.
- Cerebras 1M tokens/day but 5 RPM → single-threaded experiments.
- GitHub Models 50 RPD × 8K in/4K out → smoke tests, not agents.
### Conclusions
1. Best sustainable free coding-agent budget: Google AI Studio (Flash class), then Mistral (volume, 1 RPS) and OpenRouter with the one-time \$10 top-up (1000 free req/day, permanent).
2. Cheapest "free forever" unlock: OpenRouter \$10 one-time purchase multiplies free quota 20x; no subscription.
3. All free tiers can change or vanish (Chutes precedent); none is contractual capacity.
4. Privacy: free tiers routinely log or train on prompts (Google, Mistral, OpenRouter free endpoints, Zen free models); avoid proprietary code on them.
### Unresolved questions
- OpenCode Zen free-model rate limits: undocumented; only observable by hitting 429s.
- Google per-model free limits: no longer published as a static table; must be read per project in AI Studio.
- Z.ai free-model concurrency caps: unpublished, vary by account tier.