Luigit
repositories / pi-ext

pi-ext

bugabingas pi extensions

owned by admin

.system/research/PX-RESEARCH-KMBSJUZR-free-llm-inference-providers/index.md

Raw
Rendered preview

id: PX-RESEARCH-KMBSJUZR type: research title: Free LLM Inference Providers

Free LLM Inference Providers

Free-tier LLM inference providers that are established (multi-year track record), their access requirements, and estimated free usage. Numbers change frequently; each row cites its source and all sources were checked at research time.

Landscape

Provider Since Access requirements Free usage estimate Caveats
OpenRouter 2023 Account + API key, no card :free models: 50 req/day, 20 req/min; with ≥$10 lifetime credit purchase: 1000 req/day Free endpoints may log prompts; 429s common; lowest tier caps prompt tokens per request
OpenCode Zen 2025 OpenCode account + API key, no card Rotating free stealth/trial models (Big Pickle, MiMo-V2.5 Free, Nemotron Free, Ling Free); no published rate limit Free = "limited time"; data may be collected to improve model; NVIDIA-served ones are trial-only
Google AI Studio / Gemini API 2023 Google account, no card Per model, viewed in AI Studio; Flash class ≈ 10-15 RPM, 250K-1M TPM, 250-1500 RPD; RPD resets midnight PT No published universal table anymore; free tier data may be used for product improvement
Groq API free tier 2024 Account, no card Per model; e.g. gpt-oss / qwen3-27b: 30 RPM, 1K RPD, 8K TPM, 200K TPD; audio + guard models higher RPD Curated open models only; limits are org-level; exact values on limits page
Cerebras 2024 Account, no card Free trial: 5 RPM, 30K TPM, 1M tokens/day (gpt-oss-120b, GLM-4.7) 5 RPM breaks parallel agent tool calls; eval-only
Mistral La Plateforme Free tier 2024 Account + phone (SMS) verification, no card "Experiment" plan: ≈1 RPS, 500K TPM, ≈1B tokens/month, all models incl. Codestral 1 RPS slow for concurrency; free data may train unless opted out
GitHub Models 2024 GitHub account + PAT ≈10 RPM, ≈50 RPD on free Copilot tier; higher with paid Copilot plans Hard 8K input / 4K output token cap per request; prototyping only
Cloudflare Workers AI GA 2024/25 Cloudflare account 10,000 Neurons/day ≈ 100-300 LLM requests/day; text generation ≈300 RPM Frontier models (kimi, glm, deepseek) need paid plan or prepaid AI Gateway credits
NVIDIA NIM 2024 NVIDIA account Free prototype endpoints, ≈40 RPM most models, no per-token billing Prototype/trial use; older credit system (1000 credits signup) replaced by RPM limits
Cohere Trial keys since ~2023 Account, no card Trial key: 20 req/min chat, 1000 API calls/month total Embed 2000 inputs/min; production key requires billing
Z.ai / Zhipu GLM-4.5 free tier 2025 Account, no card GLM-4.5-Flash and GLM-4.7-Flash: $0/token for registered users, 128K-200K ctx Rate limits unpublished; concurrency caps by account tier
Hugging Face Inference Providers 2024/25 HF account, no card $0.10 credits/month free users; $2/month with PRO ($9/mo) After credits: must purchase credits; smallest allowance in this list
Vercel AI Gateway 2025 Vercel team account $5 credits per 30 days for non-paying teams; free-tier model subset, lower per-model rate limits Buying credits ends monthly free credit; BYOK needs paid tier

Discontinued free access worth knowing: Chutes removed its $5-deposit 200 req/day free tier in Aug 2025 → paid subscriptions only. NanoGPT offers one anonymous free model only; paid access from $0.10 crypto / $1 card deposit.

Reading the estimates

Free request budgets per day are the headline number, but token budgets decide usability for coding agents:

  • OpenRouter 50 req/day ≈ 1-3M tokens/day at agent-typical 20-50K tokens/request → one short coding session.
  • Gemini free Flash ≈ 250-1500 RPD with 250K-1M TPM → the largest generous allowance; viable for real daily agent use.
  • Groq 200K-500K TPD on main models → small contexts only.
  • Mistral 1B tokens/month ≈ 33M/day → largest total volume, throttled to 1 RPS.
  • Cerebras 1M tokens/day but 5 RPM → single-threaded experiments.
  • GitHub Models 50 RPD × 8K in/4K out → smoke tests, not agents.

Conclusions

  1. Best sustainable free coding-agent budget: Google AI Studio (Flash class), then Mistral (volume, 1 RPS) and OpenRouter with the one-time $10 top-up (1000 free req/day, permanent).
  2. Cheapest "free forever" unlock: OpenRouter $10 one-time purchase multiplies free quota 20x; no subscription.
  3. All free tiers can change or vanish (Chutes precedent); none is contractual capacity.
  4. Privacy: free tiers routinely log or train on prompts (Google, Mistral, OpenRouter free endpoints, Zen free models); avoid proprietary code on them.

Unresolved questions

  • OpenCode Zen free-model rate limits: undocumented; only observable by hitting 429s.
  • Google per-model free limits: no longer published as a static table; must be read per project in AI Studio.
  • Z.ai free-model concurrency caps: unpublished, vary by account tier.
---
id: PX-RESEARCH-KMBSJUZR
type: research
title: Free LLM Inference Providers
---

## Free LLM Inference Providers

Free-tier LLM inference providers that are established (multi-year track record), their access requirements, and estimated free usage.
Numbers change frequently; each row cites its source and all sources were checked at research time.

### Landscape

| Provider | Since | Access requirements | Free usage estimate | Caveats |
| --- | --- | --- | --- | --- |
| [OpenRouter](https://openrouter.ai/docs/faq) | 2023 | Account + API key, no card | `:free` models: 50 req/day, 20 req/min; with ≥\$10 lifetime credit purchase: 1000 req/day | Free endpoints may log prompts; 429s common; lowest tier caps prompt tokens per request |
| [OpenCode Zen](https://opencode.ai/docs/zen) | 2025 | OpenCode account + API key, no card | Rotating free stealth/trial models (Big Pickle, MiMo-V2.5 Free, Nemotron Free, Ling Free); no published rate limit | Free = "limited time"; data may be collected to improve model; NVIDIA-served ones are trial-only |
| [Google AI Studio / Gemini API](https://ai.google.dev/gemini-api/docs/rate-limits) | 2023 | Google account, no card | Per model, viewed in AI Studio; Flash class ≈ 10-15 RPM, 250K-1M TPM, 250-1500 RPD; RPD resets midnight PT | No published universal table anymore; free tier data may be used for product improvement |
| [Groq](https://console.groq.com/docs/rate-limits) | API free tier 2024 | Account, no card | Per model; e.g. gpt-oss / qwen3-27b: 30 RPM, 1K RPD, 8K TPM, 200K TPD; audio + guard models higher RPD | Curated open models only; limits are org-level; exact values on limits page |
| [Cerebras](https://inference-docs.cerebras.ai/) | 2024 | Account, no card | Free trial: 5 RPM, 30K TPM, 1M tokens/day (gpt-oss-120b, GLM-4.7) | 5 RPM breaks parallel agent tool calls; eval-only |
| [Mistral La Plateforme](https://help.mistral.ai/en/articles/698531-why-am-i-hitting-api-rate-limits-and-how-do-i-increase-them) | Free tier 2024 | Account + phone (SMS) verification, no card | "Experiment" plan: ≈1 RPS, 500K TPM, ≈1B tokens/month, all models incl. Codestral | 1 RPS slow for concurrency; free data may train unless opted out |
| [GitHub Models](https://docs.github.com/en/copilot/concepts/billing/copilot-requests) | 2024 | GitHub account + PAT | ≈10 RPM, ≈50 RPD on free Copilot tier; higher with paid Copilot plans | Hard 8K input / 4K output token cap per request; prototyping only |
| [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/platform/pricing/) | GA 2024/25 | Cloudflare account | 10,000 Neurons/day ≈ 100-300 LLM requests/day; text generation ≈300 RPM | Frontier models (kimi, glm, deepseek) need paid plan or prepaid AI Gateway credits |
| [NVIDIA NIM](https://build.nvidia.com/models) | 2024 | NVIDIA account | Free prototype endpoints, ≈40 RPM most models, no per-token billing | Prototype/trial use; older credit system (1000 credits signup) replaced by RPM limits |
| [Cohere](https://docs.cohere.com/docs/rate-limits) | Trial keys since ~2023 | Account, no card | Trial key: 20 req/min chat, 1000 API calls/month total | Embed 2000 inputs/min; production key requires billing |
| [Z.ai / Zhipu](https://docs.z.ai/guides/llm/glm-4.5) | GLM-4.5 free tier 2025 | Account, no card | GLM-4.5-Flash and GLM-4.7-Flash: \$0/token for registered users, 128K-200K ctx | Rate limits unpublished; concurrency caps by account tier |
| [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers/pricing) | 2024/25 | HF account, no card | \$0.10 credits/month free users; \$2/month with PRO (\$9/mo) | After credits: must purchase credits; smallest allowance in this list |
| [Vercel AI Gateway](https://vercel.com/docs/ai-gateway/pricing) | 2025 | Vercel team account | \$5 credits per 30 days for non-paying teams; free-tier model subset, lower per-model rate limits | Buying credits ends monthly free credit; BYOK needs paid tier |

Discontinued free access worth knowing: [Chutes](https://chutes.ai/pricing) removed its \$5-deposit 200 req/day free tier in Aug 2025 → paid subscriptions only.
[NanoGPT](https://nano-gpt.com/get-started) offers one anonymous free model only; paid access from \$0.10 crypto / \$1 card deposit.

### Reading the estimates

Free request budgets per day are the headline number, but token budgets decide usability for coding agents:

- OpenRouter 50 req/day ≈ 1-3M tokens/day at agent-typical 20-50K tokens/request → one short coding session.
- Gemini free Flash ≈ 250-1500 RPD with 250K-1M TPM → the largest generous allowance; viable for real daily agent use.
- Groq 200K-500K TPD on main models → small contexts only.
- Mistral 1B tokens/month ≈ 33M/day → largest total volume, throttled to 1 RPS.
- Cerebras 1M tokens/day but 5 RPM → single-threaded experiments.
- GitHub Models 50 RPD × 8K in/4K out → smoke tests, not agents.

### Conclusions

1. Best sustainable free coding-agent budget: Google AI Studio (Flash class), then Mistral (volume, 1 RPS) and OpenRouter with the one-time \$10 top-up (1000 free req/day, permanent).
2. Cheapest "free forever" unlock: OpenRouter \$10 one-time purchase multiplies free quota 20x; no subscription.
3. All free tiers can change or vanish (Chutes precedent); none is contractual capacity.
4. Privacy: free tiers routinely log or train on prompts (Google, Mistral, OpenRouter free endpoints, Zen free models); avoid proprietary code on them.

### Unresolved questions

- OpenCode Zen free-model rate limits: undocumented; only observable by hitting 429s.
- Google per-model free limits: no longer published as a static table; must be read per project in AI Studio.
- Z.ai free-model concurrency caps: unpublished, vary by account tier.