--- id: PX-RESEARCH-KMBSJUZR type: research title: Free LLM Inference Providers --- ## Free LLM Inference Providers Free-tier LLM inference providers that are established (multi-year track record), their access requirements, and estimated free usage. Numbers change frequently; each row cites its source and all sources were checked at research time. ### Landscape | Provider | Since | Access requirements | Free usage estimate | Caveats | | --- | --- | --- | --- | --- | | [OpenRouter](https://openrouter.ai/docs/faq) | 2023 | Account + API key, no card | `:free` models: 50 req/day, 20 req/min; with ≥\$10 lifetime credit purchase: 1000 req/day | Free endpoints may log prompts; 429s common; lowest tier caps prompt tokens per request | | [OpenCode Zen](https://opencode.ai/docs/zen) | 2025 | OpenCode account + API key, no card | Rotating free stealth/trial models (Big Pickle, MiMo-V2.5 Free, Nemotron Free, Ling Free); no published rate limit | Free = "limited time"; data may be collected to improve model; NVIDIA-served ones are trial-only | | [Google AI Studio / Gemini API](https://ai.google.dev/gemini-api/docs/rate-limits) | 2023 | Google account, no card | Per model, viewed in AI Studio; Flash class ≈ 10-15 RPM, 250K-1M TPM, 250-1500 RPD; RPD resets midnight PT | No published universal table anymore; free tier data may be used for product improvement | | [Groq](https://console.groq.com/docs/rate-limits) | API free tier 2024 | Account, no card | Per model; e.g. gpt-oss / qwen3-27b: 30 RPM, 1K RPD, 8K TPM, 200K TPD; audio + guard models higher RPD | Curated open models only; limits are org-level; exact values on limits page | | [Cerebras](https://inference-docs.cerebras.ai/) | 2024 | Account, no card | Free trial: 5 RPM, 30K TPM, 1M tokens/day (gpt-oss-120b, GLM-4.7) | 5 RPM breaks parallel agent tool calls; eval-only | | [Mistral La Plateforme](https://help.mistral.ai/en/articles/698531-why-am-i-hitting-api-rate-limits-and-how-do-i-increase-them) | Free tier 2024 | Account + phone (SMS) verification, no card | "Experiment" plan: ≈1 RPS, 500K TPM, ≈1B tokens/month, all models incl. Codestral | 1 RPS slow for concurrency; free data may train unless opted out | | [GitHub Models](https://docs.github.com/en/copilot/concepts/billing/copilot-requests) | 2024 | GitHub account + PAT | ≈10 RPM, ≈50 RPD on free Copilot tier; higher with paid Copilot plans | Hard 8K input / 4K output token cap per request; prototyping only | | [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/platform/pricing/) | GA 2024/25 | Cloudflare account | 10,000 Neurons/day ≈ 100-300 LLM requests/day; text generation ≈300 RPM | Frontier models (kimi, glm, deepseek) need paid plan or prepaid AI Gateway credits | | [NVIDIA NIM](https://build.nvidia.com/models) | 2024 | NVIDIA account | Free prototype endpoints, ≈40 RPM most models, no per-token billing | Prototype/trial use; older credit system (1000 credits signup) replaced by RPM limits | | [Cohere](https://docs.cohere.com/docs/rate-limits) | Trial keys since ~2023 | Account, no card | Trial key: 20 req/min chat, 1000 API calls/month total | Embed 2000 inputs/min; production key requires billing | | [Z.ai / Zhipu](https://docs.z.ai/guides/llm/glm-4.5) | GLM-4.5 free tier 2025 | Account, no card | GLM-4.5-Flash and GLM-4.7-Flash: \$0/token for registered users, 128K-200K ctx | Rate limits unpublished; concurrency caps by account tier | | [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers/pricing) | 2024/25 | HF account, no card | \$0.10 credits/month free users; \$2/month with PRO (\$9/mo) | After credits: must purchase credits; smallest allowance in this list | | [Vercel AI Gateway](https://vercel.com/docs/ai-gateway/pricing) | 2025 | Vercel team account | \$5 credits per 30 days for non-paying teams; free-tier model subset, lower per-model rate limits | Buying credits ends monthly free credit; BYOK needs paid tier | Discontinued free access worth knowing: [Chutes](https://chutes.ai/pricing) removed its \$5-deposit 200 req/day free tier in Aug 2025 → paid subscriptions only. [NanoGPT](https://nano-gpt.com/get-started) offers one anonymous free model only; paid access from \$0.10 crypto / \$1 card deposit. ### Reading the estimates Free request budgets per day are the headline number, but token budgets decide usability for coding agents: - OpenRouter 50 req/day ≈ 1-3M tokens/day at agent-typical 20-50K tokens/request → one short coding session. - Gemini free Flash ≈ 250-1500 RPD with 250K-1M TPM → the largest generous allowance; viable for real daily agent use. - Groq 200K-500K TPD on main models → small contexts only. - Mistral 1B tokens/month ≈ 33M/day → largest total volume, throttled to 1 RPS. - Cerebras 1M tokens/day but 5 RPM → single-threaded experiments. - GitHub Models 50 RPD × 8K in/4K out → smoke tests, not agents. ### Conclusions 1. Best sustainable free coding-agent budget: Google AI Studio (Flash class), then Mistral (volume, 1 RPS) and OpenRouter with the one-time \$10 top-up (1000 free req/day, permanent). 2. Cheapest "free forever" unlock: OpenRouter \$10 one-time purchase multiplies free quota 20x; no subscription. 3. All free tiers can change or vanish (Chutes precedent); none is contractual capacity. 4. Privacy: free tiers routinely log or train on prompts (Google, Mistral, OpenRouter free endpoints, Zen free models); avoid proprietary code on them. ### Unresolved questions - OpenCode Zen free-model rate limits: undocumented; only observable by hitting 429s. - Google per-model free limits: no longer published as a static table; must be read per project in AI Studio. - Z.ai free-model concurrency caps: unpublished, vary by account tier.