Free LLM API resources
This lists various services that provide free access or credits towards API-based LLM usage.
Note
Please don't abuse these services, else we might lose them.
Warning
This list explicitly excludes any services that are not legitimate (eg reverse engineers an existing chatbot)
Free Providers
Limits:
20 requests/minute
50 requests/day
Up to 1000 requests/day with $10 lifetime topup
Models share a common quota.
- Cohere North Mini Code
- Ling 3.0 Flash
- NVIDIA Nemotron 3 Nano Omni 30B A3B (Reasoning)
- NVIDIA Nemotron 3 Super 120B A12B
- NVIDIA Nemotron 3 Ultra 550B A55B
- NVIDIA Nemotron 3.5 Content Safety
- Poolside Laguna M.1
- Poolside Laguna S 2.1
- Poolside Laguna XS 2.1
- google/gemma-4-26b-a4b-it:free
- google/gemma-4-31b-it:free
- nvidia/nemotron-3-nano-30b-a3b:free
- nvidia/nemotron-nano-12b-v2-vl:free
- nvidia/nemotron-nano-9b-v2:free
- openai/gpt-oss-20b:free
Data is used for training when used outside of the UK/CH/EEA/EU.
| Model Name | Model Limits |
|---|---|
| Gemini 3.6 Flash | 250,000 tokens/minute 20 requests/day 5 requests/minute |
| Gemini 3.5 Flash | 250,000 tokens/minute 20 requests/day 5 requests/minute |
| Gemini 3 Flash | 250,000 tokens/minute 20 requests/day 5 requests/minute |
| Gemini 3.5 Flash-Lite | 250,000 tokens/minute 500 requests/day 15 requests/minute |
| Gemini 3.1 Flash-Lite | 250,000 tokens/minute 500 requests/day 15 requests/minute |
| Gemini 2.5 Flash | 250,000 tokens/minute 20 requests/day 5 requests/minute |
| Gemini 2.5 Flash-Lite | 250,000 tokens/minute 20 requests/day 10 requests/minute |
| Gemini 3.1 Flash TTS | 10,000 tokens/minute 10 requests/day 3 requests/minute |
| Gemini 2.5 Flash TTS | 10,000 tokens/minute 10 requests/day 3 requests/minute |
| Gemini Robotics-ER 1.6 | 250,000 tokens/minute 20 requests/day 5 requests/minute |
| Gemini Robotics-ER 1.5 | 250,000 tokens/minute 20 requests/day 10 requests/minute |
| Gemma 4 31B Instruct | 16,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 4 26B A4B Instruct | 16,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 3 27B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 3 12B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 3 4B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 3 1B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
Phone number verification required. Models tend to be context window limited.
Limits: 40 requests/minute
- Free tier (Experiment plan) requires opting into data training
- Requires phone number verification.
Limits: Set per-model and per-organization — check your limits page. As of July 2026 a new free account sees anywhere from 25,000 to 20,000,000 tokens/minute and 0.03 to 12.5 requests/second depending on the model.
- Currently free to use
- Monthly subscription based
- Requires phone number verification
Limits: 30 requests/minute, 2,000 requests/day
- Codestral
HuggingFace Serverless Inference limited to models smaller than 10GB. Some popular models are supported even if they exceed 10GB.
Limits: $0.10/month in credits
- Various open models across supported providers
Routes to various supported providers.
The free tier covers a subset of the model catalogue, with per-model rate limits.
Limits: $5/month
OpenAI-compatible gateway routing to various providers. Free models work without an account.
All free models may use your prompts for training.
Limits: 200 requests/hour per IP, shared across all free models
- Cohere North Mini Code
- Kilo Auto Free (Router)
- Kwaipilot KAT-Coder-Pro V2.5
- Ling 3.0 Flash
- NVIDIA Nemotron 3 Nano Omni 30B A3B (Reasoning)
- NVIDIA Nemotron 3 Super 120B A12B
- NVIDIA Nemotron 3 Ultra 550B A55B
- NVIDIA Nemotron 3.5 Content Safety
- OpenRouter Free Models (Router)
- Poolside Laguna M.1
- Poolside Laguna S 2.1
- Poolside Laguna XS 2.1
- StepFun Step 3.7 Flash
AI gateway with curated models.
Free models may use data for improvement.
- Big Pickle
- DeepSeek V4 Flash Free
- MiMo-V2.5 Free
- Laguna S 2.1 Free
- Ling-3.0-flash Free
- North Mini Code Free
- Nemotron 3 Ultra Free
| Model Name | Model Limits |
|---|---|
| gpt-oss-120b | 5 requests/minute 30,000 tokens/minute 1,000,000 tokens/hour 1,000,000 tokens/day |
| zai-glm-4.7 | 5 requests/minute 30,000 tokens/minute 1,000,000 tokens/hour 1,000,000 tokens/day |
| gemma-4-31b | 5 requests/minute 30,000 tokens/minute 1,000,000 tokens/hour 1,000,000 tokens/day |
| Model Name | Model Limits |
|---|---|
| Allam 2 7B | 7,000 requests/day 6,000 tokens/minute |
| Llama 3.1 8B | 14,400 requests/day 6,000 tokens/minute |
| Llama 3.3 70B | 1,000 requests/day 12,000 tokens/minute |
| Whisper Large v3 | 2,000 requests/day |
| Whisper Large v3 Turbo | 2,000 requests/day |
| canopylabs/orpheus-arabic-saudi | |
| canopylabs/orpheus-v1-english | |
| groq/compound | 250 requests/day 70,000 tokens/minute |
| groq/compound-mini | 250 requests/day 70,000 tokens/minute |
| meta-llama/llama-prompt-guard-2-22m | |
| meta-llama/llama-prompt-guard-2-86m | |
| openai/gpt-oss-120b | 1,000 requests/day 8,000 tokens/minute |
| openai/gpt-oss-20b | 1,000 requests/day 8,000 tokens/minute |
| openai/gpt-oss-safeguard-20b | 1,000 requests/day 8,000 tokens/minute |
| qwen/qwen3.6-27b | 1,000 requests/day 8,000 tokens/minute |
Limits:
20 requests/minute
1,000 requests/month
Models share a common monthly quota.
- c4ai-aya-expanse-32b
- c4ai-aya-vision-32b
- command-a-03-2025
- command-a-plus-05-2026
- command-a-reasoning-08-2025
- command-a-translate-08-2025
- command-a-vision-07-2025
- command-r-08-2024
- command-r-plus-08-2024
- command-r7b-12-2024
- command-r7b-arabic-02-2025
Limits: 10,000 neurons/day
- @cf/aisingapore/gemma-sea-lion-v4-27b-it
- @cf/google/gemma-4-26b-a4b-it
- @cf/ibm-granite/granite-4.0-h-micro
- @cf/moonshotai/kimi-k2.6
- @cf/moonshotai/kimi-k2.7-code
- @cf/nvidia/nemotron-3-120b-a12b
- @cf/openai/gpt-oss-120b
- @cf/openai/gpt-oss-20b
- @cf/qwen/qwen3-30b-a3b-fp8
- @cf/zai-org/glm-4.7-flash
- @cf/zai-org/glm-5.2
- DeepSeek R1 Distill Qwen 32B
- Gemma 2B Instruct (LoRA)
- Gemma 7B Instruct (LoRA)
- Llama 2 7B Chat (LoRA)
- Llama 3.1 8B Instruct (FP8)
- Llama 3.2 11B Vision Instruct
- Llama 3.2 1B Instruct
- Llama 3.2 3B Instruct
- Llama 3.3 70B Instruct (FP8)
- Llama 4 Scout Instruct
- Llama Guard 3 8B
- Mistral 7B Instruct v0.2 (LoRA)
- Mistral Small 3.1 24B Instruct
- Qwen 2.5 Coder 32B Instruct
- Qwen QwQ 32B
Providers with trial credits
Credits: $1
Models: Various open models
Credits: $30
Models: Any supported model - pay by compute time
Credits: $1
Models: Various open models
Credits: $0.5 for 1 year
Models: Various open models
Credits: $10 for 3 months
Models: Jamba family of models
Credits: $10 for 3 months
Models: Solar Pro/Mini
Credits: $15
Requirements: Phone number verification
Models: Various open models
Credits: 1 million tokens/model, valid for 90 days (Singapore endpoint only)
Models: Various open and proprietary Qwen models
Credits: $30/month on the Starter plan
Models: Any supported model - pay by compute time
Credits: $1, $25 on responding to email survey
Models: Various open models
Credits: $1
Models:
- DeepSeek V3 0324
- Llama 3.3 70B Instruct
- deepseek-ai/deepseek-r1-0528
- qwen/qwen3-coder-480b-a35b-instruct
Credits: $5 for 3 months
Models:
- deepseek-v3.1
- deepseek-v3.2
- gemma-4-31b-it
- gpt-oss-120b
- meta-llama-3.3-70b-instruct
- minimax-m2.7
Credits: 1,000,000 free tokens, plus 60 minutes of audio transcription
Models:
- BGE-Multilingual-Gemma2
- Gemma 3 27B Instruct
- Llama 3.3 70B Instruct
- Pixtral 12B (2409)
- Whisper Large v3
- devstral-2-123b-instruct-2512
- gemma-4-26b-a4b-it
- glm-5.2
- gpt-oss-120b
- holo2-30b-a3b
- mistral-medium-3.5-128b
- mistral-small-3.2-24b-instruct-2506
- qwen3-235b-a22b-instruct-2507
- qwen3-coder-30b-a3b-instruct
- qwen3-embedding-8b
- qwen3.5-397b-a17b
- qwen3.6-35b-a3b
- voxtral-small-24b-2507