Groq · checked October 10, 2026
Groq "Rate limit reached for model ... on tokens per minute (TPM)"
Groq's 429 "Rate limit reached for model" names the limit you hit: requests or tokens, per minute or per day, for one model in your organization. On the free plan the token limits bite first: 8,000 tokens per minute and 200,000 per day on gpt-oss and Qwen (checked October 10, 2026). The message ends with the exact wait.
The exact error
What you see.
HTTP 429 for a rate limit, HTTP 413 when a single request is larger than the per-minute token limit. Both carry code rate_limit_exceeded.
{"error":{"message":"Rate limit reached for model `openai/gpt-oss-120b` in organization `org_...` service tier `on_demand` on tokens per minute (TPM): Limit 8000, Used 5912, Requested 4676. Please try again in 19.41s. Need more tokens? Upgrade to Dev Tier today at https://console.groq.com/settings/billing","type":"tokens","code":"rate_limit_exceeded"}}{"error":{"message":"Request too large for model `openai/gpt-oss-20b` in organization `org_...` service tier `on_demand` on tokens per minute (TPM): Limit 8000, Requested 8039, please reduce your message size and try again. Need more tokens? Upgrade to Dev Tier today at https://console.groq.com/settings/billing","type":"tokens","code":"rate_limit_exceeded"}}Causes
Why it happens.
- 1
Tokens per minute
8,000 TPM goes quickly with long prompts, chat history and tool schemas, all of which count.
- 2
Tokens per day
200,000 per model per day; the message says "tokens per day (TPD)" and the wait can be twenty minutes or more.
- 3
413: one request is bigger than the TPM limit
"Request too large" means the request alone exceeds the per-minute limit, so retrying never works. It is not a context-window error.
- 4
Limits are per organization
Groq applies limits at the organization level, so several keys in one organization share them.
- 5
A retired model name
llama-3.3-70b-versatile and llama-3.1-8b-instant now return 404 model_not_found, which old examples still use.
Fixes
How to fix it.
429: wait the time Groq gives
Use the retry-after header (seconds, sent on 429) or the "try again in" value, then retry. Watch x-ratelimit-remaining-tokens and x-ratelimit-remaining-requests to slow down before you hit it.
413: make the request smaller
Trim history, shorten tool schemas or lower max_completion_tokens, or send the request to a model or provider with a higher limit.
Know the free limits
Free plan, checked October 10, 2026 (console.groq.com/docs/rate-limits).
model RPM RPD TPM TPD
openai/gpt-oss-120b 30 1K 8K 200K
openai/gpt-oss-20b 30 1K 8K 200K
qwen/qwen3.8-27b 30 1K 8K 200KWith freelm
Handle it automatically.
Groq limits are per model, so freelm sets aside only the model that hit its limit, for the wait Groq sends, and keeps the key's other models and your other providers in play. A 413 "Request too large" is treated as "this model cannot take this request": freelm moves to a model or provider with more room instead of retrying. freelm also paces each Groq key to 30 requests a minute.
pip install freelm # or: npm install freelm
export GROQ_API_KEY=... # plus any other free keys you have
freelm doctor # one tiny live request per keyfreelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.
Questions
Related questions.
Are Groq rate limits per API key?
No. Groq applies them per organization, so extra keys in the same organization share the limits.
Why does Groq return 413 for my prompt?
Every token in the request counts toward the per-minute limit, including history and tool definitions. If the total is over 8,000 on the free plan, the request is rejected before it runs.
Sources