Google Gemini API · checked October 10, 2026
Gemini 503 "The model is overloaded" or "experiencing high demand"
A 503 UNAVAILABLE from Gemini means the model you called is temporarily out of capacity. It is not a quota error and says nothing about your key: other Gemini models usually keep working. Responses from late 2026 read "This model is currently experiencing high demand"; older ones say "The model is overloaded. Please try again later."
The exact error
What you see.
HTTP 503. The first body is the wording reported in September and October 2026; the second is the older one, still seen on some models in mid-2026.
{"error": {"code": 503, "message": "This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.", "status": "UNAVAILABLE"}}{"error": {"code": 503, "message": "The model is overloaded. Please try again later.", "status": "UNAVAILABLE"}}Causes
Why it happens.
- 1
Capacity on one model
New and preview models, and the most popular Flash models at busy hours, run out of capacity first. Google's error reference describes it as the service being temporarily overloaded or down.
- 2
It is not your account
The same key keeps working on other models, and the error does not consume or reveal anything about your quota.
Fixes
How to fix it.
Back off, then switch model
Retry two or three times with exponential backoff and jitter, then fall back to another model such as a Flash-Lite model or a -latest alias.
import random, time
def call_with_fallback(call, models=("gemini-3.5-flash", "gemini-3.5-flash-lite", "gemini-flash-lite-latest")):
for model in models:
for attempt in range(3):
status, body = call(model)
if status != 503:
return status, body
time.sleep(2 ** attempt + random.random())
return status, bodyDo not disable the key
Treat a 503 as a temporary model problem. Logic that marks the key bad on any 5xx throws away a working key.
With freelm
Handle it automatically.
freelm treats an overloaded or high-demand 503 as a problem with that model only: it sets the model aside for up to 30 seconds without penalising the key, and the same call moves to the next Gemini model or provider. If a model is slow instead of failing, freelm races it against the next provider after a few seconds and keeps whichever answers first.
import freelm
llm = freelm.FreeLLM.from_env()
print(llm.text("hello", model="chat:fast")) # any model or alias; failover is automaticfreelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.
Questions
Related questions.
Should I retry a Gemini 503?
Yes, a few times with backoff, then switch model. The official SDKs already retry 5xx errors automatically.
Does a 503 mean my key is blocked?
No. It is a capacity error on the model you called; the same key keeps working on other models.
Sources