← All free LLM API errors

Google Gemini API · checked October 10, 2026

Gemini 503 "The model is overloaded" or "experiencing high demand"

A 503 UNAVAILABLE from Gemini means the model you called is temporarily out of capacity. It is not a quota error and says nothing about your key: other Gemini models usually keep working. Responses from late 2026 read "This model is currently experiencing high demand"; older ones say "The model is overloaded. Please try again later."

The exact error

What you see.

HTTP 503. The first body is the wording reported in September and October 2026; the second is the older one, still seen on some models in mid-2026.

{"error": {"code": 503, "message": "This model is currently experiencing high demand. Spikes in demand are usually temporary. Please try again later.", "status": "UNAVAILABLE"}}
{"error": {"code": 503, "message": "The model is overloaded. Please try again later.", "status": "UNAVAILABLE"}}

Causes

Why it happens.

  1. 1

    Capacity on one model

    New and preview models, and the most popular Flash models at busy hours, run out of capacity first. Google's error reference describes it as the service being temporarily overloaded or down.

  2. 2

    It is not your account

    The same key keeps working on other models, and the error does not consume or reveal anything about your quota.

Fixes

How to fix it.

Back off, then switch model

Retry two or three times with exponential backoff and jitter, then fall back to another model such as a Flash-Lite model or a -latest alias.

import random, time

def call_with_fallback(call, models=("gemini-3.5-flash", "gemini-3.5-flash-lite", "gemini-flash-lite-latest")):
    for model in models:
        for attempt in range(3):
            status, body = call(model)
            if status != 503:
                return status, body
            time.sleep(2 ** attempt + random.random())
    return status, body

Do not disable the key

Treat a 503 as a temporary model problem. Logic that marks the key bad on any 5xx throws away a working key.

With freelm

Handle it automatically.

freelm treats an overloaded or high-demand 503 as a problem with that model only: it sets the model aside for up to 30 seconds without penalising the key, and the same call moves to the next Gemini model or provider. If a model is slow instead of failing, freelm races it against the next provider after a few seconds and keeps whichever answers first.

import freelm
llm = freelm.FreeLLM.from_env()
print(llm.text("hello", model="chat:fast"))   # any model or alias; failover is automatic

freelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.

Questions

Related questions.

Should I retry a Gemini 503?

Yes, a few times with backoff, then switch model. The official SDKs already retry 5xx errors automatically.

Does a 503 mean my key is blocked?

No. It is a capacity error on the model you called; the same key keeps working on other models.