← All free LLM API errors

Google Gemini API · checked October 10, 2026

Gemini API 429 RESOURCE_EXHAUSTED: quota exceeded, and how to fix it

A 429 RESOURCE_EXHAUSTED from the Gemini API means one quota for one model ran out. On the free tier the usual cause is the daily cap, about 20 requests per model per project in October 2026, which resets at midnight Pacific time. The quotaId in the error tells a daily cap from a per-minute burst, and they need different fixes.

The exact error

What you see.

HTTP 429 from generateContent. The OpenAI-compatible endpoint (/v1beta/openai) returns the same object wrapped in a one-element JSON array, which some SDKs report as "429 status code (no body)". The model shown is from a July 2026 report.

{
  "error": {
    "code": 429,
    "message": "You exceeded your current quota, please check your plan and billing details. For more information on this error, head to: https://ai.google.dev/gemini-api/docs/rate-limits. ... Please retry in 29.019961092s.",
    "status": "RESOURCE_EXHAUSTED",
    "details": [
      { "@type": "type.googleapis.com/google.rpc.QuotaFailure",
        "violations": [
          { "quotaMetric": "generativelanguage.googleapis.com/generate_content_free_tier_requests",
            "quotaId": "GenerateRequestsPerDayPerProjectPerModel-FreeTier",
            "quotaDimensions": { "location": "global", "model": "gemini-2.0-flash" } }
        ] },
      { "@type": "type.googleapis.com/google.rpc.RetryInfo", "retryDelay": "29s" }
    ]
  }
}
openai.RateLimitError: Error code: 429 - [{'error': {'code': 429, 'message': 'Resource has been exhausted (e.g. check quota).', 'status': 'RESOURCE_EXHAUSTED'}}]

Causes

Why it happens.

  1. 1

    The free daily cap for that model is spent

    Free quotas are per project and per model. In October 2026 reports put the free Gemini 3.x Flash models at about 20 requests per day each. The quotaId contains PerDay, and the cap resets at midnight Pacific time.

  2. 2

    The model has no free quota at all

    Pro and preview models such as gemini-3.1-pro-preview have no free tier, so the error shows limit: 0 on the first request. No amount of waiting fixes it.

  3. 3

    A per-minute burst

    Too many requests or input tokens in one minute. The quotaId contains PerMinute, and the quota recovers within the minute.

  4. 4

    Extra keys from the same project

    Google applies limits per project, not per API key, so a second key created in the same project shares the same quota.

  5. 5

    Paid-tier spend limits

    On paid tiers a 429 can also mean a spend-based limit or shared capacity, in which case the body is the short "Resource has been exhausted" form without details.

Fixes

How to fix it.

Read the quotaId before retrying

The details array names every violated quota. A PerDay id means wait for the reset or switch model; a PerMinute id means wait the retryDelay. Google returns a short retryDelay even for daily caps, so retrying after it on a spent daily cap only fails again.

import json

def gemini_quota_kind(body: str) -> str:
    err = json.loads(body)
    err = err[0] if isinstance(err, list) else err   # OpenAI endpoint wraps it in a list
    ids = [v.get("quotaId", "") for d in err["error"].get("details", [])
           for v in d.get("violations", [])]
    if any("PerDay" in i for i in ids):
        return "daily"        # switch model or wait for midnight Pacific
    if any("PerMinute" in i for i in ids):
        return "per-minute"   # wait retryDelay, then retry
    return "unknown"

Per-minute: back off

Wait the retryDelay with exponential backoff and a little jitter. The official SDKs retry automatically, starting near one second and capping at sixty.

Daily: use another model

Each model has its own daily quota, so gemini-3.5-flash-lite, gemini-3.1-flash-lite, gemini-3.5-flash and the -latest aliases each add capacity. Otherwise wait for midnight Pacific time.

limit: 0: change model or enable billing

A model without a free tier never answers on a free key. Use a Flash or Flash-Lite model, or enable billing on the project.

Check what is left

Google no longer prints free-tier numbers in its docs; the current limits for your project are on the rate-limit page in AI Studio (aistudio.google.com/rate-limit).

With freelm

Handle it automatically.

freelm treats Gemini quotas as per model: a 429 sets aside only that model for that key, and the same call moves on to the key's other Gemini models and then to your other providers. When the quotaId names a daily cap, freelm rests the model for an hour instead of retrying after Google's 30 second hint; a per-minute 429 waits the retryDelay Google sends.

pip install freelm      # or: npm install freelm
export GEMINI_API_KEY=...   # plus any other free keys you have
freelm doctor               # one tiny live request per key

import freelm
llm = freelm.FreeLLM.from_env()
print(llm.text("Explain failover in one sentence."))

freelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.

Questions

Related questions.

When does the Gemini free quota reset?

Requests per day reset at midnight Pacific time, according to Google's rate-limit documentation checked on October 10, 2026. Per-minute quotas recover within a minute.

Does a second Gemini API key give me more quota?

Not if it belongs to the same Google Cloud project. Limits apply per project, not per key. A key from a separate project has its own quota.

Why does retryDelay say 29 seconds when the daily limit is gone?

Google returns a short RetryInfo hint even for daily caps, so a retry after it fails again. Check the quotaId: if it contains PerDay, wait for the reset or use another model.