← All free LLM API errors

Z.ai (GLM) · checked October 10, 2026

Z.ai GLM "Rate limit reached for requests" (1302) and 1305

Z.ai's 1302 "Rate limit reached for requests" (HTTP 429) means your account sent more requests to one model than its limit allows, usually too many at once: the free GLM Flash models are reported to allow one concurrent request. 1305 "The service may be temporarily overloaded" is a different problem: capacity on Z.ai's side.

The exact error

What you see.

Both come back with HTTP 429 and a string code in the error object. Messages can be localised depending on the account.

{"error":{"code":"1302","message":"Rate limit reached for requests"}}
{"error":{"code":"1305","message":"The service may be temporarily overloaded, please try again later"}}

Causes

Why it happens.

  1. 1

    1302: too many requests to one model

    Limits are per account and per model, and the free Flash models are reported to allow a single concurrent request. Official numbers are shown only in the console (z.ai/manage-apikey/rate-limits).

  2. 2

    1305: the model is overloaded

    Capacity on Z.ai's side. It has also been reported for requests sent without a User-Agent header.

  3. 3

    1113: a paid model

    The sibling error "Insufficient balance or no resource package" means the model is not free. GLM-5.3-Flash, for example, is paid.

Fixes

How to fix it.

Serialise calls per model

Send one request at a time to each free Flash model and queue the rest.

Back off on both codes

Treat 1305 as a temporary overload and retry after a short delay; treat 1302 as a signal to slow down.

Send a User-Agent

Some HTTP clients send none by default; set one explicitly.

Stay on the free models

As of October 10, 2026 the free models are GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash.

With freelm

Handle it automatically.

freelm's Z.ai provider offers only the three free Flash models, so a paid model id, and the 1113 balance error, cannot happen through it. A 1302 or 1305 sets aside only the model that returned it, for that key, and the call moves to the next Flash model or another provider. freelm sends a User-Agent on every request.

pip install freelm      # or: npm install freelm
export ZAI_API_KEY=...   # plus any other free keys you have
freelm doctor               # one tiny live request per key

freelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.

Questions

Related questions.

Which Z.ai GLM models are free?

GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash, per Z.ai's pricing page checked October 10, 2026. Other GLM models draw on a paid balance.

Why do I get 1302 with a single request?

Another request from the same account may still be running on that model, for example from another tool or a retry that has not finished. The free models allow very little concurrency.