← All free LLM API errors

Mistral · checked October 10, 2026

Mistral API "Rate limit exceeded" on the free plan

Mistral returns 429 "Rate limit exceeded" (type rate_limited, code 1300) when you pass a per-model limit on your workspace, either requests per second or tokens per minute. On the free plan, now called Free mode, limits are low and some models can be capped at zero, in which case retrying never helps. Older responses said "Requests rate limit exceeded".

The exact error

What you see.

HTTP 429. The first body is the September 2026 form; the second is the older bare form.

{"object":"error","message":"Rate limit exceeded","type":"rate_limited","param":null,"code":"1300","raw_status_code":429}
{"message":"Requests rate limit exceeded"}

Causes

Why it happens.

  1. 1

    Requests per second or tokens per minute

    Mistral sets both per model; on Free mode they are low enough that parallel calls trip them.

  2. 2

    A model capped at zero

    If the response headers show x-ratelimit-limit-req-minute: 0, that model is not available on your workspace's plan. Waiting does not help.

  3. 3

    Monthly usage spent

    When included usage runs out and pay-as-you-go is off, requests are refused until the next period.

Fixes

How to fix it.

Read the rate-limit headers

The headers show the limit and what remains for the model you called.

curl -si https://api.mistral.ai/v1/chat/completions \
  -H "Authorization: Bearer $MISTRAL_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"mistral-small-latest","messages":[{"role":"user","content":"hi"}]}' | grep -i ratelimit

Honour Retry-After and pace requests

Mistral's 429 responses carry Retry-After. Send one request at a time on Free mode and back off exponentially.

Check limits per model

The console lists your workspace's limits per model at admin.mistral.ai/plateforme/limits; the docs no longer publish free numbers.

Activate Free mode or pay-as-you-go

Both live at admin.mistral.ai/subscription. A model capped at zero needs a different plan, not a retry.

With freelm

Handle it automatically.

freelm paces each Mistral key to 2 requests a minute on the free tier, cools the key on a 429 for the time Mistral asks (or a minute when it does not say), and continues the call on your other providers. A model capped at zero shows up in freelm doctor as rate-limited on every check, which is the signal to change the workspace plan rather than wait.

pip install freelm      # or: npm install freelm
export MISTRAL_API_KEY=...   # plus any other free keys you have
freelm doctor               # one tiny live request per key

freelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.

Questions

Related questions.

What are Mistral's free API limits?

Mistral shows them per model in the console. Its docs list the dimensions, requests per second and tokens per minute, but publish no free-plan numbers as of October 2026.

Why do I get a 429 on the very first request?

Usually because that model is capped at zero for your workspace. Check the x-ratelimit-limit headers; if the limit is 0, change the plan or the model.