← All free LLM API errors

Thinking models (Gemini, gpt-oss) · checked October 10, 2026

Empty response from Gemini or gpt-oss: content null, finish_reason length

An empty reply with finish_reason "length" from a thinking model means the token budget ran out during hidden reasoning, before any visible text. Gemini counts thought tokens against max_output_tokens, and gpt-oss on Groq and OpenRouter does the same. Raise max_tokens, lower the reasoning effort, or use a model that does not think.

The exact error

What you see.

Not an HTTP error: the request succeeds with status 200. The first shape is the OpenAI-compatible response; the second is the native Gemini response.

{"choices": [{"index": 0, "message": {"role": "assistant", "content": null}, "finish_reason": "length"}],
 "usage": {"completion_tokens": 256, "completion_tokens_details": {"reasoning_tokens": 256}}}
candidates[0]: { content: {} /* no parts */, finishReason: "MAX_TOKENS" }   usageMetadata: { thoughtsTokenCount: 16 }

Causes

Why it happens.

  1. 1

    max_tokens is smaller than the reasoning

    Google's thinking guide says the output limit includes thought tokens and that a model hitting it while reasoning returns truncated or empty output, still billing the thinking. Groq's default max_completion_tokens of 1024 is flagged in its docs as possibly too low.

  2. 2

    Reasoning that cannot be turned off

    Gemini 2.5 Pro and the Gemini 3 models always reason, and gpt-oss accepts only low, medium or high effort.

  3. 3

    Hiding reasoning does not save tokens

    include_reasoning: false on Groq and exclude: true on OpenRouter hide the reasoning text but still spend the tokens and count them toward max_tokens.

  4. 4

    A hidden cap in a gateway

    A proxy or SDK that sets its own low max_tokens produces the same empty reply.

Fixes

How to fix it.

Give the model room and lower the effort

A few hundred tokens is a floor for thinking models; set the reasoning effort to the lowest the model accepts.

from openai import OpenAI
client = OpenAI(base_url="https://generativelanguage.googleapis.com/v1beta/openai/", api_key=GEMINI_API_KEY)
r = client.chat.completions.create(
    model="gemini-3.5-flash",
    messages=[{"role": "user", "content": "Summarise this in one line: ..."}],
    reasoning_effort="low",   # Gemini 3: low/medium/high (some reject "minimal")
    max_tokens=2048,
)

Per provider

Gemini on the OpenAI endpoint maps reasoning_effort to a thinking level on 3.x and accepts "none" only on 2.5 Flash models. Groq's gpt-oss takes reasoning_effort low, medium or high, while qwen3.8-27b accepts "none". OpenRouter takes reasoning: {effort: ...}; check the model's reasoning.mandatory flag before sending "none".

Detect it

If completion tokens minus reasoning tokens is about zero and finish_reason is length, the model ran out of budget while thinking: retry with a larger max_tokens rather than treating the empty text as an answer.

With freelm

Handle it automatically.

freelm's model="auto" prefers models that do not think (Flash-Lite and plain instruct models come first), and reasoning models are tagged so you get them when you ask for model="reasoning" or a concrete id. It passes reasoning_effort and max_tokens to the provider unchanged. An empty but successful reply is returned as a success, not retried, so give thinking models a budget of at least a few hundred tokens.

import freelm
llm = freelm.FreeLLM.from_env()
print(llm.text("Summarise this in one line: ...", max_tokens=512))   # auto: non-thinking models first

freelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.

Questions

Related questions.

Why does gemini-3.5-flash return empty text?

It is a thinking model, and the max_tokens you set was used up by reasoning before any visible text. Raise max_tokens and set reasoning_effort to low.

Does include_reasoning: false save tokens?

No. It only hides the reasoning text; the model still spends the tokens and they still count toward max_tokens.