Thinking models (Gemini, gpt-oss) · checked October 10, 2026
Empty response from Gemini or gpt-oss: content null, finish_reason length
An empty reply with finish_reason "length" from a thinking model means the token budget ran out during hidden reasoning, before any visible text. Gemini counts thought tokens against max_output_tokens, and gpt-oss on Groq and OpenRouter does the same. Raise max_tokens, lower the reasoning effort, or use a model that does not think.
The exact error
What you see.
Not an HTTP error: the request succeeds with status 200. The first shape is the OpenAI-compatible response; the second is the native Gemini response.
{"choices": [{"index": 0, "message": {"role": "assistant", "content": null}, "finish_reason": "length"}],
"usage": {"completion_tokens": 256, "completion_tokens_details": {"reasoning_tokens": 256}}}candidates[0]: { content: {} /* no parts */, finishReason: "MAX_TOKENS" } usageMetadata: { thoughtsTokenCount: 16 }Causes
Why it happens.
- 1
max_tokens is smaller than the reasoning
Google's thinking guide says the output limit includes thought tokens and that a model hitting it while reasoning returns truncated or empty output, still billing the thinking. Groq's default max_completion_tokens of 1024 is flagged in its docs as possibly too low.
- 2
Reasoning that cannot be turned off
Gemini 2.5 Pro and the Gemini 3 models always reason, and gpt-oss accepts only low, medium or high effort.
- 3
Hiding reasoning does not save tokens
include_reasoning: false on Groq and exclude: true on OpenRouter hide the reasoning text but still spend the tokens and count them toward max_tokens.
- 4
A hidden cap in a gateway
A proxy or SDK that sets its own low max_tokens produces the same empty reply.
Fixes
How to fix it.
Give the model room and lower the effort
A few hundred tokens is a floor for thinking models; set the reasoning effort to the lowest the model accepts.
from openai import OpenAI
client = OpenAI(base_url="https://generativelanguage.googleapis.com/v1beta/openai/", api_key=GEMINI_API_KEY)
r = client.chat.completions.create(
model="gemini-3.5-flash",
messages=[{"role": "user", "content": "Summarise this in one line: ..."}],
reasoning_effort="low", # Gemini 3: low/medium/high (some reject "minimal")
max_tokens=2048,
)Per provider
Gemini on the OpenAI endpoint maps reasoning_effort to a thinking level on 3.x and accepts "none" only on 2.5 Flash models. Groq's gpt-oss takes reasoning_effort low, medium or high, while qwen3.8-27b accepts "none". OpenRouter takes reasoning: {effort: ...}; check the model's reasoning.mandatory flag before sending "none".
Detect it
If completion tokens minus reasoning tokens is about zero and finish_reason is length, the model ran out of budget while thinking: retry with a larger max_tokens rather than treating the empty text as an answer.
With freelm
Handle it automatically.
freelm's model="auto" prefers models that do not think (Flash-Lite and plain instruct models come first), and reasoning models are tagged so you get them when you ask for model="reasoning" or a concrete id. It passes reasoning_effort and max_tokens to the provider unchanged. An empty but successful reply is returned as a success, not retried, so give thinking models a budget of at least a few hundred tokens.
import freelm
llm = freelm.FreeLLM.from_env()
print(llm.text("Summarise this in one line: ...", max_tokens=512)) # auto: non-thinking models firstfreelm is an open-source Python and Node.js library that pools the free tiers of Gemini, Groq, OpenRouter, Cloudflare Workers AI, Mistral, NVIDIA NIM, Z.ai and Cohere behind one OpenAI-compatible call. See how its failover works.
Questions
Related questions.
Why does gemini-3.5-flash return empty text?
It is a thinking model, and the max_tokens you set was used up by reasoning before any visible text. Raise max_tokens and set reasoning_effort to low.
Does include_reasoning: false save tokens?
No. It only hides the reasoning text; the model still spends the tokens and they still count toward max_tokens.
Sources