Model notes
Per-model quirks to be aware of.
These are documented behaviors that differ from the generic OpenAI request shape.
kimi-k2.6
kimi-k2.6 ignores temperature, top_p, and penalty parameters.
MiniMax models
MiniMax models ignore presence_penalty and frequency_penalty parameters.
gpt-oss-20b
gpt-oss-20b requires max_tokens ≥ 500 to produce output. Accepts reasoning_effort.
gpt-5.x and gpt-6-astra
gpt-5.5, gpt-5.6-sol, and gpt-6-astra are reasoning models. The gateway translates max_tokens to the model's max_completion_tokens field for you, but temperature, top_p, penalty params, and logprobs are silently dropped — only default sampling is supported.
thinking and reasoning_effort are model-specific
These are not gateway-wide flags, and the gateway does not reject them with an error. Support is per-model:
claude-*acceptsthinking; reasoning content comes back as areasoningfield (see Anthropic (Claude) models).gpt-oss-*,gpt-5.*, andgpt-6-astraacceptreasoning_effort.gpt-5.*andgpt-6-astrasilently dropthinkingif sent.- Every other provider forwards these fields as-is to its upstream API, which may ignore or error on them depending on that provider's own request validation.
Omit both fields unless you're targeting one of the reasoning models above.