Skip to main content
GPT-OSS 120B is OpenAI’s open-weight mixture-of-experts (MoE) reasoning model. It thinks step by step before answering and exposes an adjustable reasoning effort, so you can trade depth for cost and latency. The model ID is openai/gpt-oss-120b. Pricing is $0.15 per 1M input tokens and $0.60 per 1M output tokens, with a 128K-token context window.
The smaller openai/gpt-oss-20b was removed from serverless inference on September 15, 2026, with Qwen/Qwen3.5-9B as its listed replacement. It remains available for fine-tuning and dedicated endpoints. See Deprecations for details, and Recommended models for Together’s current picks by use case.

Call GPT-OSS 120B

Reasoning is on by default. The thinking trace arrives on reasoning and the final answer arrives on content. The two share the same completion budget, so give max_tokens headroom and parse content only.

Set the reasoning effort

reasoning_effort accepts "low", "medium", and "high". The default is "medium".
  • "low": shallow reasoning. Use for short, high-volume calls.
  • "medium": balanced depth, the default. Use for most work.
  • "high": deep reasoning. Use for hard multi-step problems, and set max_tokens generously, since the trace can run to tens of thousands of tokens on hard prompts.
Reasoning cannot be disabled entirely. Values outside this set are accepted silently instead of returning an error, so validate the value in your own code.
For broader guidance on reasoning controls and prompting, see Reasoning.

Stream the reasoning trace and the answer

Set stream=True to render the trace and the answer as they arrive. Handle both channels, and skip chunks where choices is empty, since Together emits a final usage-only chunk.

Call tools

Declare functions in tools. When the model returns tool_calls, append the assistant message to history, append one tool message per call with the matching tool_call_id, then call again.
For the full pattern, including parallel calls and best practices, see Function calling.

Constrain the output to a schema

Pass a JSON schema through response_format with "strict": True to constrain the final content. The whole thinking trace is spent before the first schema-constrained token is emitted, so keep max_tokens generous. Parse content only, never reasoning.
For schema design guidance, see Structured outputs.

Usage tips

Next steps

Reasoning

Control reasoning depth and handle reasoning output across models.

Function calling

Build tool-calling loops against any function-calling model.

Recommended models

See Together’s current picks for every use case.

Serverless models

Browse every model, context length, and price on serverless inference.