> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# GPT-OSS quickstart

> Call OpenAI's GPT-OSS 120B open-weight reasoning model on Together.

GPT-OSS 120B is OpenAI's open-weight mixture-of-experts (MoE) reasoning model. It thinks step by step before answering and exposes an adjustable reasoning effort, so you can trade depth for cost and latency.

The model ID is `openai/gpt-oss-120b`. Pricing is \$0.15 per 1M input tokens and \$0.60 per 1M output tokens, with a 128K-token context window.

<Note>
  The smaller `openai/gpt-oss-20b` was removed from serverless inference on September 15, 2026, with `Qwen/Qwen3.5-9B` as its listed replacement. It remains available for [fine-tuning](/docs/fine-tuning/supported-models) and [dedicated endpoints](/docs/dedicated-endpoints/models). See [Deprecations](/docs/deprecations) for details, and [Recommended models](/docs/inference/recommended-models) for Together's current picks by use case.
</Note>

## Call GPT-OSS 120B

Reasoning is on by default. The thinking trace arrives on `reasoning` and the final answer arrives on `content`. The two share the same completion budget, so give `max_tokens` headroom and parse `content` only.

<CodeGroup>
  ```python Python theme={null}
  from together import Together

  client = Together()

  completion = client.chat.completions.create(
      model="openai/gpt-oss-120b",
      messages=[
          {
              "role": "user",
              "content": "If all roses are flowers and some flowers are red, can we conclude that some roses are red?",
          }
      ],
      temperature=1.0,
      top_p=1.0,
      max_tokens=8192,
  )

  print(completion.choices[0].message.content)
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const together = new Together();

  const completion = await together.chat.completions.create({
    model: "openai/gpt-oss-120b",
    messages: [
      {
        role: "user",
        content:
          "If all roses are flowers and some flowers are red, can we conclude that some roses are red?",
      },
    ],
    temperature: 1.0,
    top_p: 1.0,
    max_tokens: 8192,
  });

  console.log(completion.choices[0].message.content);
  ```

  ```bash cURL theme={null}
  curl -X POST "https://api.together.ai/v1/chat/completions" \
       -H "Authorization: Bearer $TOGETHER_API_KEY" \
       -H "Content-Type: application/json" \
       -d '{
          "model": "openai/gpt-oss-120b",
          "messages": [
            {"role": "user", "content": "If all roses are flowers and some flowers are red, can we conclude that some roses are red?"}
          ],
          "temperature": 1.0,
          "top_p": 1.0,
          "max_tokens": 8192
       }'
  ```
</CodeGroup>

## Set the reasoning effort

`reasoning_effort` accepts `"low"`, `"medium"`, and `"high"`. The default is `"medium"`.

* `"low"`: shallow reasoning. Use for short, high-volume calls.
* `"medium"`: balanced depth, the default. Use for most work.
* `"high"`: deep reasoning. Use for hard multi-step problems, and set `max_tokens` generously, since the trace can run to tens of thousands of tokens on hard prompts.

Reasoning cannot be disabled entirely. Values outside this set are accepted silently instead of returning an error, so validate the value in your own code.

<CodeGroup>
  ```python Python theme={null}
  from together import Together

  client = Together()

  completion = client.chat.completions.create(
      model="openai/gpt-oss-120b",
      messages=[
          {
              "role": "user",
              "content": "Plan a migration from REST to gRPC for a payments service.",
          }
      ],
      reasoning_effort="high",
      max_tokens=32768,
  )

  print(completion.choices[0].message.content)
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const together = new Together();

  const completion = await together.chat.completions.create({
    model: "openai/gpt-oss-120b",
    messages: [
      {
        role: "user",
        content: "Plan a migration from REST to gRPC for a payments service.",
      },
    ],
    reasoning_effort: "high",
    max_tokens: 32768,
  });

  console.log(completion.choices[0].message.content);
  ```
</CodeGroup>

For broader guidance on reasoning controls and prompting, see [Reasoning](/docs/inference/chat/reasoning).

## Stream the reasoning trace and the answer

Set `stream=True` to render the trace and the answer as they arrive. Handle both channels, and skip chunks where `choices` is empty, since Together emits a final usage-only chunk.

<CodeGroup>
  ```python Python theme={null}
  from together import Together

  client = Together()

  stream = client.chat.completions.create(
      model="openai/gpt-oss-120b",
      messages=[{"role": "user", "content": "Explain why the sky is blue."}],
      max_tokens=8192,
      stream=True,
  )

  in_answer = False
  for chunk in stream:
      if not chunk.choices:
          continue
      delta = chunk.choices[0].delta
      thinking = getattr(delta, "reasoning", None)
      if thinking:
          print(thinking, end="", flush=True)
      if delta.content:
          if not in_answer:
              print("\n--- answer ---")
              in_answer = True
          print(delta.content, end="", flush=True)
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";
  import type { ChatCompletionChunk } from "together-ai/resources/chat/completions";

  const together = new Together();

  type ReasoningDelta = ChatCompletionChunk.Choice.Delta & {
    reasoning?: string;
  };

  const stream = await together.chat.completions.create({
    model: "openai/gpt-oss-120b",
    messages: [{ role: "user", content: "Explain why the sky is blue." }],
    max_tokens: 8192,
    stream: true,
  });

  let inAnswer = false;
  for await (const chunk of stream) {
    if (!chunk.choices?.length) continue;
    const delta = chunk.choices[0]?.delta as ReasoningDelta;

    if (delta?.reasoning) process.stdout.write(delta.reasoning);

    if (delta?.content) {
      if (!inAnswer) {
        process.stdout.write("\n--- answer ---\n");
        inAnswer = true;
      }
      process.stdout.write(delta.content);
    }
  }
  ```
</CodeGroup>

## Call tools

Declare functions in `tools`. When the model returns `tool_calls`, append the assistant message to history, append one `tool` message per call with the matching `tool_call_id`, then call again.

<CodeGroup>
  ```python Python theme={null}
  import json

  from together import Together

  client = Together()

  tools = [
      {
          "type": "function",
          "function": {
              "name": "get_weather",
              "description": "Get the current weather for a city.",
              "parameters": {
                  "type": "object",
                  "properties": {
                      "city": {
                          "type": "string",
                          "description": "City name, e.g. Paris",
                      },
                      "unit": {
                          "type": "string",
                          "enum": ["celsius", "fahrenheit"],
                      },
                  },
                  "required": ["city"],
                  "additionalProperties": False,
              },
          },
      }
  ]


  def get_weather(city, unit="celsius"):
      return {
          "city": city,
          "temperature": 21,
          "unit": unit,
          "conditions": "sunny",
      }


  messages = [{"role": "user", "content": "What's the weather in Paris?"}]

  for _ in range(5):
      response = client.chat.completions.create(
          model="openai/gpt-oss-120b",
          messages=messages,
          tools=tools,
          max_tokens=8192,
      )
      choice = response.choices[0]
      message = choice.message

      messages.append(message.model_dump(exclude_none=True))

      if choice.finish_reason != "tool_calls" or not message.tool_calls:
          print(message.content)
          break

      for call in message.tool_calls:
          try:
              args = json.loads(call.function.arguments)
          except json.JSONDecodeError:
              args = {}
          messages.append(
              {
                  "role": "tool",
                  "tool_call_id": call.id,
                  "content": json.dumps(get_weather(**args)),
              }
          )
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const together = new Together();

  const tools = [
    {
      type: "function" as const,
      function: {
        name: "get_weather",
        description: "Get the current weather for a city.",
        parameters: {
          type: "object",
          properties: {
            city: { type: "string", description: "City name, e.g. Paris" },
            unit: { type: "string", enum: ["celsius", "fahrenheit"] },
          },
          required: ["city"],
          additionalProperties: false,
        },
      },
    },
  ];

  function getWeather(city: string, unit = "celsius") {
    return { city, temperature: 21, unit, conditions: "sunny" };
  }

  const messages: any[] = [
    { role: "user", content: "What's the weather in Paris?" },
  ];

  for (let i = 0; i < 5; i++) {
    const response = await together.chat.completions.create({
      model: "openai/gpt-oss-120b",
      messages,
      tools,
      max_tokens: 8192,
    });

    const choice = response.choices[0];
    const message = choice.message;

    messages.push(message);

    if (choice.finish_reason !== "tool_calls" || !message.tool_calls?.length) {
      console.log(message.content);
      break;
    }

    for (const call of message.tool_calls) {
      let args: { city: string; unit?: string };
      try {
        args = JSON.parse(call.function.arguments);
      } catch {
        args = { city: "" };
      }
      messages.push({
        role: "tool",
        tool_call_id: call.id,
        content: JSON.stringify(getWeather(args.city, args.unit)),
      });
    }
  }
  ```
</CodeGroup>

For the full pattern, including parallel calls and best practices, see [Function calling](/docs/inference/function-calling/overview).

## Constrain the output to a schema

Pass a JSON schema through `response_format` with `"strict": True` to constrain the final `content`. The whole thinking trace is spent before the first schema-constrained token is emitted, so keep `max_tokens` generous. Parse `content` only, never `reasoning`.

<CodeGroup>
  ```python Python theme={null}
  import json

  from together import Together

  client = Together()

  completion = client.chat.completions.create(
      model="openai/gpt-oss-120b",
      max_tokens=4096,
      messages=[{"role": "user", "content": "Ada Lovelace was 36 years old."}],
      response_format={
          "type": "json_schema",
          "json_schema": {
              "name": "person",
              "strict": True,
              "schema": {
                  "type": "object",
                  "properties": {
                      "name": {"type": "string"},
                      "age": {"type": "integer"},
                  },
                  "required": ["name", "age"],
                  "additionalProperties": False,
              },
          },
      },
  )

  person = json.loads(completion.choices[0].message.content)
  print(person)
  # -> {'name': 'Ada Lovelace', 'age': 36}
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const together = new Together();

  const completion = await together.chat.completions.create({
    model: "openai/gpt-oss-120b",
    max_tokens: 4096,
    messages: [{ role: "user", content: "Ada Lovelace was 36 years old." }],
    response_format: {
      type: "json_schema",
      json_schema: {
        name: "person",
        strict: true,
        schema: {
          type: "object",
          properties: {
            name: { type: "string" },
            age: { type: "integer" },
          },
          required: ["name", "age"],
          additionalProperties: false,
        },
      },
    },
  });

  const person = JSON.parse(completion.choices[0].message.content ?? "{}");
  console.log(person);
  // -> { name: 'Ada Lovelace', age: 36 }
  ```
</CodeGroup>

For schema design guidance, see [Structured outputs](/docs/inference/chat/structured-outputs).

## Usage tips

| Tip | Rationale |
| - | - |
| **Give `max_tokens` real headroom** | The trace shares the completion budget with the answer. A tight cap spends the allowance on reasoning and returns an empty or truncated answer. |
| **Parse `content` only** | Never run JSON parsing over the thinking trace on `reasoning`. |
| **Match `reasoning_effort` to the task** | `"low"` for routine calls, `"high"` for hard problems. Reasoning tokens bill as output tokens. |
| **Temperature = 1.0, top\_p = 1.0** | OpenAI's recommended defaults for the GPT-OSS family. |
| **Think in goals, not steps** | Give high-level objectives and let the model determine the methodology. Micromanaging steps limits its reasoning. |

## Next steps

<CardGroup cols={2}>
  <Card title="Reasoning" icon="brain" href="/docs/inference/chat/reasoning">
    Control reasoning depth and handle reasoning output across models.
  </Card>

  <Card title="Function calling" icon="tool" href="/docs/inference/function-calling/overview">
    Build tool-calling loops against any function-calling model.
  </Card>

  <Card title="Recommended models" icon="star" href="/docs/inference/recommended-models">
    See Together's current picks for every use case.
  </Card>

  <Card title="Serverless models" icon="stack-2" href="/docs/serverless/models">
    Browse every model, context length, and price on serverless inference.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.