> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Hugging Face inference quickstart

> Call Together models from Hugging Face Inference clients.

Together AI is an [inference provider](https://huggingface.co/docs/inference-providers) on the Hugging Face Hub. You can call Together models through the Hugging Face client libraries, the OpenAI client, or directly from model page widgets on the Hub.

## Authentication and billing

When using Together AI through Hugging Face, you have two options for authentication:

* **Direct requests:** Add your Together API key to your Hugging Face user account settings. Inference requests are sent directly to Together AI, and billing is handled by your Together AI account.
* **Routed requests:** If you don't configure a Together API key, your requests are routed through Hugging Face and authenticated with a Hugging Face token. Billing for routed requests is applied to your Hugging Face account at standard provider API rates. You don't need a Together AI account for this option.

To add a Together API key to your Hugging Face settings:

1. Go to your [Hugging Face inference provider settings](https://huggingface.co/settings/inference-providers).
2. Add your Together API key under the Together AI provider.
3. Optionally, set your preferred provider order, which controls the display order in model widgets and code snippets.

<Info>
  You can browse all [Together AI models on the Hub](https://huggingface.co/models?inference_provider=together\&sort=trending) and try them directly in the model page widget.
</Info>

## Installation

Install the Hugging Face client for your language:

<CodeGroup>
  ```bash Python theme={null}
  pip install "huggingface_hub>=0.29.0"
  ```

  ```bash TypeScript theme={null}
  npm install @huggingface/inference
  ```
</CodeGroup>

## Chat completions

Pass `provider="together"` to route the request to Together AI. The API key can be a Together AI key (direct requests) or a Hugging Face token (routed requests).

<CodeGroup>
  ```python Python theme={null}
  from huggingface_hub import InferenceClient

  client = InferenceClient(
      provider="together",
      api_key="your_api_key",  # Together AI key or Hugging Face token
  )

  completion = client.chat.completions.create(
      model="moonshotai/Kimi-K3",
      messages=[{"role": "user", "content": "What is the capital of France?"}],
      max_tokens=500,
  )

  print(completion.choices[0].message)
  ```

  ```typescript TypeScript theme={null}
  import { InferenceClient } from "@huggingface/inference";

  const client = new InferenceClient("your_api_key"); // Together AI key or Hugging Face token

  const chatCompletion = await client.chatCompletion({
    model: "moonshotai/Kimi-K3",
    messages: [{ role: "user", content: "What is the capital of France?" }],
    provider: "together",
    max_tokens: 500,
  });

  console.log(chatCompletion.choices[0].message);
  ```
</CodeGroup>

<Note>
  In `@huggingface/inference` v3 and later, the `HfInference` class is a deprecated alias slated for removal. Use `InferenceClient` instead.
</Note>

Kimi K3 reasons by default, so the message includes a `reasoning_content` field alongside `content`. The Hugging Face clients only forward their own parameter set, so you can't pass Together-specific parameters like `reasoning` through them. If you need to control reasoning, call the [Together API directly](/docs/inference/chat/reasoning).

You can swap in any [Together AI text model on the Hub](https://huggingface.co/models?inference_provider=together\&other=text-generation-inference\&sort=trending).

## OpenAI client library

You can also reach Together models through Hugging Face's OpenAI-compatible router with the [OpenAI Python client](https://github.com/openai/openai-python). The router requires a Hugging Face token. Append `:together` to the model ID to pin the provider:

```python Python theme={null}
from openai import OpenAI

client = OpenAI(
    base_url="https://router.huggingface.co/v1",
    api_key="your_hf_token",
)

completion = client.chat.completions.create(
    model="moonshotai/Kimi-K3:together",
    messages=[{"role": "user", "content": "What is the capital of France?"}],
    max_tokens=500,
)

print(completion.choices[0].message)
```

You can copy a ready-made snippet from any [model page](https://huggingface.co/moonshotai/Kimi-K3?inference_api=true\&inference_provider=together\&language=python) on the Hub.

## Other modalities

The Hub's provider mapping for Together AI currently covers text generation, speech recognition, and embeddings. For image generation and text-to-speech, call the Together API directly. Start with the [image generation overview](/docs/inference/images/overview).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.