> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Third-party integrations

> Use Together AI models through partner SDKs and integrations.

The Together API is [OpenAI-compatible](/docs/inference/openai-compatibility), so most third-party SDKs work by pointing them at `https://api.together.ai/v1` and supplying your [Together API key](/docs/api-keys-authentication). The integrations below ship dedicated Together support, with first-class clients, helpers, or providers.

For agent frameworks (CrewAI, LangGraph, DSPy, PydanticAI, AutoGen, Agno, Composio), see the dedicated pages under [Framework integrations](/docs/agent-integrations). If you're migrating from the v1 Python SDK, see the [Python SDK v2 migration](/docs/pythonv2-migration-guide) appendix.

## Integration guides

Each of these integrations has a dedicated guide with installation steps and runnable examples:

<CardGroup cols={2}>
  <Card title="Hugging Face" icon="robot-face" href="/docs/quickstart-using-hugging-face-inference">
    Call Together models through Hugging Face Inference clients and the Hub's provider routing.
  </Card>

  <Card title="Vercel AI SDK" icon="brand-vercel" href="/docs/using-together-with-vercels-ai-sdk">
    Generate text, stream, call tools, and get structured outputs with the `@ai-sdk/togetherai` provider.
  </Card>

  <Card title="Mastra" icon="route-2" href="/docs/using-together-with-mastra">
    Point Mastra's model router at Together models to build agents and workflows.
  </Card>

  <Card title="Next.js" icon="brand-nextjs" href="/docs/nextjs-chat-quickstart">
    Build a streaming chat app with the Together SDK, the Vercel AI SDK, or Mastra.
  </Card>
</CardGroup>

## LangChain

[LangChain](https://www.langchain.com/) is a framework for building context-aware, reasoning applications powered by LLMs. The `langchain-together` package provides chat models and embeddings.

```bash Shell theme={null}
pip install --upgrade langchain-together
```

```python Python theme={null}
from langchain_together import ChatTogether

chat = ChatTogether(
    model="moonshotai/Kimi-K3",
    extra_body={"reasoning": {"enabled": False}},
)

for chunk in chat.stream("Tell me fun things to do in NYC"):
    print(chunk.content, end="", flush=True)
```

For more LangChain integration details, see the [LangChain provider docs](https://python.langchain.com/docs/integrations/providers/together/).

## LlamaIndex

[LlamaIndex](https://www.llamaindex.ai/) is a data framework for connecting custom data sources to LLMs. Together AI works with LlamaIndex through the `OpenAILike` LLM and dedicated embedding classes.

```bash Shell theme={null}
pip install llama-index llama-index-llms-openai-like
```

```python Python theme={null}
import os
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="moonshotai/Kimi-K3",
    api_base="https://api.together.ai/v1",
    api_key=os.environ["TOGETHER_API_KEY"],
    is_chat_model=True,
    is_function_calling_model=True,
    temperature=0.1,
)

response = llm.complete("Explain large language models in 500 words.")
print(response)
```

For more LlamaIndex integration details, see the [LlamaIndex Together LLM docs](https://docs.llamaindex.ai/en/stable/examples/llm/together/).

## LiteLLM

[LiteLLM](https://docs.litellm.ai/) is an open-source gateway with a single OpenAI-format interface for over 100 LLM providers, available as a Python SDK and a proxy server. LiteLLM 1.100.0 and later ships official Together AI support, with a model registry that syncs against the serverless catalog daily.

```bash Shell theme={null}
pip install "litellm>=1.100.0"
export TOGETHERAI_API_KEY=<your_key>
```

```python Python theme={null}
import litellm

response = litellm.completion(
    model="together_ai/zai-org/GLM-5.3",
    messages=[
        {
            "role": "user",
            "content": "What are some fun things to do in New York?",
        }
    ],
)
print(response.choices[0].message.content)
```

Streaming (`stream=True`), function calling (`tools`), JSON mode (`response_format`), and `reasoning_effort` all pass through in the OpenAI format. Reasoning models return their reasoning in a separate `reasoning_content` field, so `content` holds only the final answer.

To run Claude Code on Together models through the LiteLLM proxy, see the [LiteLLM + Claude Code guide](/docs/using-together-with-litellm).

## Helicone

[Helicone](https://www.helicone.ai/) is an open-source LLM observability platform. Route Together requests through Helicone's gateway by overriding `base_url` and adding the auth header.

<CodeGroup>
  ```python Python theme={null}
  import os
  from together import Together

  client = Together(
      api_key=os.environ["TOGETHER_API_KEY"],
      base_url="https://together.hconeai.com/v1",
      default_headers={
          "Helicone-Auth": f"Bearer {os.environ['HELICONE_API_KEY']}",
      },
  )

  stream = client.chat.completions.create(
      model="Qwen/Qwen3.5-9B",
      messages=[
          {
              "role": "user",
              "content": "What are some fun things to do in New York?",
          }
      ],
      stream=True,
  )

  for chunk in stream:
      if chunk.choices:
          print(chunk.choices[0].delta.content or "", end="", flush=True)
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const client = new Together({
    apiKey: process.env.TOGETHER_API_KEY,
    baseURL: "https://together.hconeai.com/v1",
    defaultHeaders: {
      "Helicone-Auth": `Bearer ${process.env.HELICONE_API_KEY}`,
    },
  });

  const stream = await client.chat.completions.create({
    model: "Qwen/Qwen3.5-9B",
    messages: [
      { role: "user", content: "What are some fun things to do in New York?" },
    ],
    stream: true,
  });

  for await (const chunk of stream) {
    process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  }
  ```

  ```bash cURL theme={null}
  curl https://together.hconeai.com/v1/chat/completions \
    -H "Authorization: Bearer $TOGETHER_API_KEY" \
    -H "Helicone-Auth: Bearer $HELICONE_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "Qwen/Qwen3.5-9B",
      "messages": [
        {"role": "user", "content": "What are some fun things to do in New York?"}
      ],
      "stream": true
    }'
  ```
</CodeGroup>

## Agent frameworks

Each framework below has a dedicated guide with installation, model selection, and runnable examples:

* [CrewAI](/docs/crewai): Open-source orchestration for multi-agent workflows.
* [LangGraph](/docs/langgraph): Stateful, multi-actor applications built on LangChain.
* [DSPy](/docs/dspy): Modular AI systems written in code instead of prompt strings.
* [PydanticAI](/docs/pydanticai): Typed agent framework from the Pydantic team.
* [AutoGen (AG2)](/docs/autogen): Conversational multi-agent systems.
* [Agno](/docs/agno): Open-source library for multimodal agents.
* [Composio](/docs/composio): Tool-use platform for connecting agents to external services.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.