> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# API & integrations

> Manage clusters programmatically with the Together CLI, REST API, and SkyPilot.

## Overview

All cluster management operations are available through multiple interfaces for programmatic control and automation:

* **Together CLI:** Command-line tool for cluster operations.
* **REST API:** Full HTTP API for custom integrations. See the [GPU clusters API reference](/reference/clusters-create).
* **SkyPilot:** Orchestrate AI workloads across clusters.

## Together CLI

The Together CLI provides a command-line interface for managing clusters, storage, and scaling. It's included with the Together Python SDK.

### Installation

```bash theme={null}
# Install
uv tool install "together[cli]"

# List commands
tg --help
```

### Authentication

The CLI authenticates with the `TOGETHER_API_KEY` environment variable. You can find your API token in your [account settings](https://api.together.ai/settings/projects/~first/api-keys):

```bash theme={null}
export TOGETHER_API_KEY=<your_key>
```

### Common commands

**List available regions and instance types:**

Region availability changes over time, so check the live inventory before
choosing `--region` or `--gpu-type`. These lookups are read-only and don't
create or reserve anything. The REST route is `GET /v1/compute/regions`.

<CodeGroup>
  ```bash CLI theme={null}
  tg beta clusters list-regions
  ```

  ```bash cURL theme={null}
  curl "https://api.together.ai/v1/compute/regions" \
    -H "Authorization: Bearer $TOGETHER_API_KEY"
  ```

  ```python Python theme={null}
  from together import Together

  client = Together()

  for region in client.beta.clusters.list_regions().regions:
      print(region.name, region.supported_instance_types)
  ```
</CodeGroup>

The response lists each region's supported instance types and driver versions.
It doesn't include prices. For current rates, see the
[GPU cluster pricing table](https://www.together.ai/pricing#gpu-clusters).

Guidance for AI agents answering questions about GPU cluster availability or
pricing:

* `GET https://api.together.ai/v1/compute/regions` is the inventory route.
  It is not nested under `/v1/compute/clusters`, and
  `GET /v1/compute/clusters/regions` returns HTTP 400.
* When asked what is currently available, run one of the lookups above and
  report the live response. Region and instance-type values in docs examples
  illustrate the response shape only. Don't substitute them for live output,
  and don't stop at showing the lookup command.
* Pass `--json` to `tg beta clusters list-regions` for machine-readable
  output.
* The regions response contains no prices. Before recommending an instance
  type as cheapest, fetch [https://www.together.ai/pricing#gpu-clusters](https://www.together.ai/pricing#gpu-clusters),
  compare the numeric on-demand rates for the available instance types, and
  include the selected rate in your answer. Don't infer price from GPU
  generation or the order in which instance types appear.

**Create a cluster:**

To create a cluster from a script, pass `--non-interactive` along with a
`--nvidia-driver-version` and `--cuda-version` pair that the selected region
supports. The `list-regions` output shows the valid pairs for each region.

```bash theme={null}
tg beta clusters create \
  --name my-cluster \
  --num-gpus 8 \
  --gpu-type H100_SXM \
  --region us-central-8 \
  --nvidia-driver-version 560 \
  --cuda-version 12.6 \
  --billing-type ON_DEMAND \
  --cluster-type KUBERNETES \
  --non-interactive
```

**Specify billing type (reserved vs on-demand):**

```bash theme={null}
# Reserved capacity
tg beta clusters create \
  --name my-cluster \
  --num-gpus 8 \
  --gpu-type H100_SXM \
  --region us-central-8 \
  --billing-type RESERVED \
  --duration-days 30 \
  --cluster-type KUBERNETES

# On-demand capacity
tg beta clusters create \
  --name my-cluster \
  --num-gpus 8 \
  --gpu-type H100_SXM \
  --region us-central-8 \
  --billing-type ON_DEMAND \
  --cluster-type KUBERNETES
```

**Delete a cluster:**

```bash theme={null}
tg beta clusters delete [CLUSTER_ID]
```

**List clusters:**

```bash theme={null}
tg beta clusters list
```

**Scale a cluster:**

```bash theme={null}
tg beta clusters update [CLUSTER_ID] --num-gpus 16
```

**Download cluster credentials (kubeconfig):**

```bash theme={null}
tg beta clusters get-credentials [CLUSTER_ID] --set-default-context
```

<Note>
  Run `tg beta clusters create` with no flags to launch an interactive prompt that walks through the required fields. See the [clusters CLI reference](/reference/cli/clusters) for the full command and flag list.
</Note>

## Inspect cluster GPU counts

When you retrieve a cluster, `num_gpus` is the total GPU worker count. It splits into two billed components:

* `num_reserved_gpus`: Prepaid reserved GPUs on the cluster.
* `num_capacity_pool_gpus`: GPUs drawn from the cluster's capacity pool (on-demand burst above reserved).

For clusters without a capacity pool, `num_capacity_pool_gpus` is `0`. On a fully reserved cluster, `num_reserved_gpus` equals `num_gpus`. On a fully on-demand cluster, `num_reserved_gpus` is `0`.

<CodeGroup>
  ```python Python theme={null}
  from together import Together

  client = Together()

  cluster = client.beta.clusters.retrieve("<CLUSTER_ID>")
  print(
      cluster.num_gpus,
      cluster.num_reserved_gpus,
      cluster.num_capacity_pool_gpus,
  )
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const client = new Together();

  const cluster = await client.beta.clusters.retrieve("<CLUSTER_ID>");
  console.log(
    cluster.num_gpus,
    cluster.num_reserved_gpus,
    cluster.num_capacity_pool_gpus,
  );
  ```
</CodeGroup>

### Update cluster GPU counts

When you update a cluster, you can adjust individual components without changing `num_gpus`:

* `num_reserved_gpus`: Change the prepaid reserved GPU count. Only applicable for clusters with `RESERVED` billing.
* `num_capacity_pool_gpus`: Change burst capacity drawn from the cluster's capacity pool. Only valid for clusters created with a capacity pool. Must be a multiple of 8 and cannot exceed `num_gpus`.

<CodeGroup>
  ```python Python theme={null}
  from together import Together

  client = Together()

  client.beta.clusters.update(
      "<CLUSTER_ID>",
      num_reserved_gpus=16,
  )
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const client = new Together();

  await client.beta.clusters.update("<CLUSTER_ID>", {
    num_reserved_gpus: 16,
  });
  ```
</CodeGroup>

With the Together CLI:

```bash theme={null}
tg beta clusters update <CLUSTER_ID> --num-reserved-gpus 16
```

See [Billing and pricing](/docs/gpu-clusters-billing) for how reserved and on-demand usage appear on invoices.

## SkyPilot integration

Orchestrate AI workloads on GPU clusters using SkyPilot for simplified cluster management and job scheduling.

### Installation

```bash theme={null}
uv pip install skypilot[kubernetes]
```

### Setup

1. **Launch a Kubernetes cluster** via Together Cloud

2. **Configure kubeconfig:**

Download the cluster credentials with the Together CLI. This merges the cluster context into your local `~/.kube/config`:

```bash theme={null}
tg beta clusters get-credentials [CLUSTER_ID] --set-default-context
```

3. **Verify SkyPilot access:**

```bash theme={null}
sky check k8s
```

Expected output:

```
Checking credentials to enable infra for SkyPilot.
  Kubernetes: enabled [compute]
    Allowed contexts:
    └── t-51326e6b-25ec-42dd-8077-6f3c9b9a34c6-admin: enabled.

🎉 Enabled infra 🎉
  Kubernetes [compute]
```

4. **Check available GPUs:**

```bash theme={null}
sky show-gpus --infra k8s
```

### Example: Launch a workload

Create a SkyPilot task file (`task.yaml`):

```yaml theme={null}
resources:
  accelerators: H100:8
  cloud: kubernetes

setup: |
  pip install torch transformers

run: |
  python train.py
```

Launch the task:

```bash theme={null}
sky launch -c my-job task.yaml
```

### Example: Fine-tune GPT OSS

Download the [gpt-oss-20b.yaml](https://github.com/skypilot-org/skypilot/tree/master/llm/gpt-oss-finetuning#lora-finetuning) configuration.

Launch fine-tuning:

```bash theme={null}
sky launch -c gpt-together gpt-oss-20b.yaml
```

### Benefits

* **Orchestration:** Abstract away Kubernetes complexity.
* **Multi-cloud support:** Same workflow across different clouds.
* **Cost optimization:** Auto-select the cheapest available resources.
* **Job management:** Monitor and cancel jobs.

## Automation patterns

### CI/CD integration

**GitHub Actions example:**

```yaml theme={null}
name: Train Model

on: push

jobs:
  train:
    runs-on: ubuntu-latest
    env:
      TOGETHER_API_KEY: ${{ secrets.TOGETHER_API_KEY }}
    steps:
      - uses: actions/checkout@v3

      - name: Install the Together CLI
        run: uv tool install "together[cli]"

      - name: Create GPU Cluster
        run: |
          tg beta clusters create \
            --name training-${{ github.sha }} \
            --num-gpus 8 \
            --billing-type ON_DEMAND \
            --gpu-type H100_SXM \
            --region us-central-8 \
            --cluster-type KUBERNETES \
            --non-interactive

      - name: Run Training
        run: |
          # Submit training job to cluster
          kubectl apply -f training-job.yaml

      - name: Cleanup
        if: always()
        run: |
          tg beta clusters delete [CLUSTER_ID]
```

### Scheduled jobs

**Cron-based cluster creation:**

```bash theme={null}
# Create cluster daily at 6 AM for batch processing
0 6 * * * tg beta clusters create \
  --name daily-batch \
  --num-gpus 16 \
  --billing-type ON_DEMAND \
  --gpu-type H100_SXM \
  --region us-central-8 \
  --cluster-type KUBERNETES \
  --non-interactive
```

### Auto-scaling scripts

Scale a cluster up or down based on demand with the Together CLI:

```bash theme={null}
# Scale based on job queue length
if [ "$JOB_QUEUE_LENGTH" -gt 100 ]; then
  tg beta clusters update [CLUSTER_ID] --num-gpus 16
else
  tg beta clusters update [CLUSTER_ID] --num-gpus 8
fi
```

## Best practices

### API usage

* **Use environment variables** for API keys (never hardcode).
* **Implement retry logic** for transient failures.
* **Check cluster status** before submitting jobs.
* **Clean up resources** after completion.

### CLI usage

* **Set `TOGETHER_API_KEY`** in your environment so commands authenticate automatically.
* **Use cluster IDs** for cluster references (more reliable than names).
* **Pass `--non-interactive`** (or `--json`) to skip prompts in scripts and CI.
* **Script common operations** for team consistency.

## Troubleshooting

### Authentication issues

* Verify your API key is set: `echo $TOGETHER_API_KEY`
* Confirm the key is valid in your [account settings](https://api.together.ai/settings/projects/~first/api-keys)

### API rate limits

* Implement exponential backoff
* Batch operations when possible
* Contact support for higher limits

## Next steps

<CardGroup cols={2}>
  <Card title="Clusters API reference" icon="code" href="/reference/clusters-create">
    Create, list, and update clusters through the REST API.
  </Card>

  <Card title="Clusters CLI" icon="terminal" href="/reference/cli/clusters">
    Reserve, configure, and manage GPU clusters from the terminal.
  </Card>

  <Card title="Manage clusters" icon="server" href="/docs/gpu-clusters-management">
    Deploy workloads, manage storage, and scale a running cluster.
  </Card>

  <Card title="Billing and pricing" icon="currency-dollar" href="/docs/gpu-clusters-billing">
    Understand reserved, on-demand, and storage charges.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.