> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Supported models

> Every base model available for fine-tuning, with context length and batch size limits.

The tables below list every model available through the fine-tuning API. Context lengths are the maximum for that model in SFT and DPO modes. Batch sizes refer to packed batches for text formats. See [data preparation](/docs/fine-tuning/data-preparation) for details on packing.

<Warning>
  Some models can be fine-tuned but cannot be deployed as dedicated endpoints. To verify deployability before training, confirm the base model appears in the [supported models](/docs/dedicated-endpoints/models) list for dedicated model inference (or run `tg beta models configs <BASE_MODEL>`). If it isn't listed there, the fine-tune can't be hosted on a dedicated endpoint.
</Warning>

[Fill out this form](https://www.together.ai/forms/model-requests) to request a model that isn't in the list.

<Columns cols={4}>
  <Card title="LoRA fine-tuning" icon="adjustments-horizontal" horizontal href="#lora-fine-tuning" />

  <Card title="Full fine-tuning" icon="stack-2" horizontal href="#full-fine-tuning" />

  <Card title="Vision-language" icon="eye" horizontal href="#vision-language-models" />

  <Card title="LoRA target modules" icon="target" horizontal href="#lora-target-modules" />
</Columns>

## LoRA fine-tuning

| Organization | Model | API ID | Context (SFT) | Context (DPO) | Max batch (SFT) | Max batch (DPO) | Min batch | Grad accum | Default LoRA rank | Max LoRA rank |
| - | - | - | - | - | - | - | - | - | - | - |
| DeepSeek | DeepSeek V4 Flash 0731 | `deepseek-ai/DeepSeek-V4-Flash-0731` | 131072 | 32768 | 1 | 1 | 1 | 8 | 64 | 128 |
| DeepSeek | DeepSeek V3.1 | `deepseek-ai/DeepSeek-V3.1` | 65536 | 32768 | 2 | 2 | 2 | 8 | 16 | 16 |
| DeepSeek | DeepSeek V4 Flash | `deepseek-ai/DeepSeek-V4-Flash` | 131072 | 32768 | 1 | 1 | 1 | 8 | 64 | 128 |
| Z.ai | GLM 5.3 | `zai-org/GLM-5.3` | 50688 | 25344 | 1 | 1 | 1 | 1 | 16 | 16 |
| Z.ai | GLM 5.2 | `zai-org/GLM-5.2` | 50688 | 25344 | 1 | 1 | 1 | 1 | 16 | 16 |
| Z.ai | GLM 5.1 | `zai-org/GLM-5.1` | 50688 | 25344 | 1 | 1 | 1 | 1 | 16 | 16 |
| Qwen | Qwen3.8 27B | `Qwen/Qwen3.8-27B` | 131072 | 65536 | 2 | 2 | 2 | 4 | 64 | 128 |
| Qwen | Qwen3.6 35B A3B | `Qwen/Qwen3.6-35B-A3B` | 65536 | 32768 | 8 | 8 | 8 | 1 | 64 | 128 |
| Qwen | Qwen3.5 9B | `Qwen/Qwen3.5-9B` | 65536 | 49152 | 8 | 8 | 8 | 1 | 64 | 128 |
| Qwen | Qwen3.5 4B | `Qwen/Qwen3.5-4B` | 131072 | 65536 | 8 | 8 | 8 | 1 | 64 | 128 |
| Qwen | Qwen3.6 27B | `Qwen/Qwen3.6-27B` | 131072 | 65536 | 2 | 2 | 2 | 4 | 64 | 128 |
| Qwen | Qwen3.5 397B A17B | `Qwen/Qwen3.5-397B-A17B` | 32768 | 16384 | 16 | 16 | 16 | 1 | 64 | 128 |
| Qwen | Qwen3.5 122B A10B | `Qwen/Qwen3.5-122B-A10B` | 65536 | 32768 | 16 | 16 | 16 | 1 | 64 | 128 |
| Qwen | Qwen3.5 35B A3B | `Qwen/Qwen3.5-35B-A3B` | 65536 | 32768 | 8 | 8 | 8 | 1 | 64 | 128 |
| Qwen | Qwen3.5 27B | `Qwen/Qwen3.5-27B` | 131072 | 65536 | 2 | 2 | 2 | 4 | 64 | 128 |
| Moonshot AI | Kimi K2.7 Code | `moonshotai/Kimi-K2.7-Code` | 32768 | 16384 | 4 | 4 | 4 | 8 | 16 | 16 |
| Moonshot AI | Kimi K2.6 | `moonshotai/Kimi-K2.6` | 32768 | 16384 | 4 | 4 | 4 | 8 | 16 | 16 |
| Google | Gemma 4 31B IT | `google/gemma-4-31B-it` | 49152 | 24576 | 4 | 4 | 4 | 2 | 64 | 128 |
| Google | Gemma 4 31B IT VLM | `google/gemma-4-31B-it-VLM` | 24576 | 12288 | 8 | 8 | 8 | 1 | 64 | 128 |
| Google | Gemma 4 26B A4B IT | `google/gemma-4-26B-A4B-it` | 49152 | 24576 | 4 | 4 | 4 | 2 | 64 | 128 |
| OpenAI | GPT-OSS 20B | `openai/gpt-oss-20b` | 131072 | 65536 | 1 | 1 | 1 | 8 | 64 | 128 |
| OpenAI | GPT-OSS 120B | `openai/gpt-oss-120b` | 65536 | 32768 | 2 | 2 | 2 | 8 | 64 | 128 |
| Meta | Llama 4 Scout 17B 16E Instruct | `meta-llama/Llama-4-Scout-17B-16E-Instruct` | 65536 | 12288 | 8 | 8 | 8 | 1 | 64 | 128 |
| Meta | Llama 4 Scout 17B 16E Instruct VLM | `meta-llama/Llama-4-Scout-17B-16E-Instruct-VLM` | 32768 | 32768 | 8 | 8 | 8 | 1 | 64 | 128 |
| Meta | Llama 4 Maverick 17B 128E Instruct | `meta-llama/Llama-4-Maverick-17B-128E-Instruct` | 16384 | 24576 | 16 | 16 | 16 | 1 | 64 | 128 |
| Meta | Llama 4 Maverick 17B 128E Instruct VLM | `meta-llama/Llama-4-Maverick-17B-128E-Instruct-VLM` | 16384 | 16384 | 16 | 16 | 16 | 1 | 64 | 128 |
| Meta | Llama 3.3 70B Instruct Reference | `meta-llama/Llama-3.3-70B-Instruct-Reference` | 24576 | 12288 | 8 | 8 | 8 | 1 | 64 | 128 |
| Meta | Meta Llama 3.1 8B Instruct Reference | `meta-llama/Meta-Llama-3.1-8B-Instruct-Reference` | 131072 | 65536 | 8 | 8 | 8 | 1 | 64 | 128 |
| Mistral | Mixtral 8x7B Instruct v0.1 | `mistralai/Mixtral-8x7B-Instruct-v0.1` | 32768 | 16384 | 8 | 8 | 8 | 1 | 64 | 128 |

## Full fine-tuning

| Organization | Model | API ID | Context (SFT) | Context (DPO) | Max batch (SFT) | Max batch (DPO) | Min batch |
| - | - | - | - | - | - | - | - |
| Qwen | Qwen3.8 27B | `Qwen/Qwen3.8-27B` | 131072 | 65536 | 16 | 8 | 4 |
| Qwen | Qwen3.5 9B | `Qwen/Qwen3.5-9B` | 65536 | 49152 | 8 | 8 | 8 |
| Qwen | Qwen3.5 4B | `Qwen/Qwen3.5-4B` | 131072 | 65536 | 8 | 8 | 8 |
| Qwen | Qwen3.6 27B | `Qwen/Qwen3.6-27B` | 131072 | 65536 | 16 | 8 | 4 |
| Qwen | Qwen3.5 27B | `Qwen/Qwen3.5-27B` | 131072 | 65536 | 16 | 8 | 4 |
| Google | Gemma 4 31B IT | `google/gemma-4-31B-it` | 49152 | 24576 | 8 | 8 | 8 |
| Google | Gemma 4 31B IT VLM | `google/gemma-4-31B-it-VLM` | 24576 | 12288 | 16 | 16 | 16 |
| Meta | Llama 3.3 70B Instruct Reference | `meta-llama/Llama-3.3-70B-Instruct-Reference` | 24576 | 12288 | 32 | 32 | 32 |
| Meta | Meta Llama 3.1 8B Instruct Reference | `meta-llama/Meta-Llama-3.1-8B-Instruct-Reference` | 131072 | 65536 | 8 | 8 | 8 |
| Mistral | Mixtral 8x7B Instruct v0.1 | `mistralai/Mixtral-8x7B-Instruct-v0.1` | 32768 | 16384 | 16 | 16 | 16 |

## Vision-language models

For the list of models that support vision-language fine-tuning on image and text data, along with the dataset schema and the `train_vision` parameter, see [vision fine-tuning](/docs/fine-tuning/vision).

## LoRA target modules

See [LoRA vs. full fine-tuning](/docs/fine-tuning/lora-vs-full#default-target-modules) for the default target modules per model. Pass `lora_trainable_modules="all-linear"` to train every linear layer.

## Model limits from the CLI

To get more detailed fine-tuning constraints for a specific model, including learning-rate bounds, batch-size limits, and the LoRA rank and target modules, run [`tg fine-tuning model-limits <model>`](/reference/cli/finetune#model-limits).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.