> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Structured extraction with vision

> Combine image input with a JSON schema to extract typed data from screenshots, documents, and photos, or run OCR to pull text out of images.

You can combine vision input with structured outputs to extract typed data from an image. Pass an `image_url` content block and a `response_format` with a JSON schema; the model returns JSON that conforms to the schema.

For example, you could extract a project name and a column count from a screenshot of a Trello board:

<CodeGroup>
  ```python Python theme={null}
  import json
  from together import Together
  from pydantic import BaseModel, Field

  client = Together()


  class ImageDescription(BaseModel):
      project_name: str = Field(
          description="The name of the project shown in the image"
      )
      col_num: int = Field(description="The number of columns in the board")


  image_url = "https://napkinsdev.s3.us-east-1.amazonaws.com/next-s3-uploads/d96a3145-472d-423a-8b79-bca3ad7978dd/trello-board.png"

  extract = client.chat.completions.create(
      messages=[
          {
              "role": "user",
              "content": [
                  {
                      "type": "text",
                      "text": "Extract a JSON object from the image.",
                  },
                  {"type": "image_url", "image_url": {"url": image_url}},
              ],
          }
      ],
      model="moonshotai/Kimi-K3",
      reasoning={"enabled": False},
      response_format={
          "type": "json_schema",
          "json_schema": {
              "name": "image_description",
              "schema": ImageDescription.model_json_schema(),
          },
      },
  )

  print(json.dumps(json.loads(extract.choices[0].message.content), indent=2))
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";
  import { z } from "zod";

  const together = new Together();

  const schema = z.object({
    projectName: z.string().describe("The name of the project shown in the image"),
    columnCount: z.number().describe("The number of columns in the board"),
  });

  const imageUrl =
    "https://napkinsdev.s3.us-east-1.amazonaws.com/next-s3-uploads/d96a3145-472d-423a-8b79-bca3ad7978dd/trello-board.png";

  const extract = await together.chat.completions.create({
    messages: [
      {
        role: "user",
        content: [
          { type: "text", text: "Extract a JSON object from the image." },
          { type: "image_url", image_url: { url: imageUrl } },
        ],
      },
    ],
    model: "moonshotai/Kimi-K3",
    reasoning: { enabled: false },
    response_format: {
      type: "json_schema",
      json_schema: {
        name: "image_description",
        schema: z.toJSONSchema(schema),
      },
    },
  });

  console.log(JSON.parse(extract.choices[0].message.content));
  ```
</CodeGroup>

Example output:

```json JSON theme={null}
{
  "projectName": "Project A",
  "columnCount": 4
}
```

## OCR: extract text from documents

Optical character recognition (OCR) is the same pattern applied to documents. Vision models read the text in an image while understanding its context and structure, which makes them useful for processing receipts, invoices, and other structured documents. See [llamaOCR.com](https://llamaocr.com/) for a working example.

To extract everything a document contains as plain text, send the image without a schema:

<CodeGroup>
  ```python Python theme={null}
  from together import Together

  client = Together()

  prompt = "You are an expert at extracting information from receipts. Extract all the content from the receipt."

  imageUrl = "https://napkinsdev.s3.us-east-1.amazonaws.com/next-s3-uploads/1627e746-7eda-46d3-8d08-8c8eec0d6c9c/nobu.jpg?x-id=PutObject"


  stream = client.chat.completions.create(
      model="zai-org/GLM-5.3-Flash",
      messages=[
          {
              "role": "user",
              "content": [
                  {"type": "text", "text": prompt},
                  {
                      "type": "image_url",
                      "image_url": {
                          "url": imageUrl,
                      },
                  },
              ],
          }
      ],
      reasoning={"enabled": False},
      stream=True,
  )

  for chunk in stream:
      print(
          chunk.choices[0].delta.content or "" if chunk.choices else "",
          end="",
          flush=True,
      )
  ```

  ```typescript TypeScript theme={null}
  import Together from "together-ai";

  const together = new Together();

  async function main() {
    const billUrl =
      "https://napkinsdev.s3.us-east-1.amazonaws.com/next-s3-uploads/1627e746-7eda-46d3-8d08-8c8eec0d6c9c/nobu.jpg?x-id=PutObject";

    const response = await together.chat.completions.create({
      model: "zai-org/GLM-5.3-Flash",
      messages: [
        {
          role: "system",
          content:
            "You are an expert at extracting information from receipts. Extract all the content from the receipt.",
        },
        {
          role: "user",
          content: [
            { type: "text", text: "Extract receipt information" },
            { type: "image_url", image_url: { url: billUrl } },
          ],
        },
      ],
      reasoning: { enabled: false },
    });

    if (response?.choices?.[0]?.message?.content) {
      console.log(response.choices[0].message.content);
      return (response.choices[0].message.content);
    }

    throw new Error("Failed to extract receipt information");
  }

  main();
  ```
</CodeGroup>

The model returns the receipt's contents as formatted text:

```text Text theme={null}
**Restaurant Information:**
- Name: Noby
- Location: Los Angeles
- Address: 903 North La Cienega
- Phone Number: 310-657-5111

**Receipt Details:**
- Date: 04/16/2011
- Time: 9:19 PM
- Server: Daniel
- Guest Count: 15

**Ordered Items:**
1. **Pina Martini** - $14.00
2. **Jasmine Calpurnina** - $14.00
...

**Billing Information:**
- **Subtotal** - $3830.00
- **Tax** - $766.00
- **Total** - $5043.72
- **Balance Due** - $5043.72
```

### Extract structured data from a receipt

For receipt processing (as seen on [usebillsplit.com](https://www.usebillsplit.com/)), combine OCR with a schema to get back typed JSON instead of free text:

<CodeGroup>
  ```python Python theme={null}
  import json
  import together
  from pydantic import BaseModel, Field
  from typing import Optional

  client = together.Together()


  class Receipt(BaseModel):
      businessName: Optional[str] = Field(
          None, description="Name of the business on the receipt"
      )
      date: Optional[str] = Field(
          None, description="Date when the receipt was created"
      )
      total: Optional[float] = Field(
          None, description="Total amount on the receipt"
      )
      tax: Optional[float] = Field(None, description="Tax amount on the receipt")


  def extract_receipt_info(image_url: str) -> dict:
      response = client.chat.completions.create(
          model="zai-org/GLM-5.3-Flash",
          messages=[
              {
                  "role": "system",
                  "content": "You are an expert at extracting information from receipts. Extract the relevant information and format it as JSON.",
              },
              {
                  "role": "user",
                  "content": [
                      {"type": "text", "text": "Extract receipt information"},
                      {"type": "image_url", "image_url": {"url": image_url}},
                  ],
              },
          ],
          reasoning={"enabled": False},
          response_format={
              "type": "json_schema",
              "json_schema": {
                  "name": "receipt",
                  "schema": Receipt.model_json_schema(),
              },
          },
      )

      if response and response.choices and response.choices[0].message.content:
          try:
              return json.loads(response.choices[0].message.content)
          except json.JSONDecodeError:
              return {"error": "Failed to parse response as JSON"}

      return {"error": "Failed to extract receipt information"}


  receipt_url = "https://napkinsdev.s3.us-east-1.amazonaws.com/next-s3-uploads/1627e746-7eda-46d3-8d08-8c8eec0d6c9c/nobu.jpg?x-id=PutObject"
  print(json.dumps(extract_receipt_info(receipt_url), indent=2))
  ```

  ```typescript TypeScript theme={null}
  import { z } from "zod";
  import Together from "together-ai";

  const together = new Together();

  async function main() {
    const billUrl =
      "https://napkinsdev.s3.us-east-1.amazonaws.com/next-s3-uploads/1627e746-7eda-46d3-8d08-8c8eec0d6c9c/nobu.jpg?x-id=PutObject";

    const receiptSchema = z.object({
      businessName: z
        .string()
        .optional()
        .describe("Name of the business on the receipt"),
      date: z.string().optional().describe("Date when the receipt was created"),
      total: z.number().optional().describe("Total amount on the receipt"),
      tax: z.number().optional().describe("Tax amount on the receipt"),
    });

    const response = await together.chat.completions.create({
      model: "zai-org/GLM-5.3-Flash",
      messages: [
        {
          role: "system",
          content:
            "You are an expert at extracting information from receipts. Extract the relevant information and format it as JSON.",
        },
        {
          role: "user",
          content: [
            { type: "text", text: "Extract receipt information" },
            { type: "image_url", image_url: { url: billUrl } },
          ],
        },
      ],
      reasoning: { enabled: false },
      response_format: {
        type: "json_schema",
        json_schema: {
          name: "receipt",
          schema: z.toJSONSchema(receiptSchema),
        },
      },
    });

    if (response?.choices?.[0]?.message?.content) {
      const output = JSON.parse(response.choices[0].message.content);
      console.dir(output);
      return output;
    }

    throw new Error("Failed to extract receipt information");
  }

  main();
  ```
</CodeGroup>

The response contains only the fields the schema asks for:

```json JSON theme={null}
{
  "businessName": "Noby",
  "date": "04/16/2011",
  "total": 5043.72,
  "tax": 766
}
```

## Best practices

1. **Structured data definition:** Define clear schemas for your expected output, so you can validate and process the extracted data.
2. **Model selection:** Choose the appropriate model based on your use case. Try [Together's vision models](/docs/serverless/models#vision-models) to match your workload.
3. **Error handling:** Always implement robust error handling for cases where the OCR might fail or return unexpected results.
4. **Validation:** Implement validation for the extracted data to ensure accuracy and completeness.

For the full structured-outputs reference, see [Structured outputs](/docs/inference/chat/structured-outputs).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.