> ## Documentation Index
> Fetch the complete documentation index at: https://docs.together.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Get a deployment

> Retrieves a deployment's desired configuration, placement, runtime information, and current provisioning status.



## OpenAPI

````yaml openapi.yaml GET /projects/{projectId}/endpoints/{endpointId}/deployments/{id}
openapi: 3.1.0
info:
  title: Together APIs
  description: The Together REST API. See https://docs.together.ai for more details.
  version: 2.0.0
  termsOfService: https://www.together.ai/terms-of-service
  contact:
    name: Together Support
    url: https://www.together.ai/contact
  license:
    name: MIT
    url: https://github.com/togethercomputer/openapi/blob/main/LICENSE
servers:
  - url: https://api.together.ai/v1
    description: Default environment for APIs
  - url: https://api-inference.together.ai/v2
    description: Optimized environment for inference
security:
  - bearerAuth: []
paths:
  /projects/{projectId}/endpoints/{endpointId}/deployments/{id}:
    get:
      tags:
        - DeploymentService
      summary: Get a deployment
      description: >-
        Retrieves a deployment's desired configuration, placement, runtime
        information, and current provisioning status.
      operationId: DeploymentService_GetDeployment
      parameters:
        - name: projectId
          in: path
          required: true
          schema:
            description: Project identifier.
            type: string
        - name: endpointId
          in: path
          required: true
          schema:
            description: Endpoint identifier.
            type: string
        - name: id
          in: path
          required: true
          schema:
            description: Deployment identifier.
            type: string
      responses:
        '200':
          description: OK
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DE.Deployment'
        default:
          description: Default error response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorData'
      servers:
        - url: https://api.together.ai/v2
components:
  schemas:
    DE.Deployment:
      description: >-
        Serving workload that binds a model and immutable config to an endpoint
        and manages its replicas.
      type: object
      required:
        - id
        - projectId
        - endpointId
        - name
        - createdAt
        - updatedAt
        - modelId
        - modelRevisionId
        - model
        - autoscaling
        - configId
        - config
        - etag
        - hardware
        - trafficMode
        - status
      properties:
        id:
          readOnly: true
          type: string
          description: Unique deployment identifier.
        projectId:
          readOnly: true
          type: string
          description: ID of the project that owns the deployment.
        endpointId:
          type: string
          description: ID of the endpoint that contains the deployment.
        name:
          type: string
          description: >-
            Project- and endpoint-qualified deployment name in the form
            `<project_slug>/<endpoint_name>/<deployment_name>`. Pass it as
            `model` in an inference request to target this deployment directly
            instead of using the endpoint's traffic split.
        createdAt:
          readOnly: true
          type: string
          description: Timestamp when the deployment was created.
          format: date-time
        updatedAt:
          readOnly: true
          type: string
          description: Timestamp when the deployment was last updated.
          format: date-time
        modelId:
          type: string
          description: >-
            Deprecated. Use `model`. Model identifier being served, populated
            during migration.
        modelRevisionId:
          type: string
          description: >-
            Deprecated. Use `model` with a /revisions/{revisionId} segment. Pin
            to a specific model revision.
        model:
          type: string
          description: >-
            Pinned model resource in the form
            `projects/{projectId}/models/{modelId}/revisions/{revisionId}`.
        autoscaling:
          allOf:
            - $ref: '#/components/schemas/DE.AutoscalingResponse'
          description: >-
            Replica bounds, timing windows, and metrics that control horizontal
            scaling.
        configId:
          type: string
          description: >-
            Deprecated. Use `config`. Config revision identifier, populated
            during migration.
        config:
          type: string
          description: >-
            Immutable config revision in the form
            `projects/{projectId}/configs/{configRevisionId}`.
        speculatorId:
          readOnly: true
          type: string
          description: >-
            Deprecated. Use `speculator`. Speculative decoding model identifier
            derived from the deployment config.
        speculatorRevisionId:
          readOnly: true
          type: string
          description: >-
            Deprecated. Use `speculator`. ID of the speculative decoding
            draft-model revision pinned at creation time.
        speculator:
          readOnly: true
          type: string
          description: >-
            Pinned draft-model resource used for speculative decoding, in the
            same form as `model`. Omitted when speculative decoding is disabled.
        estimatedEffectiveTrafficShare:
          readOnly: true
          type: number
          format: double
          description: >-
            Estimated fraction in [0, 1] of endpoint traffic that reaches this
            deployment under the current routing configuration. Absent or
            unrouted deployments are 0.
        maxConcurrentRequestsPerReplica:
          type: string
          description: >-
            Maximum number of inference requests that may be in flight to a
            single replica. If omitted, the platform uses one less than the
            config's per-replica concurrency limit to reserve a health-check
            slot. Values above that maximum are reduced on create and update; 0
            means unlimited when the config limit is 1 or less. Changes take
            effect without restarting replicas.
        inactiveTimeout:
          type: integer
          description: >-
            Minutes without an inference request before the deployment stops
            automatically. Omitted or 0 means automatic stopping is disabled.
        etag:
          type: string
          description: >-
            Opaque version tag for optimistic concurrency control. Supply on
            update/delete to ensure consistent read-modify-write. If not set,
            the write overwrites based on current state.
        hardware:
          readOnly: true
          type: string
          description: >-
            Hardware selected by the deployment config, including GPU type and
            count.
        trafficMode:
          readOnly: true
          enum:
            - TRAFFIC_MODE_LIVE
            - TRAFFIC_MODE_SHADOW
          type: string
          description: >-
            Whether the deployment serves client-visible responses or only
            mirrored shadow traffic.
        runtimeInfo:
          readOnly: true
          allOf:
            - $ref: '#/components/schemas/DE.RuntimeInfo'
          description: >-
            Serving engine and feature support derived from the immutable
            config.
        desiredReplicas:
          readOnly: true
          type: integer
          description: >-
            Number of replicas the autoscaler currently wants across all
            regions. Not settable on

            any request; steer it through `autoscaling.minReplicas` and
            `autoscaling.maxReplicas`.
        status:
          readOnly: true
          allOf:
            - $ref: '#/components/schemas/DE.DeploymentStatus'
          description: Current lifecycle state and observed replica counts.
        placement:
          allOf:
            - $ref: '#/components/schemas/DE.Placement'
          description: Region constraints used to schedule the deployment's replicas.
    ErrorData:
      type: object
      required:
        - error
      properties:
        error:
          type: object
          properties:
            message:
              type: string
              nullable: false
            type:
              type: string
              nullable: false
            param:
              type: string
              nullable: true
              default: null
            code:
              type: string
              nullable: true
              default: null
          required:
            - type
            - message
    DE.AutoscalingResponse:
      description: Resolved autoscaling configuration returned for a deployment.
      allOf:
        - $ref: '#/components/schemas/DE.Autoscaling'
        - type: object
          required:
            - minReplicas
            - maxReplicas
    DE.RuntimeInfo:
      type: object
      properties:
        engineType:
          type: string
          description: Serving engine, such as `vllm`, `trtllm`, or `sglang`.
        engineVersion:
          type: string
          description: Version of the serving engine.
        functionCallingSupported:
          type: boolean
          description: Whether the runtime accepts tool and function-calling requests.
        structuredOutputSupported:
          type: boolean
          description: >-
            Whether the runtime can constrain generation to a structured output
            schema.
      description: Runtime information derived from the deployment's configuration.
    DE.DeploymentStatus:
      type: object
      required:
        - state
        - message
      properties:
        state:
          enum:
            - DEPLOYMENT_STATE_PROVISIONING
            - DEPLOYMENT_STATE_READY
            - DEPLOYMENT_STATE_SCALING
            - DEPLOYMENT_STATE_DEGRADED
            - DEPLOYMENT_STATE_FAILED
            - DEPLOYMENT_STATE_STOPPED
            - DEPLOYMENT_STATE_STOPPING
          type: string
          description: High-level lifecycle state.
        readyReplicas:
          type: integer
          description: Total replicas actively serving traffic across all clusters.
        message:
          type: string
          description: Human-readable explanation of the current state.
        scheduledReplicas:
          type: integer
          description: Replicas the scheduler has placed on clusters.
        details:
          allOf:
            - $ref: '#/components/schemas/DE.StatusDetails'
          description: >-
            Status totals broken down by dimension. Omitted when the breakdown
            cannot be determined.
      description: >-
        Current status of a deployment, derived at read time from internal
        state.
    DE.Placement:
      description: Placement controls where a deployment is scheduled.
      oneOf:
        - type: object
          x-stainless-variantName: Inline
          required:
            - inline
          properties:
            inline:
              description: Inline placement parameters evaluated at deploy time.
              allOf:
                - $ref: '#/components/schemas/DE.InlinePlacement'
        - type: object
          x-stainless-variantName: Profile
          required:
            - profile
          properties:
            profile:
              type: string
              description: UID of a saved placement profile.
    DE.Autoscaling:
      type: object
      properties:
        minReplicas:
          type: integer
          description: >-
            Minimum number of replicas. Omit on update to preserve the current
            value. Set both `minReplicas` and `maxReplicas` to `0` to stop the
            deployment.
        maxReplicas:
          type: integer
          description: >-
            Maximum number of replicas. Defaults to `minReplicas`; omitting it
            on update preserves the current value.
        scaleDownWindow:
          pattern: ^-?(?:0|[1-9][0-9]{0,11})(?:\.[0-9]{1,9})?s$
          type: string
          description: >-
            Time a lower replica recommendation must remain stable before
            scaling down. Defaults to `5m`.
        scaleUpWindow:
          pattern: ^-?(?:0|[1-9][0-9]{0,11})(?:\.[0-9]{1,9})?s$
          type: string
          description: Stabilization window before scaling up.
        scaleToZeroWindow:
          pattern: ^-?(?:0|[1-9][0-9]{0,11})(?:\.[0-9]{1,9})?s$
          type: string
          description: >-
            Idle period after which the deployment automatically stops and
            releases its replicas.
        scalingMetrics:
          type: array
          items:
            $ref: '#/components/schemas/DE.ScalingMetric'
          description: >-
            Metrics and targets that drive replica recommendations. When
            omitted, the platform uses concurrent in-flight requests per
            replica.
        scaleUp:
          allOf:
            - $ref: '#/components/schemas/DE.ScalingRules'
          description: >-
            Rate limits applied when scaling up. Stabilization remains
            controlled by `scaleUpWindow`.

            Omitted fields are preserved on update; a non-empty policy list
            replaces the previous list.

            To clear policies or reset the selector, explicitly mask that leaf
            field. Empty lists in

            parent-only updates are treated as omitted.
        scaleDown:
          allOf:
            - $ref: '#/components/schemas/DE.ScalingRules'
          description: >-
            Rate limits applied when scaling down. Stabilization remains
            controlled by `scaleDownWindow`.

            Omitted fields are preserved on update; a non-empty policy list
            replaces the previous list.

            To clear policies or reset the selector, explicitly mask that leaf
            field. Empty lists in

            parent-only updates are treated as omitted.
      description: Autoscaling configuration for a deployment.
    DE.StatusDetails:
      type: object
      required:
        - region
      properties:
        region:
          type: array
          items:
            $ref: '#/components/schemas/DE.RegionStatus'
          description: >-
            Regions where the deployment is actually scheduled or serving
            replicas, sorted by region.
      description: Deployment status broken down by each supported dimension.
    DE.InlinePlacement:
      type: object
      properties:
        regions:
          type: array
          items:
            type: string
          description: >-
            Regions where the deployment is allowed to run. Multiple regions
            allow best-effort replica spreading.
        constraint:
          enum:
            - ENFORCEMENT_REQUIRED
            - ENFORCEMENT_PREFERRED
          type: string
          description: How strictly the regions list is enforced.
        compliancePolicy:
          description: Compliance regimes required for clusters that run the deployment.
          allOf:
            - $ref: '#/components/schemas/DE.CompliancePolicy'
      description: >-
        Inline placement parameters expanded into scheduling rules by the
        server.
    DE.ScalingMetric:
      type: object
      properties:
        name:
          enum:
            - active_sessions
            - cache_hit_rate
            - decoding_speed
            - e2e_latency
            - gpu_utilization
            - inflight_requests
            - throughput_per_replica
            - token_utilization
            - ttft
          type: string
          description: Autoscaling metric name from the server allowlist.
        type:
          enum:
            - METRIC_TARGET_TYPE_VALUE
            - METRIC_TARGET_TYPE_UTILIZATION
            - METRIC_TARGET_TYPE_AVERAGE_VALUE
          type: string
          description: >-
            Whether `target` is an absolute value, a utilization percentage, or
            a per-replica average.
        target:
          type: number
          description: >-
            Target interpreted according to `type`. Utilization uses a
            percentage from 0 to 100, value uses an absolute measurement, and
            average value uses a per-replica measurement.
        percentile:
          type: string
          description: >-
            Percentile to evaluate for latency-based metrics: `p50`, `p90`,
            `p95`, or `p99`.
      description: Metric and target used by the autoscaler to recommend a replica count.
      required:
        - name
        - type
        - target
    DE.ScalingRules:
      type: object
      properties:
        policies:
          type: array
          items:
            $ref: '#/components/schemas/DE.ScalingPolicy'
          description: >-
            Non-empty lists replace the existing policies. To clear policies,
            include

            `autoscaling.scaleDown.policies` or `autoscaling.scaleUp.policies`
            in the update mask

            and supply an empty scaling rules object or `policies: []`.
        selectPolicy:
          enum:
            - SCALING_POLICY_SELECT_MAX
            - SCALING_POLICY_SELECT_MIN
            - SCALING_POLICY_SELECT_DISABLED
          type: string
          description: >-
            `SCALING_POLICY_SELECT_MIN` chooses the policy allowing the smallest
            replica change;

            `SCALING_POLICY_SELECT_MAX` chooses the largest. These are caps, not
            guaranteed changes.

            `SCALING_POLICY_SELECT_DISABLED` holds this direction steady while
            replica bounds still apply.

            Omitted preserves the existing selector on update. When no selector
            is configured,

            authored policies use MAX; with no policies configured, the platform
            defaults apply.

            To reset the selector, include `autoscaling.scaleDown.selectPolicy`
            or

            `autoscaling.scaleUp.selectPolicy` in the update mask and omit
            `selectPolicy`.

            Clear both policies and `selectPolicy` to restore inherited
            defaults.
      description: Rate limits applied after stabilization and before replica bounds.
      example:
        policies:
          - type: SCALING_POLICY_TYPE_PERCENT
            value: 25
            periodSeconds: 60
          - type: SCALING_POLICY_TYPE_PODS
            value: 10
            periodSeconds: 60
        selectPolicy: SCALING_POLICY_SELECT_MIN
    DE.RegionStatus:
      type: object
      required:
        - region
        - scheduledReplicas
        - readyReplicas
      properties:
        region:
          type: string
          description: >-
            Region name using the same vocabulary accepted by inline placement
            regions.
        scheduledReplicas:
          type: integer
          description: Replicas the scheduler has placed in this region.
        readyReplicas:
          type: integer
          description: Replicas serving traffic in this region.
      description: Realized scheduled and ready replica counts for one deployment region.
    DE.CompliancePolicy:
      type: object
      properties:
        hipaa:
          type: boolean
          description: Restrict placement to HIPAA-attested clusters.
      description: Compliance regimes required by a deployment placement policy.
    DE.ScalingPolicy:
      type: object
      description: Replica rate-limit policy applied over a trailing window.
      required:
        - type
        - value
        - periodSeconds
      properties:
        type:
          enum:
            - SCALING_POLICY_TYPE_PODS
            - SCALING_POLICY_TYPE_PERCENT
          type: string
          description: >-
            Whether `value` is a replica count or a percentage of the replica
            count at the start

            of the trailing period. Scaling events within that period count
            against the allowance;

            percentages are rounded to whole replicas.
        value:
          type: integer
          minimum: 1
          maximum: 2147483647
          description: Positive replica count or percentage used as the rate-limit amount.
        periodSeconds:
          type: integer
          minimum: 1
          maximum: 1800
          description: Trailing rate-limit window in seconds, from 1 to 1800.
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      x-bearer-format: bearer
      x-default: default

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.