Skip to main content
The fastest way to serve a model on Anaconda Platform is to deploy it directly from the model catalog. Deployments from the catalog serve the model as-is, with no application logic around it, which makes them well suited for evaluating and comparing models before you commit to one. The platform packages the model into a long-running service, schedules it on an inference compute pool, and gives you an authenticated API endpoint and a built-in chat interface to try it out, with no code or configuration files required. To build a service around a model instead, see Writing your first deployment. For background on how deployments work, see What is a deployment? For more inference patterns, including CLI-based deployment for vLLM (safetensors) models and the @vllm and @llamacpp decorators, see the inference examples repository:
outerbounds/inference-examples
Loading repository data...

Deploying a model

  1. In the left-hand navigation, select Resources, then select Models.
  2. Select a model from the list to open its details page.
  3. Click Deploy Model.
  4. In the Deploy model dialog, choose the Project context for the deployment. Deployments created from the model catalog are scoped to the perimeter: the model catalog operates at perimeter level, so the context picker shows the perimeter while you are on the Models page, and the deployment appears in the perimeter’s Deployments list rather than under a specific project branch.
  5. Under Resources, set the CPU, memory, and disk for the deployment. All three fields start at zero, and you must set them before you can create the deployment. Use the file size and estimated RAM values shown under the selected file as your guide: memory must exceed the estimated RAM, and disk must exceed the file size with room for the serving runtime. The dialog validates your values against the available compute pools as you type. When your resources fit a pool, the dialog shows which pool the deployment will be scheduled on. If no pool can fit your values, choose a smaller model file or reduce the requested resources.
    The Deploy model dialog showing Project, Model, and File fields with resource inputs for CPU, memory, disk, GPU, and shared memory, and a list of compute pools
  6. Click Create.
The platform opens the Deployments view. The deployment registers within a few seconds, then takes several minutes to provision and start.
If the Deployments view shows “Deployment not found” immediately after creating a deployment, refresh the page. The deployment list updates as the platform registers the new deployment.

Accessing the deployment

When the deployment is ready, it appears in the Deployments list with a green status indicator. Select it to open the details view, which shows:
  • The API endpoint URL, labeled “available at”. This is the endpoint your applications call.
  • Additional UI routes. For certain model file types (.gguf models served with llamacpp), this includes a built-in chat interface for trying the model interactively. Safetensors models served with vLLM expose the API endpoint only. Both the endpoint and the chat UI require visitors to be signed in to the platform.
  • The serving image, compute pool, and resources the deployment is running with.
  • Charts for requests per minute and autoscaling activity.
The deployment details view for a deployed model showing the endpoint URL, additional UI routes, serving image, compute pool, and requests per minute chart
The deployment name gets a random suffix (for example, gemma-2-2b-it-dm2kkh) so you can deploy the same model more than once without name conflicts.

Trying the model

The deployment’s API endpoint is the primary interface: send requests to the endpoint URL with your application or tools like curl. For gguf models served with llamacpp, the deployment also includes a built-in chat interface. To open it, click the Additional UI routes link in the deployment details. You will be asked to sign in if you are not already signed in. From the chat UI, you can send messages to the model and inspect its responses without writing any code. Each response shows its token count, latency, and throughput in tokens per second, which gives you a quick read on how the deployment performs under your sizing choices.
The chat interface is available only for .gguf models served with llamacpp. Safetensors models served with vLLM expose the API endpoint only.
The chat interface for a deployed model showing a message input field with the model name

Deleting the deployment

To tear down a deployment you no longer need, use the CLI:
Deletion is asynchronous. The deployment might show a terminating state briefly before it disappears from the list.

Next steps