CosmicAC Logo
vLLM

Create a vLLM Managed Inference Job in the web interface

Serve a language model behind an OpenAI-compatible endpoint by creating a vLLM Managed Inference Job in the CosmicAC web interface.

Create a vLLM Managed Inference Job to serve a language model behind an OpenAI-compatible chat endpoint. The form has six sections, and you click Continue to move from one section to the next. For a description of each field in the form, see vLLM Managed Inference Job configuration.

Changing some of the prefilled serving values can cause the deployment to fail. These values come from the model's recommended configuration, which holds the settings that the model needs to deploy.

CosmicAC provides recommended configurations for a set of models. For the list, see Recommended configuration values. For other models, register a vLLM model to create a recommended configuration from its vLLM recipe, or add a recommended configuration with your own values.

Prerequisites

Before you start, make sure that you have the following.

  • A running CosmicAC deployment. See Set up CosmicAC.
  • Access to the CosmicAC web interface.

Steps

Open the new job form

In the left navigation, click Jobs, and then click New Job.

Select the job type

In the What kind of job? section, select Managed Inference, and then click Continue.

Enter the basics

In the Basics section, enter a Job name. Use lowercase letters and hyphens.

In Tags, add at least one tag. To add a tag, type it, and then press Enter or type a comma.

Click Continue.

Select a model

In the Model to serve section, select a vLLM Model. CosmicAC prefills the Serving configuration from the model's recommended configuration, and labels it Prefilled from with the model name. The configuration has the following fields.

  • Runtime image (CUDA): the vLLM serving image.
  • Data type: the numeric precision the model runs at.
  • Quantisation: how to compress the model weights.
  • GPU memory utilisation: the fraction of GPU memory to use, with 85%, 90%, and 95% presets.
  • Max model length: the maximum context length, with 8k, 16k, 32k, and 128k presets.
  • Max concurrent sequences: the maximum requests handled at once, with 64, 128, 256, and 512 presets.
  • Reasoning parser: the parser that separates thinking tokens from the final response.
  • Video & image input: whether the model accepts multimodal input, for vision-language models only.

You can edit every field except the runtime image, which comes from the model's recommended configuration. To serve a different image, update the recommended model configuration.

Name the endpoint

Under Endpoint, enter an Endpoint name. The name must be unique across the active Managed Inference Jobs in the deployment. A failed job keeps its name until you delete the job.

Below the name, Will be reachable at shows the endpoint URL.

Set the root disk and environment variables

Under Instance resources, set the Root disk (GB). Select a preset or enter a value. Increase it for a large checkpoint that exceeds the cluster default.

Under Environment variables, review the prefilled variables, and then edit them if needed. To add another, click Add variable, and then enter a Name and Value. To remove one, click the X beside it. For the names CosmicAC reads, see vLLM serving options.

Require an API key

Under API key required, select Require Authorization header, and then click Continue. To create an API key that authenticates requests to the endpoint, see Create an API key in the web interface.

Select the hardware

In the Hardware section, select a Location first. The GPU list stays empty until you select one.

Select a GPU from the ones available in that location. Each card shows the GPU's VRAM, CPU, and RAM. Set the GPU count, and then set the CUDA / driver. Below the count, Recommended for this model shows the model's GPU count. On one node is the most free GPUs on any single node, and In this region is the total free across the location. A GPU count of 16 spreads one replica over two nodes. See Multi-node replicas.

Set Replicas to 1, 2, or 4. CosmicAC doesn't autoscale replicas, and a job keeps the count it starts with. Total shows the GPUs the job claims, the GPU count multiplied by the replica count. For how requests reach the replicas, see Replicas.

Click Continue.

Select the notification events

In the Notifications section, turn on each job lifecycle event you want this job to report. CosmicAC turns all four on by default.

  • job.failed: the job transitions to Failed, and the event carries the failure reason.
  • job.degraded: healthy replicas drop below the count you set, and the endpoint stays live.
  • job.recovered: the job returns to Active from Degraded or Failed.
  • job.restart_storm: any replica restarts three times within 10 minutes.

These preferences cover this job alone. An event you turn on here reaches your webhook only if it's also turned on in Settings > Notifications, which also controls the model health and usage window events for the whole deployment. See Set up webhook notifications.

Click Continue.

Review and create the job

In the Review & launch section, check that it reports Ready to create, and then click Create job. If it reports issues instead, click Edit on the row that names the problem, fix it, and then return to this section.

If your configuration departs from the model's recommended parameters, the form shows a Job creation warning that names the values that differ, and Create job stays disabled until you select Create anyway with this configuration. A job created that way can still fail to start.

The model's recommended model configuration sets which parameters warn and in which direction. See Set the job creation warnings.

Open the endpoint

In the Provisioning status dialog, click Open job. When the job is running, click the Endpoint tab to view the endpoint URL. To send it a request, see Connect to a vLLM Managed Inference endpoint.

Help and troubleshooting

Job stuck in Creating or Starting

If a job stays in Creating or Starting, check the status of its KubeVirt virtual machine instance (VMI).

  1. Find the job's container ID. Click Jobs in the left navigation, and then click the job. The Containers section lists the Container ID for each container.

  2. Find the VMI for the container.

    CosmicAC creates one VMI for each container and names it <container-id>-n0. A multi-node job has one VMI per node.

    From a machine with kubectl access to your Kubernetes cluster, run the following command.

    kubectl get vmi -n <namespace>

    Replace <namespace> with the namespace configured in K8S_NAMESPACE.

  3. Check the VMI status.

Next steps

On this page