CosmicAC Logo
Architecture

GPU Container architecture

How a GPU Container Job runs on your cluster, from the job starting to the shell you open on it.

A GPU Container Job runs your workload inside a KubeVirt virtual machine instance (VMI) on a single node. The VMI claims one or more whole GPUs through passthrough. You open an interactive shell into it and use it like a remote machine.

The following diagram shows how a job request reaches your cluster, and the separate path that your shell takes into the VMI. A node can hold many pods, each running a different job.

How a job starts

When you submit a job from cosmicac-ui or cosmicac-cli, cosmicac-app-node authenticates the request and forwards it to cosmicac-wrk-ork. The orchestrator allocates the GPUs that the job needs on one node that the job's team can use, and then hands the job to cosmicac-wrk-server-k8s-nvidia. That worker creates the job's Kubernetes resources through your cluster's API server, and Kubernetes creates a pod that holds the VMI. cosmicac-wrk-agent-instance runs inside that VMI.

How a shell connects

After the VMI is running, cosmicac-cli asks cosmicac-app-node for the container's shell key. cosmicac-app-node checks that you can access the job, and then returns the key.

cosmicac-cli uses the key to connect directly to cosmicac-wrk-agent-instance over hyperswarm-ssh, which is part of the Holepunch peer-to-peer stack. cosmicac-app-node isn't part of that connection, so the interactive session doesn't depend on the control path that submitted the job.

Next steps

On this page