Draft

This documentation is in draft and under active review. Figures for the Fall 2026 GPU reservation system are not yet published, and some pages describe behaviour that has not been verified against a primary source. Check with us before relying on anything here.

Running Jobs

Launching containers, the modes a job can run in, watching one, and the Kubernetes underneath.


Documented once, and linked from wherever it is needed. One fact, one anchor: a flag table documented here is not restated on an audience page.

Every page in this directory is an initial draft. Each opens with a note naming what its writer could not settle. Please read those before treating any page as final.

Page Covers
launch.sh Reference Every flag, the three resource tiers, and why requests are half of limits.
Job Modes, Runtime Limits & Configuration Interactive, background and batch; the 6- and 12-hour limits; running several jobs at once; and configuring by environment variable.
Watching a Running Job CPU, memory and the GPU — from the browser, from a shell, and from TensorBoard.
Checkpointing & Logging Long Runs Writing state an interruption cannot corrupt, and the termination warning.
Kubernetes kubectl in your own namespace, services from a manifest, and the events a session emits.

← Documentation index