Command-line tool for distributed LLM inference benchmarks on SLURM clusters using SGLang, vLLM, TensorRT LLM, TileRT and AMD's ATOM. Replace complex shell scripts and 50+ CLI flags with a declarative schema: 2 YAML recipe: engine: names the engine, roles: describes each worker role, and services: covers everything launched next to the workers.
Documentation: https://nvidia.github.io/srt-slurm/ (source in docs/; agents: llms.txt)
git clone https://github.com/NVIDIA/srt-slurm.git
cd srt-slurm
uv sync --no-dev
# One-time setup: downloads etcd, NATS, uv and the tachometer binaries for the
# compute nodes' architecture and writes srtslurm.yaml (account, partition, GPUs per node)
make setup ARCH=aarch64 # or ARCH=x86_64
uv run srtctl dry-run -f examples/sglang/dynamo-agg.yaml # render without submitting
uv run srtctl apply -f examples/sglang/dynamo-agg.yaml # submitInstallation covers cluster-specific setup; examples/ has runnable recipes, one per frontend and topology.
srtctl ships a skill that teaches Claude Code, Codex or Cursor how to set up a checkout, write srtslurm.yaml, author and validate recipes, submit them, and read the results. Install it into the checkout and hand the agent the checkout:
uv run srtctl skill --target claude # .claude/skills/srtctl/SKILL.md
uv run srtctl skill --target codex # .codex/skills/srtctl/SKILL.md
uv run srtctl skill --target cursor # .cursor/rules/srtctl.mdcThe srtctl-mcp server, the recipe JSON Schema for editors, and llms.txt are described in the documentation home.
- Write a recipe: Recipe Guide, then the topic pages it links (engines, topology, frontends, benchmarks, services, ...)
- Every field, type, default and allowed value: Schema Reference (generated from the code)
- Every command and flag: CLI Guide and CLI Reference (generated)
- Run and debug: Monitoring, SLURM FAQ
- Analyze: Profiling, DSight
- Old recipes: Legacy (v1) layout;
srtctl migraterewrites them