via Remote Rocketship
$120K - 160K a year
Build and operate automation and tools for large-scale GPU cluster production operations.
8+ years experience in production infrastructure, strong programming in Python or Go, Linux and Kubernetes knowledge, troubleshooting distributed systems.
Job Description: • Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners (NCP) and on-prem environments. • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations. • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations. • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows. • Participate in on-call, incident response, debugging, and durable follow-up work. • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready. Requirements: • 8+ years of experience building or operating production infrastructure. • Strong programming skills in Python, Go, or similar. • Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation. • Ability to troubleshoot distributed systems in production. • Clear communication and ability to work across teams. • BS/MS in Computer Science or equivalent experience. Benefits: • equity • benefits
This job posting was last updated on 6/1/2026