via Ladders
$184K - 357K a year
Design and develop scalable automation and orchestration systems for cloud infrastructure workflows.
Senior-level experience with programming, cloud infrastructure, container orchestration, and distributed systems.
We are looking for a Senior Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA's high-performance GPU infrastructure. You will play a critical role in designing scalable automation solutions, integrating diverse systems, and enabling seamless workflows across global cloud operations. What You'll Be Doing: • Design and develop APIs to orchestrate and integrate operational workflows. • Build state management and workflow automation systems that streamline infrastructure lifecycle processes. • Collaborate across teams to codify business processes into scalable, self-measuring systems. • Develop extensible, schema-driven platforms for reducing manual toil and ensuring operational consistency. • Drive integrations with container orchestration tools like Kubernetes and observability systems such as Prometheus, OpenTelemetry, Grafana. • Optimize the reliability and efficiency of cloud operations through automated workflows and telemetry systems. • Lead and ship impactful technical projects, ensuring quality and scalability at every stage. What we need to see: • 8+ years of industry experience with a Bachelor's or Master's degree (or equivalent experience), or 2+ years with a PhD. • Expertise in designing, building, and operating services in a high reliability environment. • Proficiency in programming languages such as Go, Java, or Python. • Strong understanding of cloud infrastructure (AWS, GCP, Azure) and container technologies like Docker and Kubernetes. • Experience with high-scale distributed systems, including architectural patterns for APIs and data pipelines. • Outstanding communication and collaboration skills, with a focus on solving complex operational challenges. • A passion for automating manual processes and driving system efficiency. Ways to Stand Out from the Crowd: • A track record of designing workflow orchestration systems for large-scale infrastructure. • Proven experience in reducing operational inefficiencies through automation and integration. • Strong debugging and problem-solving skills in distributed environments. • Prior experience or strong familiarity with the operational aspects of the NVIDIA AI/ML software stack (e.g., CUDA, cuDNN, containerization) Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until July 27, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
This job posting was last updated on 7/27/2026