Find your dream job faster with JobLogr
AI-powered job search, resume help, and more.
Try for Free
JO

Jobgether

via Lever.co

All our jobs are verified from trusted employers and sources. We connect to legitimate platforms only.

Senior Site Reliability Engineer

Anywhere
Full-time
Posted 8/25/2026
Direct Apply
Key Skills:
Site Reliability Engineering
Python
Infrastructure as Code

Compensation

Salary Range

$80K - 140K a year

Responsibilities

Develop and scale Python automation frameworks and infrastructure-as-code utilities for AI hardware, lead incident response, and manage telemetry pipelines.

Requirements

5+ years in SRE or systems engineering with a technical bachelor's degree and proficiency in Python, networking, and observability tools.

Full Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States. This is an opportunity to join a critical AI Hardware SRE team responsible for the reliability of next-generation dedicated AI infrastructure. You will help scale and optimize high-density hardware and software environments across regional data centers. The role combines automation, observability, infrastructure engineering, networking, and real-time incident response. You will build Python-based tooling, infrastructure-as-code utilities, telemetry pipelines, and intelligent monitoring solutions. Your work will directly improve uptime, performance, scalability, and operational efficiency for business-critical systems. You will collaborate with engineering teams, infrastructure vendors, and field technicians to solve complex reliability challenges. This role is ideal for an experienced SRE who thrives on ownership, ambiguity, automation, and production-scale infrastructure. \n Accountabilities Develop and scale robust Python-based tooling, infrastructure-as-code utilities, and automation frameworks to eliminate operational toil and streamline fleet-wide provisioning. Build automated workflows and API integrations across corporate ticketing systems to accelerate resolution of hardware and network incidents. Apply modern AI and LLM-based development tools to improve technical execution, automate scripting, and evaluate complex infrastructure systems. Work with advanced private cloud and compute technologies to improve availability, latency, scalability, and overall health across high-density hardware environments. Design and implement telemetry pipelines, Prometheus and Grafana dashboards, and AI-driven anomaly detection for bare-metal and virtualized infrastructure. Define operational KPIs, monitoring standards, telemetry baselines, alerting thresholds, and operational readiness criteria for new services and infrastructure deployments. Participate in a 24x7x365 on-call rotation, leading real-time incident response and managing high-severity service disruptions through automated PagerDuty and Slack workflows. Develop detailed technical runbooks, lead incident response bridges, and drive blameless post-mortems that identify systemic improvements and prevent recurring issues. Partner with infrastructure vendors and coordinate on-site field technicians to support hardware reliability, break-fix activities, and uptime objectives. Collaborate across engineering and infrastructure teams to identify reliability gaps, establish best practices, and deliver production-grade solutions to ambiguous technical challenges. Requirements 5+ years of relevant Site Reliability Engineering, infrastructure engineering, systems engineering, or related experience, along with a Bachelor’s degree in Computer Science or a related technical field. Exceptional proficiency in Python and experience developing scalable operational tooling, API integrations, automation frameworks, and infrastructure utilities. Hands-on experience with modern observability technologies such as Prometheus, Grafana, OpenTelemetry, and Loki, as well as familiarity with time-series monitoring and telemetry systems. Strong understanding of advanced networking concepts, including high-bandwidth routing and switching, BGP, and dual-stack IPv4/IPv6 environments. Experience designing and launching new services with clear operational readiness requirements, telemetry baselines, monitoring strategies, and alerting thresholds. Extensive experience creating technical runbooks, leading complex incident response processes, and conducting comprehensive, blameless post-mortems. Strong understanding of distributed infrastructure, high-density compute environments, private cloud technologies, and large-scale content or infrastructure delivery challenges. Ability to leverage AI-assisted development tools and LLM-based approaches to accelerate engineering workflows and solve technical problems effectively. Proven ability to take ownership of ambiguous and complex technical challenges, coordinate cross-functional teams, and drive solutions through to production. Strong communication and collaboration skills, with the ability to work effectively with engineering teams, vendors, and field operations. Willingness to participate in a 24x7x365 on-call rotation and respond effectively to high-severity production incidents. Benefits Comprehensive benefits designed to support employee health, well-being, financial security, and life beyond work. Flexible working options that allow employees to work from home, in an office, or through a combination of both, depending on role and business needs. Opportunity to work on cutting-edge AI hardware, private cloud, distributed infrastructure, and edge technologies. Exposure to large-scale, business-critical systems serving global digital experiences. Collaborative environment with opportunities to work alongside experienced infrastructure, engineering, and technology professionals. Opportunities to develop expertise in SRE, observability, automation, AI-assisted engineering, networking, and high-density compute. Support for professional growth and continued development within a technology-focused environment. \n How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

This job posting was last updated on 8/25/2026

Ready to have AI work for you in your job search?

Sign-up for free and start using JobLogr today!

Get Started »
JobLogr badgeTinyLaunch BadgeJobLogr - AI Job Search Tools to Land Your Next Job Faster than Ever | Product Hunt