via Clearance Jobs
$80K - 130K a year
Design, implement, and maintain scalable infrastructure, monitoring systems, ETL pipelines, and CI/CD workflows with on-call incident management.
Requires 6+ years SRE/DevOps experience, proficiency with Terraform, AWS/GCP, CI/CD, monitoring tools, on-call operations, and knowledge of cloud-native services.
On behalf of our client, ClearanceJobs Workforce Solutions is seeking a Site Reliability Engineer. This full-time (direct hire) position located in San Francisco, CA. Work is performed 100% on-site. Our client is willing to consider candidates located near Omaha, NE, Boston, MA, and in Washington, D.C. metro area with expectation of quarterly travel to CA and NE. Candidates must be able to work on our client’s W2 (annual salary + benefits package). Due to government contract requirements, United States Citizenship and an active DOD Secret security clearance is required. No 3rd party Corp to Corp inquiries will be considered. Summary: Our client is a cutting-edge startup focused on delivering AI-based weather forecasting solutions to enhance climate resilience and public safety. To deliver advanced weather forecasts, they use numerous cutting-edge computing environments, including cloud-based hyperscalers, on-premise GPU clusters, field-deployed computers, and large supercomputing centers. They work closely with the Department of Defense (DoD) and must adhere to strict compliance and security standards. The team thrives in a dynamic, fast-paced environment, and every member plays a critical role in driving our mission forward. Responsibilities: • Scaling Production Environment: o Design and implement scalable infrastructure solutions to support growing business needs. o Optimize system performance and availability through capacity planning and performance tuning. • Monitoring Stack Improvement: o Develop and maintain whitebox (application-level) and blackbox (system-level) monitoring systems. o Ensure comprehensive observability through the integration of metrics, logging, and tracing. o Utilize tools such as Datadog and PagerDuty to establish reliable alerting and incident response processes. • ETL Pipeline Management: o Design, launch, and maintain robust ETL pipelines to support data-driven operations. o Collaborate with data teams to ensure data quality and pipeline reliability. • CI/CD and Build Systems: o Implement and manage continuous integration and continuous deployment pipelines. o Improve developer productivity by maintaining reliable build systems and workflows. • On-call Responsibilities: o Participate in on-call rotations to ensure high availability and timely incident resolution. o Develop and automate incident response playbooks to minimize downtime. Required: • 6+ years of experience in SRE or DevOps roles. • Proficiency with infrastructure as code tools, particularly Terraform. • Experience with AWS &/or Google Cloud • Strong background in software engineering and CI/CD pipeline management. • Experience with on-call operations and incident management. • Familiarity with monitoring and alerting tools such as Datadog and PagerDuty. • Knowledge of cloud-native services, including SNS/SQS and Redis. • Experience with ML experiment tracking and GPU optimization is a plus. Desired: • Expertise in managing and scaling distributed systems. • Strong understanding of networking, security, and Linux systems. • Experience in automating infrastructure and deployment processes. • Familiarity with message queues (SNS/SQS) and caching systems (Redis). • Knowledge of ML workflows and GPU resource management.
This job posting was last updated on 8/17/2026