via Gem
$90K - 140K a year
Improve platform reliability and performance by managing AWS infrastructure, automating tasks, leading incident response, mentoring staff, and collaborating with product teams.
Over 10 years software engineering experience, 5 years in SRE/DevOps, strong AWS, Terraform, observability skills, and ability to contribute to application code.
Senior Site Reliability Engineer Improve the reliability, performance, and operational maturity of a platform that supports the future of educational fundraising. CONTRACT-TO-HIRE REMOTE - UNITED STATES SEATTLE / WEST COAST PREFERRED About the role Our client is looking for a hands-on Senior Site Reliability Engineer to improve the reliability, performance, and operational maturity of our platform. With our migration to AWS complete, this role will focus on strengthening our production environment: improving observability, automating infrastructure and operational work, enhancing incident response, and partnering with product engineers to build resilient systems. This is a high-impact role for someone who understands both infrastructure and application development. You will work across the stack, contribute code when appropriate, and help ensure our systems remain secure, scalable, and dependable as we grow. What you'll do • Operate, maintain, and improve our production infrastructure in AWS. • Build and maintain infrastructure as code using Terraform. • Improve monitoring, alerting, dashboards, and service-level indicators using New Relic or comparable observability platforms. • Reduce alert noise and build systems that identify problems before customers are affected. • Participate in the 24/7 on-call rotation and help coordinate the response to production incidents. • Lead blameless postmortems and ensure corrective actions result in durable improvements. • Partner with product engineers to diagnose performance and reliability issues throughout the application stack. • Improve application resilience through appropriate use of timeouts, retries, queuing, backpressure, and idempotency. • Improve CI/CD pipelines and deployment practices using platforms such as GitHub Actions, GitLab CI, or CircleCI. • Automate repetitive operational work and reduce engineering toil. • Create and maintain runbooks, system diagrams, troubleshooting guides, and production documentation. • Support capacity planning, performance testing, database reliability, and production-readiness reviews. • Collaborate with Security and Engineering teams on infrastructure hardening, access controls, logging, and compliance-related operational practices. • Mentor engineers and promote effective reliability practices across the Engineering organization What we're looking for • 10+ years of overall software engineering, infrastructure, or systems experience, including at least 5 years in an SRE, Platform Engineering, DevOps, or production operations role. • Previous professional software development experience and the ability to read, debug, and contribute to application code. • Strong, hands-on experience operating production workloads in AWS. • Experience building and maintaining infrastructure with Terraform or a similar infrastructure-as-code tool. • Strong observability skills using New Relic, Datadog, or another modern monitoring platform. • Experience with incident response, on-call operations, postmortems, and production troubleshooting. • Experience building or maintaining CI/CD pipelines. • Working knowledge of networking, Linux, distributed systems, and relational databases. • Strong judgment when balancing immediate operational needs with long-term maintainability. • Clear communication skills and the ability to collaborate effectively across engineering disciplines. • A track record of using automation to improve reliability and create leverage for other engineers. Bonus points • Experience with Ruby or Ruby on Rails. • Strong PostgreSQL administration or performance-tuning experience. • Experience operating enterprise SaaS products at scale. • Familiarity with SLOs, SLIs, error budgets, capacity modeling, and load testing. • Experience with payments, fintech, or other highly regulated systems. • Experience supporting SOC 2 or similar security and compliance programs. Role details • Contract-to-hire. • Remote within the United States. • Seattle-area or West Coast candidates are preferred to support occasional in-person collaboration, but exceptional candidates elsewhere should also be considered. • Participation in a shared on-call rotation is required.
This job posting was last updated on 9/8/2026