via Workable
$90K - 140K a year
Architect and maintain Kubernetes infrastructure, implement observability systems, and define SLOs to ensure system reliability.
Deep experience with Linux, AWS, Kubernetes, Terraform, programming in PHP, Go, or Bash, and reliability engineering principles.
At Laravel, we don’t just build tools; we build the foundation that empowers millions of developers to ship their dreams. We are looking for a Senior Site Reliability Engineer to help us scale that mission by ensuring our global infrastructure remains as elegant and reliable as the code we write. If you are energized by the challenge of building robust observability systems, running Kubernetes clusters across multiple regions, and solving complex operational puzzles with code, you’ve found your next home. Location / Timezone: Remote, between Central Europe and US East timezones for optimal collaboration with the team. Description of the Role As a Senior Site Reliability Engineer, you will be an early contributor to our dedicated SRE function, reporting directly to Florian Beer. This is a high-impact, autonomous role where you will design and implement the systems that support Laravel Cloud, Nightwatch, and Forge. You will act as a bridge between development and operations, advocating for a blameless culture and shared responsibility for reliability across the entire organization. Your 12-Month Mission Imagine we are all at Laracon in 12 months' time. You are telling the team about your first year, and the impact is undeniable: First 30 Days: You will have selected an SLO monitoring solution. Day 60: You will have guided teams towards identifying SLIs and how/why we work with SLOs at Laravel. Day 90: You have established clear, data-driven SLOs across engineering teams, giving us a unified language for reliability. Year One: You have worked together with our engineering teams to identify SLIs, review existing SLAs, created SLOs, and educated teams on how SLO’s error budgets guide decisions. What You Will Do Architect Reliability: Establish SRE as a core function at Laravel, building the fundamentals from the ground up. System Design: Design, build, and maintain multi-region Kubernetes infrastructure and global distributed systems. Automation: Solve operational challenges through software, reducing manual intervention (toil) for our product teams. Observability: Design and implement monitoring, logging, and alerting systems using tools like Prometheus, Grafana, and Loki. Collaboration: Partner with product leads and SecOps to make reliability a shared responsibility. Requirements - What You Will Bring Infrastructure Mastery: Deep experience with Linux system administration and cloud platforms, specifically AWS. Orchestration & IaC: Proficiency with Kubernetes, Docker, and managing infrastructure via Terraform. Programming Skills: The ability to solve problems with software and scripting using e.g. PHP, Bash, or Go. Systems Thinking: A "smart and passionate" approach to troubleshooting, with the ability to deconstruct complex systems into triagable components. Reliability Mindset: Experience with SLO/SLI/SLA definition, capacity planning, and performance tuning. Soft Skills: A commitment to documentation, cross-team collaboration, and an automation-first mindset. Requirements - Bonus Skills Framework Familiarity: Previous experience working with the Laravel framework and our existing product suite (Cloud, Forge, Vapor, etc.) is highly preferred. Advanced Observability: Experience with Prometheus, Grafana Mimir, and Grafana Loki for metrics storage and alerting. Small tight-knit team where every developer counts Fully remote and globally distributed working environment Option to attend Laracon conferences around the world Health care plan (Medical, Dental & Vision) Paid time off (Vacation, Sick & Public holidays) Family leave (Maternity, Paternity) Pension plans (As locally applicable) Performance based bonus plan Company equity
This job posting was last updated on 9/30/2026