via Remote Rocketship
$120K - 180K a year
Lead architecture, incident command, and scaling strategy for multi-region AWS infrastructure and observability platforms.
Requires 9+ years experience, deep Kubernetes and AWS expertise, Terraform mastery, senior incident command, and strong communication skills.
Job Description: • Define architecture and operational standards for Refinery as a Service and Honeycomb Private Cloud across multiple AWS accounts and regions • Architect Terraform modules, Helm charts, and deployment automation • Set technical direction for instrumentation, monitoring, and operation of managed infrastructure • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization across multi-region AWS environments • Build platforms and automation that enable the FRE team to scale • Serve as final technical escalation point for novel, high-stakes customer situations • Resolve infrastructure and observability issues involving distributed systems, Kubernetes, AWS networking, and polyglot service meshes • Partner with customer SRE, platform, and engineering leadership on escalations and architecture redesigns • Provide senior incident command for managed services • Build playbooks, diagnostic frameworks, and tooling • Shape Honeycomb's OpenTelemetry open-source strategy and represent Honeycomb in OpenTelemetry SIGs • Build reference architectures and integration guides • Lead contributions to Honeycomb open-source projects • Act as final technical authority for Solutions Architects on complex deals and production troubleshooting • Lead architecture reviews, SLO workshops, instrumentation deep-dives, POCs, and pilots • Drive roadmap prioritization based on strategic customer needs • Build internal tools and UIs for the FRE function • Drive alignment across Solutions Architecture, Customer Success, Support, Product, and Engineering • Manage trade-offs across the FRE charter • Mentor IC3/IC4 engineers on technical scope and career development • Communicate with Honeycomb leadership and customer C-suite Requirements: • 9+ years in engineering, SRE, infrastructure, DevOps, or equivalent, with demonstrated Staff-level technical scope and impact • Deep hands-on Kubernetes experience, with EKS strongly preferred • Strong AWS expertise across EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, and Route53 • Fluency in multi-account architecture design, service quotas, and cost optimization strategy • Senior incident command experience, including incident response, triage, and postmortem process improvements • Infrastructure as Code mastery with Terraform, Helm, Chef, and Ansible • Deep observability expertise in structured logging, distributed tracing, metrics, SLOs/SLIs, and instrumentation lifecycle • Strong command of OpenTelemetry SDK, Collector architecture, processors, exporters, and semantic conventions, with public community contribution and leadership • Proficiency in at least two of Go, Python, Java, TypeScript/Node.js, and .NET • Excellent executive communication skills • Ability to set direction in ambiguous, high-pressure situations and build structures enabling independent operation • Visa sponsorship or visa transfers are not currently supported • Must verify identity and eligibility to work Benefits: • Generous equity with employee-friendly stock program • Transparent pay based on levels relative to experience • Unlimited PTO • Distributed-first mindset and culture • Home office, co-working, and internet stipend • Full benefits coverage for employees, with additional coverage available for dependents • Up to 16 weeks of paid parental leave, regardless of path to parenthood • Annual development allowance
This job posting was last updated on 8/17/2026