3 open positions available
Lead architecture, incident command, and scaling strategy for multi-region AWS infrastructure and observability platforms. | Requires 9+ years experience, deep Kubernetes and AWS expertise, Terraform mastery, senior incident command, and strong communication skills. | Job Description: • Define architecture and operational standards for Refinery as a Service and Honeycomb Private Cloud across multiple AWS accounts and regions • Architect Terraform modules, Helm charts, and deployment automation • Set technical direction for instrumentation, monitoring, and operation of managed infrastructure • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization across multi-region AWS environments • Build platforms and automation that enable the FRE team to scale • Serve as final technical escalation point for novel, high-stakes customer situations • Resolve infrastructure and observability issues involving distributed systems, Kubernetes, AWS networking, and polyglot service meshes • Partner with customer SRE, platform, and engineering leadership on escalations and architecture redesigns • Provide senior incident command for managed services • Build playbooks, diagnostic frameworks, and tooling • Shape Honeycomb's OpenTelemetry open-source strategy and represent Honeycomb in OpenTelemetry SIGs • Build reference architectures and integration guides • Lead contributions to Honeycomb open-source projects • Act as final technical authority for Solutions Architects on complex deals and production troubleshooting • Lead architecture reviews, SLO workshops, instrumentation deep-dives, POCs, and pilots • Drive roadmap prioritization based on strategic customer needs • Build internal tools and UIs for the FRE function • Drive alignment across Solutions Architecture, Customer Success, Support, Product, and Engineering • Manage trade-offs across the FRE charter • Mentor IC3/IC4 engineers on technical scope and career development • Communicate with Honeycomb leadership and customer C-suite Requirements: • 9+ years in engineering, SRE, infrastructure, DevOps, or equivalent, with demonstrated Staff-level technical scope and impact • Deep hands-on Kubernetes experience, with EKS strongly preferred • Strong AWS expertise across EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, and Route53 • Fluency in multi-account architecture design, service quotas, and cost optimization strategy • Senior incident command experience, including incident response, triage, and postmortem process improvements • Infrastructure as Code mastery with Terraform, Helm, Chef, and Ansible • Deep observability expertise in structured logging, distributed tracing, metrics, SLOs/SLIs, and instrumentation lifecycle • Strong command of OpenTelemetry SDK, Collector architecture, processors, exporters, and semantic conventions, with public community contribution and leadership • Proficiency in at least two of Go, Python, Java, TypeScript/Node.js, and .NET • Excellent executive communication skills • Ability to set direction in ambiguous, high-pressure situations and build structures enabling independent operation • Visa sponsorship or visa transfers are not currently supported • Must verify identity and eligibility to work Benefits: • Generous equity with employee-friendly stock program • Transparent pay based on levels relative to experience • Unlimited PTO • Distributed-first mindset and culture • Home office, co-working, and internet stipend • Full benefits coverage for employees, with additional coverage available for dependents • Up to 16 weeks of paid parental leave, regardless of path to parenthood • Annual development allowance
Lead architecture, escalation, and platform engineering for managed infrastructure and customer-facing technical leadership. | 9+ years engineering with staff-level impact, Kubernetes/EKS expertise, AWS multi-account design, Terraform mastery, observability, OpenTelemetry leadership, and executive communication skills. | Little more about the team: The Field Reliability Engineer (FRE) is an extension of Honeycomb's extensive brand and technical expertise focused on our customer base via engagements, support and our managed services. In this role you parachute into the most complex, highest-stakes technical situations our customers face - unblocking them when they're stuck, creating tooling, guiding our Solution Architecture team through complex decisions, and ensuring prospects and customers get the most out of Honeycomb and the broader observability ecosystem. You're equal parts platform engineer and customer engineer. You build and operate the managed infrastructure our largest customers depend on, and you're the technical backstop the SA team calls when a deal gets deep into infrastructure, data pipelines, or production debugging. This isn't a support role - it's a technical leadership role where you solve problems that don't have runbooks yet, build the platforms and tooling so the next person can have a runbook, and directly impact revenue by unblocking our most strategic deals. What You'll Do Platform Engineering - Architecture & Standards Ownership • Define the architecture and operational standards for Refinery as a Service (RaaS) and Honeycomb Private Cloud (HnyPC) - decisions other engineers build within - across multiple AWS accounts and regions. • Architect the Terraform modules, Helm charts, and deployment automation that other FREs build on and extend, not just consume. • Set the technical direction for how Honeycomb instruments, monitors, and operates its own managed infrastructure - using Honeycomb to monitor Honeycomb. • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization across multi-region AWS environments. • Build platforms and automation that change how the FRE team operates at scale - enabling the team to grow without proportional headcount. Technical Escalation & Unblocking • Serve as the final technical escalation point for the most novel, highest-stakes customer situations - problems with no precedent in existing runbooks. • Resolve deep infrastructure and observability issues spanning distributed systems, Kubernetes clusters, AWS networking (ALBs, PrivateLink, NLBs, VPCs), and polyglot service meshes, often in real time under revenue-critical pressure. • Partner directly with customer SRE, platform, and engineering leadership to navigate multi-week escalations and architecture redesigns tied to the company's largest relationships. • Provide senior incident command for managed services (RaaS, HnyPC) - the point of last escalation when Tier 2 support and IC3/IC4 engineers need staff-level judgment. • Build the playbooks, diagnostic frameworks, and tooling that let IC3/IC4 engineers operate independently in situations that previously required staff involvement. Open Source & Ecosystem Leadership • Shape Honeycomb's open source strategy in the OpenTelemetry ecosystem - drive initiatives, not just contributions. • Represent Honeycomb at the community level in OpenTelemetry SIGs, setting direction on collectors, exporters, and instrumentation libraries the broader ecosystem depends on. • Spot whitespace in the OTel ecosystem that creates structural customer friction, and lead the effort to close it - upstream or through Honeycomb tooling. • Build reference architectures and integration guides that set the standard for effective instrumentation across common customer environments (Kubernetes, ECS, serverless). • Lead feature and architecture contributions to Honeycomb's own open source projects (Refinery, Honeycomb Collector Distro) that support managed service capabilities at scale. Technical Backstop for the Field • Be the final technical authority Solutions Architects call when a deal goes deeper than any existing playbook - join live production troubleshooting, validate architecture decisions, and provide the infrastructure credibility that closes the company's most technically demanding evaluations. • Own the infrastructure and data pipeline narrative on the company's most strategic accounts, in partnership with SA leadership. • Lead architecture reviews, SLO workshops, and instrumentation deep-dives for the most complex customer environments (multi-cluster Kubernetes, hybrid cloud, high-cardinality workloads) - often advising customer VPs and C-suite directly. • Step into the highest-stakes customer-facing POCs and pilots as technical lead, standing up collector pools, configuring Refinery pipelines, and proving out integrations in the customer's actual environment. • Drive prioritization at the roadmap level by identifying strategic gaps between Honeycomb's product capabilities and what the company's largest customers need. Internal Tooling, Mentorship & Cross-Functional Leadership • Build internal tools and UIs that change how the FRE function operates - deployment dashboards, rule management interfaces, and monitoring tooling used company-wide. • Drive alignment across Solutions Architecture, Customer Success, Support, Product, and Engineering by synthesizing field signal into a coherent picture others can act on. • Manage trade-offs across the FRE charter - managed services, technical escalation, open source, and field backstop - making prioritization calls that others follow. • Mentor IC3/IC4 engineers on career development, not just technical skills - helping them build the scope, judgment, and stakeholder navigation needed for their next level. • Communicate at the executive level with both Honeycomb leadership and customer C-suite when the situation demands it. What We're Looking For Required • 9+ years in engineering, SRE, infrastructure, DevOps, or equivalent - with demonstrated Staff-level (or equivalent) technical scope and impact, not just tenure. • Deep hands-on experience with Kubernetes (EKS strongly preferred) - you've deployed, scaled, and operated production clusters at scale, and set standards for how others do the same. • Strong AWS expertise across core services (EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, Route53), with fluency in multi-account architecture design, service quotas, and cost optimization strategy. • A track record of senior incident command - not just participation in on-call, but owning incident response, triage, and postmortem process improvements at the function level. • Infrastructure as Code mastery (Terraform, Helm, Chef, Ansible) - you've architected modules and standards that other engineers build on, not just written your own. • Deep observability expertise: structured logging, distributed tracing, metrics, SLOs/SLIs, and the full instrumentation lifecycle, with a record of setting standards for others. • Strong command of OpenTelemetry (SDK, Collector architecture, processors, exporters, semantic conventions) or equivalent, including public community contribution and leadership. • Proficiency in at least two of: Go, Python, Java, TypeScript/Node.js, .NET - enough to read, instrument, and debug customer code at a deep level. • Excellent executive communication skills - equally credible with a customer's staff SRE and their VP or C-suite, and with Honeycomb's own leadership. • Demonstrated ability to set direction in ambiguous, high-pressure situations with no precedent - and to build the structure that lets others operate independently afterward. Nice to Haves • Background leading customer-facing engineering functions - solutions architecture, field engineering, or technical consulting - where you've owned the technical answer a deal hinged on. • Track record of building platforms and organizational systems, not just tools, that let a team scale without proportional headcount. • Public recognition or a leadership role in the CNCF/OpenTelemetry ecosystem - maintainer status, SIG leadership, or conference speaking. • Familiarity with Honeycomb or event-based observability approaches. • Experience operating telemetry pipelines at scale (OTel Collectors, Refinery/sampling, tail-based sampling, pipeline reliability). • Experience with managed SaaS deployments, private cloud offerings, or multi-tenant infrastructure operations. • Prior experience at an observability, monitoring, or developer tools vendor. • Experience mentoring senior or staff-track engineers on scope and career progression.
Design and deliver production-grade AI agents that investigate, reason, and act on live observability data within the Canvas workspace. | Requires deep expertise in building and shipping LLM-based systems that are reliable in production environments with strong product judgment and end-to-end engineering ability. | What We’re Building Honeycomb is a service for the near and present future, defining observability and raising expectations of what developer tools can do! We’re working with well known companies like HelloFresh, Slack, LaunchDarkly, and Vanguard and more across a range of industries. This is an exciting time in our trajectory, we’ve closed Series D funding, scaled past the 200-person mark, and were named to Forbes’ America’s Best Startups of 2022 and 2023! If you want to see what we’ve been up to, please check out these blog posts and Honeycomb.io press releases. Who We Are We come for the impact, and stay for the culture! We’re a talented, opinionated, passionate, fiercely inclusive, and responsible group of bees. We have conviction and we strive to live our values every day. We want our people to do what they truly love amongst a team of highly talented (but humble) peers. How We Work We are a fully distributed company, which means we believe it is not where you sit, but how you deliver that matters most. We invest in our people and care about how you orient to our culture and processes. At the same time we imbue a lot of trust, autonomy, and accountability from Day 1. #LI-Remote About the role AI agents at Honeycomb investigate, reason, and act on real observability data. They live in Canvas: the agentic workspace where engineers go to understand their systems. The Agentic Intelligence team has shipped Canvas, the Honeycomb MCP server, and the Canvas Agent and Canvas Skills surfaces. What we're looking for today is someone who brings deep agent expertise and uses it to expand what the team can build: new agents, new surface area in Canvas, memory, spatial awareness, improved performance on our Bedrock loop. Honeycomb's data store is fast and accepts high cardinality data; that's what makes agents built on top of it different from anything built on a conventional observability backend. This role is about taking advantage of that building agents that can do things no other observability product can do because the underlying data makes it possible. Some of this work will start as a prototype. The expectation is that the code makes it through the full arc, from the rough first version through to something that holds up in production. What you'll do Design and deliver production-grade agents. Build agents that investigate, reason, and act on live observability data inside Canvas. These agents must be trustworthy to engineers in high pressure situations, including mid-incident. Take one from rough first version to something that holds up under production traffic. Own the agent work; support the whole product. Scope, build, ship, and maintain the agents including the evals that tell you whether they got better or are just different. This role is agent-focused and also includes some fullstack development. Build agents only Honeycomb can build. Use a data store that returns high-cardinality queries in seconds to reason over signal a conventional backend can't serve at this fidelity correlating across services, drilling into a single trace, comparing before and after a deploy. Extend the surface, and decide what's next. Ship new capability into Canvas, the MCP server, and Canvas Skills memory, spatial awareness, a faster Bedrock loop and make the case for what comes after with working code. Distinguish hype from signal in a field with plenty of both. Define what "good" means for agents here. Set the bar: measurable against real evals, maintainable, and honest about their limits. Example projects Multiple agents collaborating on one shared Canvas investigation each claiming a hypothesis, publishing findings, and narrowing the search space for the others so it resolves faster (blog). Auto-investigation the moment an SLO burn alert fires the agent forms hypotheses and prepares visualizations before a human looks, cutting mean-time-to-insight for on-call (o11ycon 2026). Skills that encode a team's domain expertise e.g. Kubernetes thresholds so agents and human colleagues can lean on them (o11ycon 2026). What you'll bring: AI and agent engineering experience. You've shipped LLM-based systems people relied on in production not demos, not fine-tuned models in a research context. You know where agent systems break and how to design around it. End-to-end ownership. On a small team there's no handoff queue. You can take something from rough prototype to production-grade without needing someone behind you to do the durable engineering. Current judgment, not just past experience. You have informed opinions about what's shifted in agent design in the last six to twelve months that would change how you'd build today. Agent architecture depth. You understand how a fast, high-cardinality data store changes what an agent can reason about, and how to design for that. Product judgment. You can look at what the agent layer does today and see what it should do next and make that case with a prototype, not a deck. Even better Observability or developer-tools background. Engineers are your users; you'll ramp faster with fluency in that world, and the work is better. Familiarity with eval frameworks, agent tooling, RAG, and prompt engineering. Base Salary based on level of experience $183,340—$206,000 USD What you'll get when you join the Hive: A stake in our success - generous equity with employee-friendly stock program It’s not about how strong of a negotiator you are - our pay is based on transparent levels relative to experience Time to recharge with unlimited PTO A distributed-first mindset and culture (really!) Home office, co-working, and internet stipend Full benefits coverage for employees, with additional coverage available for dependents Up to 16 weeks of paid parental leave, regardless of path to parenthood Annual development allowance And much more... Please note we cannot currently sponsor or support visa transfers at this time. Additionally, in compliance with applicable law, all persons hired will be required to verify identity and eligibility to work. Phishing and Recruitment Scam Warning: We take your security seriously. Please be aware that recruitment scams are increasingly common and scammers may create email addresses or websites to impersonate Honeycomb employees. To help protect you: All communications will come from an @honeycomb.io email address We occasionally work with external recruiting agencies. These partners will use legitimate business email addresses—never personal accounts like Gmail or Yahoo. Our recruiting process will never ask you to provide financial or sensitive personal information, including but not limited to: Social security or tax identification numbers Credit card numbers Bank account information Diversity & Accommodations: We're committed to building a diverse, inclusive, and equitable workplace—where people of all backgrounds, identities, experiences, and abilities are welcomed, valued, and supported. We recognize that there is no single path to success and embrace nontraditional career journeys and diverse perspectives as key to building stronger, more innovative teams. We strive to ensure an inclusive experience throughout every stage of our hiring process and are happy to provide reasonable accommodations as needed. If you require accommodations or accessible formats at any point during our hiring process, please let your recruiter know. As an equal opportunity employer our hiring process is designed to put you at ease and help you show your best work. If there’s anything we can do to improve your experience, we’re always open to feedback. Privacy Notice: If you apply for a job at Honeycomb and your application is unsuccessful (or you withdraw from the process or decline our offer), Honeycomb will retain your information after your application for a period of time in accordance with local laws. We retain this information for various reasons, including in case we face a legal challenge in respect of a recruitment decision, to consider you for other current or future jobs at Honeycomb, and to help us better understand, analyze and improve our recruitment processes. For more information regarding our privacy practices please see the Honeycomb Privacy Notice. If you do not want us to retain your information for consideration for other roles, or want us to update it, please contact privacy@honeycomb.io. Please note, however, that we may retain some information if required by law or as necessary to protect ourselves from legal claims.
Create tailored applications specifically for Honeycomb.io with our AI-powered resume builder
Get Started for Free