4 open positions available
Design, deploy, and operate highly available database clusters and lead incident responses. | 7+ years in SRE/DevOps with 3+ years managing production databases and a bachelor's degree. | IMPORTANT: Please be aware, scammers may try to impersonate Zello by reaching out regarding job opportunities. We will never ask you for bank account information, checks, or other sensitive information as part of our hiring process. All correspondence will come from the zello.com email domain. If you’re unsure, please email recruiting@zello.com with questions. About Zello Zello is a voice-first communication platform, powered by our industry-leading push-to-talk technology, to improve collaboration and productivity for desk-less workers. With over 175+ million users, we’re the #1 rated push-to-talk app in the world, delivering 9 billion (yes, with a B) messages a month. At Zello, our company values are at the heart of what we do everyday. We’re proud to serve the frontline, we’re privileged to connect people in times of crisis across the globe, and we’re honored to support first responders. And this is where you come in. We're seeking a Senior Site Reliability Engineer who can own our data tier at high availability while also pulling weight across the broader platform. As Zello scales, the line between "database problem" and "platform problem" keeps blurring. We want someone who can sit on either side of it. This hire owns our data tier reliability (MySQL, MongoDB, ScyllaDB, Elasticsearch, Redis) and contributes to monitoring, on-call, and our ongoing cloud modernization efforts. About Zello Zello is the leading push-to-talk communication platform, enabling instant voice communication for frontline workers across hospitality, logistics, transportation, construction, and public safety. When a hotel manager radios housekeeping or a trucker calls dispatch, they're on Zello — and they need it to work every time. The Platform team builds and operates the infrastructure that makes that possible. Databases sit at the center of that promise: every channel, every message, every login depends on them. The Role You'll join the Platform team and report to the Director of Platform Engineering. You'll own the reliability of our MySQL and MongoDB footprint across Google Cloud, work alongside application engineers on performance and schema decisions, and contribute to the broader platform, observability with Prometheus, Loki, and Tempo; on-call; incident response;. This role suits someone who likes operating real production systems, doesn't get stage fright in incidents, and writes the runbook for the next person who hits the same problem. We're investing in AI to compress incident response, build agents and tooling that speed up root-cause analysis, and lift developer productivity across engineering. We want someone curious about what that looks like for an SRE and excited to help shape it. After a Successful First Year, You Will Have: Operated Zello's MySQL and MongoDB clusters to documented availability targets, with automated backups, regularly tested restores, and failover the on-call team trusts under real incident pressure. Cut latency or capacity cost on at least one critical database workload through measurable performance work — index strategy, query tuning, schema changes, or sharding. Extended our Observability coverage so incidents are diagnosed in minutes rather than hours, with dashboards and alerts the team actually uses. Owned a slice of the Platform on-call rotation and led postmortems that turned recurring incidents into permanent fixes. What You'll Do Design, deploy, and operate highly available MySQL and MongoDB clusters across our cloud environments; replication, sharding, backups, point-in-time recovery, upgrades, and disaster recovery. Tune query performance, schema, and index strategy in partnership with application engineers and push fixes upstream into the application when that's the right answer. Extend our observability stack — Prometheus, Loki, and Tempo — so the data tier is as well instrumented as the application tier, and traces actually reach the root cause. Participate in the Platform on-call rotation, lead incident response for data-tier issues, and write postmortems that drive durable change. Improve disaster recovery, security posture, and compliance for our database footprint — encryption, access control, audit logging, backup integrity. Evaluate and operate ScyllaDB/Cassandra and Elasticsearch where they fit the workload, and bring an opinion on when they don't. Write the automation, tooling, and operators that take repetitive work off the team's plate. Use AI to compress incident response and root-cause analysis; building agents, automation, and developer-enablement tooling that scale the team's reliability work Who You Are You've operated highly available MySQL and MongoDB in production at scale; replication, sharding, backups, point-in-time recovery, and failover drills you've actually run, not just designed on paper. You diagnose database performance end-to-end; query plan, indexes, locking, OS, storage, network — and can point to specific incidents where you found and fixed root cause that others had missed. You've shipped meaningful work on at least two of bare metal Linux, containerized workloads (Docker, Kubernetes, or similar), and a major cloud (GCP preferred; AWS or Azure equivalent is fine). You instrument what you build. You've used Prometheus, OpenTelemetry, or comparable systems to close real incidents, and you've written the dashboard the next on-call engineer will actually open. You write code that runs in production: Python, Go, Bash, or similar for automation, tooling, or operators. You don't hand off scripting to someone else. You communicate clearly under pressure and after the fact. Your postmortems are blameless, specific, and lead to changes that stick — and the people you've worked with describe collaborating with you as straightforward. You bring an opinion on managed vs. self-managed databases, and can defend the trade-off based on availability, cost, and operational burden. 7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, with at least 3 years owning production databases. BSc in Computer Science or equivalent practical experience. ScyllaDB/Cassandra or Elasticsearch experience is a plus You've used AI tooling: copilots, agents, or custom automation to expedite incident response, root-cause analysis, or developer workflows. We hire for potential, passion for our mission, and a knack for solving difficult problems over checking every qualification box. We have competitive pay, equity with significant upside, and intentionally design our benefits to encourage healthy and well-balanced employees, flexible schedules and time off. We even offer a sabbatical after every five years of service so you’re able to pursue and enjoy what matters most to you. And of course, we wouldn’t be a technology company without a ping-pong table and free snacks in our break room. Join us! Zello provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. All Zello personnel are required to comply with defined security, privacy, and compliance requirements applicable to their role along with requirements that are applicable to all Zello personnel.
Operate as a senior individual contributor on backend services, managing critical distributed systems and incident response. | Extensive C++ production experience, deep distributed systems knowledge, on-call readiness, and proactive observability practices. | IMPORTANT: Please be aware, scammers may try to impersonate Zello by reaching out regarding job opportunities. We will never ask you for bank account information, checks, or other sensitive information as part of our hiring process. All correspondence will come from the zello.com email domain. If you’re unsure, please email recruiting@zello.com with questions. About Zello Zello is a voice-first communication platform, powered by our industry-leading push-to-talk technology, to improve collaboration and productivity for desk-less workers. With over 175+ million users, we’re the #1 rated push-to-talk app in the world, delivering 9 billion (yes, with a B) messages a month. At Zello, our company values are at the heart of what we do everyday. We’re proud to serve the frontline, we’re privileged to connect people in times of crisis across the globe, and we’re honored to support first responders. And this is where you come in. This role is on Zello's backend services team. The team that owns and operates the distributed systems carrying every conversation our customers depend on. You'll join as the third engineer on a high-ownership team, taking full production responsibility for our most critical services from day one. There are no layers of management between you and the work that matters. After a successful first year, you will Have owned full production responsibility for at least one critical existing service; on-call, SLOs, incident response and measurably improved its reliability through lower P0 frequency or shorter MTTR. Have extended observability coverage with Grafana and Loki dashboards and alerting, shifting the team's posture from reactive to proactive on at least one reliability dimension. Have shipped at least three non-trivial reliability, performance, or architecture improvements into production. Have contributed a meaningful component to backend 2.0. Our ground-up rearchitecture toward an AI-native engineering posture. What you'll do Operate as a senior individual contributor on the backend services team, owning production for distributed systems at real-time scale. Carry the on-call rotation for Zello's most critical infrastructure. Lead incident response, write runbooks, and turn each incident into a durable improvement. Build observability where it's missing. Set SLOs, instrument services with Grafana and Loki, and shift team posture from reactive to proactive. Ship production C++ today; contribute to Rust and Go components as they come online. Work hands-on across our database stack: Cassandra, MongoDB, MySQL and choosing the right tool for the workload. Contribute to backend 2.0: a ground-up modernization of the platform anchored on AI-native engineering workflows (AGENTS.md, CLAUDE.md, .cursor/rules). Who you are You've carried real production weight on distributed systems at scale, and you talk about reliability as something you own, not something you measure. You've shipped C++ in production for years and read long-lived backend codebases without flinching. Rust or Go in your toolkit is a plus. You've gone deep on at least one of Cassandra, MongoDB, or MySQL, and can name the tradeoffs because you've felt them in production. You use AI-assisted development as a default and have shaped CLAUDE.md, AGENTS.md, .cursor/rules, or equivalent for a real team. You thrive on a small team where every hire moves the needle. You drop ego in code review. You take scope, you don't wait for it. Your definition of "done" includes tests, alerts, observability, and runbooks. This Role Is Not A management or tech-lead-without-keyboard role. The team is three people and you're shipping production code week one. A pure greenfield architect role. Backend 2.0 is real, but reliability and ops on existing critical services come first. A "ship your AI tooling and call it a day" role. AI-native posture is the floor, not the ceiling. A role with layers of approval between idea and ship. If you want managerial cover before you act, this isn't the seat. We hire for potential, passion for our mission, and a knack for solving difficult problems over checking every qualification box. We have competitive pay, equity with significant upside, and intentionally design our benefits to encourage healthy and well-balanced employees, flexible schedules and time off. We even offer a sabbatical after every five years of service so you’re able to pursue and enjoy what matters most to you. And of course, we wouldn’t be a technology company without a ping-pong table and free snacks in our break room. Join us! Zello provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. All Zello personnel are required to comply with defined security, privacy, and compliance requirements applicable to their role along with requirements that are applicable to all Zello personnel.
Provide second-level technical support by investigating complex issues, managing escalations, supporting enterprise integrations, and mentoring team members. | 3–5 years technical support experience in SaaS or communications technology with proficiency in networking, APIs, mobile platforms, and log interpretation. | IMPORTANT: Please be aware, scammers may try to impersonate Zello by reaching out regarding job opportunities. We will never ask you for bank account information, checks, or other sensitive information as part of our hiring process. All correspondence will come from the zello.com email domain. If you’re unsure, please email recruiting@zello.com with questions. About Zello Zello is a voice-first communication platform, powered by our industry-leading push-to-talk technology, to improve collaboration and productivity for desk-less workers. With over 175+ million users, we’re the #1 rated push-to-talk app in the world, delivering 9 billion (yes, with a B) messages a month. At Zello, our company values are at the heart of what we do everyday. We’re proud to serve the frontline, we’re privileged to connect people in times of crisis across the globe, and we’re honored to support first responders. And this is where you come in. Overview The Support Engineer provides second-level technical support, bridging the gap between the Product Advocate (L1) team and Engineering. You will investigate advanced technical issues—including app-level, network, and API integrations—working directly with enterprise customers and developers to resolve complex problems. Your work ensures Zello remains a reliable, high-performing solution for organizations that depend on it daily. You will report to the Product Advocate Manager and collaborate with the Product, Engineering, Sales and Customer Success teams. Mission Deliver deep technical expertise to diagnose and resolve complex product issues, supporting both enterprise customers and the Product Advocate team. Focus areas include: Advanced Troubleshooting and Root-Cause Analysis Escalation Management and Cross-Team Coordination API, SDK, and Integration Support Enterprise Implementation Support Responsibilities Advanced Troubleshooting & Root-Cause Analysis Investigate complex technical issues beyond the scope of L1 support. Use diagnostic tools, logs, and APIs to isolate and identify root causes. Reproduce and document product bugs for Engineering. Provide troubleshooting support for PAs on hybrid software/hardware solutions and on-premise server products. Technical Support for Enterprise and Developer Accounts Support enterprise deployments, integrations, and custom configurations. Assist third-party developers integrating Zello SDKs and APIs. Help customers design robust solutions using Zello technology. Assist with implementation of MDM solutions and SSO for enterprise customers. Escalation & Collaboration Serve as the primary liaison between L1 Support and Engineering. Ensure accurate, complete escalation documentation and follow-up until resolution. Aid in relationship management by acting as key technical resource for ongoing Enterprise and Partnership communications. Mentor Product Advocates in advanced troubleshooting and technical concepts. Continuous Improvement Identify recurring issues and propose fixes or automation tools. Contribute to internal knowledge bases and troubleshooting guides. Qualifications 3–5 years of technical support or related experience in a SaaS or communications technology company. Strong understanding of APIs, networking fundamentals, and mobile platforms (Android/iOS). Proficiency in reading and interpreting logs, JSON, and basic scripting. Excellent written and verbal communication skills. Customer-first mindset with attention to clarity and accuracy. Career Path Support Engineers can advance into roles such as: Senior Support Engineer (L3) – Expert in Zello’s full product ecosystem, mentor and process owner. Solutions Engineer – Supporting pre-sales and custom integrations. Software Engineer – Transitioning into Engineering for those contributing to code-level debugging and automation. We hire for potential, passion for our mission, and a knack for solving difficult problems over checking every qualification box. We have competitive pay, equity with significant upside, and intentionally design our benefits to encourage healthy and well-balanced employees, flexible schedules and time off. We even offer a sabbatical after every five years of service so you’re able to pursue and enjoy what matters most to you. And of course, we wouldn’t be a technology company without a ping-pong table and free snacks in our break room. Join us! Zello provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. All Zello personnel are required to comply with defined security, privacy, and compliance requirements applicable to their role along with requirements that are applicable to all Zello personnel.
Design, implement, and maintain scalable infrastructure and platform tooling that accelerates development and increases reliability. Collaborate with product engineers and security to ensure the platform is secure, efficient, and easy to use. | A seasoned engineer with 5+ years of experience in designing, building, and operating scalable systems is required. Strong programming skills in Python and familiarity with GCP, AWS, and Kubernetes are essential. | IMPORTANT: Please be aware, scammers may try to impersonate Zello by reaching out regarding job opportunities. We will never ask you for bank account information, checks, or other sensitive information as part of our hiring process. All correspondence will come from the zello.com email domain. If you’re unsure, please email recruiting@zello.com with questions. About Zello Zello is a voice-first communication platform, powered by our industry-leading push-to-talk technology, to improve collaboration and productivity for desk-less workers. With over 175+ million users, we’re the #1 rated push-to-talk app in the world, delivering 9 billion (yes, with a B) messages a month. At Zello, our company values are at the heart of what we do everyday. We’re proud to serve the frontline, we’re privileged to connect people in times of crisis across the globe, and we’re honored to support first responders. And this is where you come in. The Platform Engineering team at Zello builds and maintains the foundational systems that power our products and improve the developer experience for our engineering organization. As a Senior Software Engineer on this team, you will design, implement, and maintain scalable infrastructure and platform tooling that accelerates development, increases reliability, and streamlines operations. You will collaborate closely with product engineers and security to ensure our platform is secure, efficient, and easy to use. This role requires strong problem-solving skills, a high level of technical ownership, and a passion for creating tools and systems that empower others to move faster with confidence. After a successful first year, you will Designed, implemented, and launched at least one major platform service or infrastructure enhancement that measurably improved developer productivity. Delivered developer self-service capabilities through tooling and an internal developer portal, reducing time-to-deploy for product teams. Improved reliability of Zello’s systems by leading application changes that enhanced detection, alerting, and automated response to failures. What you'll do Design, build, and maintain platform services that support Zello’s development teams Develop automation, tooling, and processes to improve the developer experience Partner with other engineers and security to define and evolve infrastructure standards and patterns. Participate in on-call rotations to ensure uptime and responsiveness for critical services Identify issues before they become problems and take ownership from concept to delivery. Mentor other engineers by sharing expertise, reviewing code, and driving adoption of best practices. Who you are A seasoned engineer with 5+ years experience designing, building, and operating scalable, reliable systems in production Have strong programming skills in Python; familiarity with Go is a plus Have experience working with GCP and/or AWS Familiarity with distributed systems deployed on Kubernetes You enjoy building the tools, infrastructure, and abstractions that empower other engineers to move faster We hire for potential, passion for our mission, and a knack for solving difficult problems over checking every qualification box. We have competitive pay, equity with significant upside, and intentionally design our benefits to encourage healthy and well-balanced employees, flexible schedules and time off. We even offer a sabbatical after every five years of service so you’re able to pursue and enjoy what matters most to you. And of course, we wouldn’t be a technology company in Austin without a ping-pong table and free snacks in our break room. Join us! Zello provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. #LI-Hybrid
Create tailored applications specifically for Zello with our AI-powered resume builder
Get Started for Free