Find your dream job faster with JobLogr
AI-powered job search, resume help, and more.
Try for Free
Qode

Qode

via Workable

All our jobs are verified from trusted employers and sources. We connect to legitimate platforms only.

Senior Site Reliability Engineer

Anywhere
Full-time
Posted 4/15/2026
Direct Apply
Key Skills:
Observability
API Design
System Architecture

Compensation

Salary Range

$90K - 150K a year

Responsibilities

Design and maintain observability solutions ensuring high availability and performance of distributed systems, lead incident response and automation efforts.

Requirements

7–10+ years in SRE or production support with proficiency in observability tools, cloud-native environments, microservices, and scripting in Python or Go.

Full Description

Job Title: Senior Site Reliability Engineer (Observability & Transaction Reliability)Location: Austin, TXType: Full-Time About IncedoIncedo is a global AI and data transformation firm helping organizations drive measurable business impact from digital investments. We operate at the intersection of business and technology, combining AI, data, and digital engineering to deliver scalable, high-impact solutions.With over 4,000 professionals across the U.S., Canada, Latin America, and India, Incedo partners with Fortune 500 and high-growth organizations across banking, payments, wealth management, telecom, and life sciences. Role OverviewWe are seeking a Senior Site Reliability Engineer (SRE) to drive reliability, observability, and performance across business-critical distributed systems.This is a hands-on engineering role with strong ownership, focused on building and scaling observability platforms, improving transaction visibility, and enhancing system resilience. You will work closely with engineering, platform, and infrastructure teams to ensure high availability, performance, and operational excellence across microservices, APIs, and cloud-native systems.The ideal candidate combines deep technical expertise in SRE practices with a passion for automation, monitoring, and continuous improvement. Key ResponsibilitiesObservability & Monitoring Design, implement, and maintain observability solutions across distributed systems Build and optimize logging, metrics, and tracing pipelines using tools like Dynatrace, Datadog, Splunk, ELK, Grafana, and OpenTelemetry Enable end-to-end transaction tracing across microservices and APIs Develop dashboards and alerting strategies for proactive issue detection Reliability & Incident Management Own service reliability, uptime, and operational performance for critical systems Lead incident response, root cause analysis (RCA), and postmortems Reduce MTTD and MTTR through automation and improved observability Create and maintain runbooks and incident response playbooks Performance Engineering Monitor and optimize system performance (latency, throughput, error rates) Partner with application and database teams to troubleshoot bottlenecks Use distributed tracing and telemetry data to identify and resolve issues Implement performance testing and tuning strategies Resiliency & Automation Build and maintain fault-tolerant, highly available systems Implement resiliency patterns (failover, retries, circuit breakers, self-healing) Drive chaos engineering practices to validate system reliability Automate operational tasks using scripting (Python, Go, etc.) SRE Best Practices & Governance Define and enforce SLOs, SLIs, and error budgets aligned to business goals Promote SRE principles across engineering teams Partner with DevOps and platform teams to improve CI/CD reliability Contribute to building a culture of operational excellence and accountability Required Qualifications 7–10+ years of experience in Site Reliability Engineering or Production Support Engineering Strong hands-on experience with observability tools (Dynatrace, Datadog, Splunk, ELK, Grafana, OpenTelemetry, Jaeger) Experience supporting cloud-native environments (AWS, Azure, or GCP) Deep understanding of microservices architecture and distributed systems Proficiency in scripting/programming (Python, Go, Java, or similar) Experience with monitoring, alerting, and incident management in production environments Preferred Qualifications Experience implementing OpenTelemetry at scale Background in chaos engineering and resiliency testing Familiarity with AIOps or intelligent monitoring platforms Experience in financial services, banking, or wealth management environments Dynatrace certification (Associate or Professional) What Success Looks Like Measurable reduction in MTTD and MTTR Increased proactive detection of issues through monitoring Improved system uptime, performance, and reliability Strong adoption of SRE best practices across engineering teams Why Join Us Work on high-impact, mission-critical systems Drive modern SRE and observability practices at scale Collaborate with top-tier engineering and architecture teams Opportunity to influence reliability strategy across the organization

This job posting was last updated on 4/16/2026

Ready to have AI work for you in your job search?

Sign-up for free and start using JobLogr today!

Get Started »
JobLogr badgeTinyLaunch BadgeJobLogr - AI Job Search Tools to Land Your Next Job Faster than Ever | Product Hunt