2 open positions available
Design, build, and operate agentic AI systems while driving adoption of agent-assisted development and owning production monitoring and SRE practices. | 6+ years software engineering experience with 2+ years senior or staff level, strong hands-on AWS serverless, agentic AI frameworks, and modern CI/CD. | We operate a multi-tenant automotive SaaS platform serving thousands of dealer groups across the United States. Our backend — event-driven serverless on AWS — orchestrates everything from dealer onboarding to inventory management to real-time transaction processing. That platform works. Now we need to make it think — and keep it modern, secure, and safe to ship. We also run a monolith that needs upkeep while we build the next-gen platform and redesign its components, a few at a time, into microfrontends. This is a hands-on engineering role at its core. Day to day you build, ship, and operate agentic AI systems — autonomous, tool-using agents (AgentCore, MCP servers) that observe platform state, reason over dealer context, and take action through production APIs. But you are also the engineer who keeps the whole platform current and trustworthy: you continuously hunt and fix security vulnerabilities, upgrade the frameworks and runtimes underneath us, drive adoption of an agentic development harness to speed up delivery, and lead modernization of our existing systems. You work across the org, not inside one team's walls. You own production monitoring and the SRE practices that keep us reliable, and you are the champion for safe releases — defining the checks and balances that gate production, setting the engineering standards, and making sure they are actually enforced. When something misbehaves in production, your instrumentation, your guardrails, and your release discipline are what make it fail safe instead of fail loud. Reports to: SVP, Engineering. What You'll Own - Hands-on agentic development — designing, building, and operating agentic AI systems (AWS Bedrock AgentCore, MCP servers) every day: agent code, tool interfaces, evaluation harnesses, and production AI workflows. - Driving adoption of the agentic development harness — making agent-assisted development a first-class way the org ships, and measurably improving speed of delivery across teams. - Security vulnerability management — continuously monitoring for vulnerabilities (dependencies, CVEs, container images, IAM/config drift), remediating quickly, and keeping frameworks, libraries, and runtimes patched and upgraded. - Platform modernization — creating and driving initiatives to modernize existing platforms: retiring legacy patterns, adopting current frameworks, and staying close to the leading edge. - Production monitoring & SRE — owning observability (OpenTelemetry, CloudWatch), SLOs and error budgets, on-call and incident response. - Third-party integration observability — a standardized way to organize all third-party integration calls, report on them with precision, and run anomaly detection on call volumes with continuous monitoring, so a spike, drop, or failure surfaces before it becomes an outage or a cost problem. - On-call cookbook — keeping applications well documented so an on-call engineer can quickly query a cookbook, understand the likely problem areas, and know where to look. - Release governance — champion across releases: defining and enforcing the checks and balances (CI/CD quality gates, security and test gates, rollback criteria, progressive delivery) required to move change into production. - Setting and enforcing engineering standards — establishing the patterns, guardrails, and review bar the org follows, and holding the line so they are consistently applied. - Working across the org — partnering with every engineering team to raise the technical, security, and reliability bar broadly. Tech & Tools - Cloud: Lambda, EventBridge, DynamoDB, S3, ECS Fargate, Aurora, API Gateway, CloudWatch, Secrets Manager. - AI & Agentic: AWS Bedrock AgentCore, MCP servers, LangChain/LangGraph. - Languages: Python and Java (Spring Boot), TypeScript/React (Next.js), legacy PHP/Laravel. - Security: Dependabot/Snyk, SAST/DAST, container image scanning, secrets management, IAM hardening. - Reliability & Observability: OpenTelemetry, CloudWatch, SLO/error-budget tooling, PagerDuty, Datadog. - CI/CD & Release: CircleCI, CloudFormation, progressive delivery, quality/security gates, rollback automation. - Integration Surfaces: REST, SOAP/XML, EventBridge, SES, Playwright. How You'll Use AI This is not a "we are AI-curious" company. This is the role that makes agent-assisted development the norm for everyone else — building and operating production agents, triaging CVEs and generating patch/upgrade PRs, driving framework upgrades, reviewing PRs across the stack, standing up observability, and drafting ADRs and runbooks. Hands-On Expectations Roughly 60–70% building and operating, 20–30% on design and standards, and ~10% on cross-team enablement. First 12 Months - Months 1–3: Immerse in the codebase, ship your first meaningful changes, audit our security posture, dependency/framework currency, and release pipeline, and publish a baseline of engineering standards and release checks. - Months 4–6: Ship agentic automation into at least one production workflow, roll out the agentic development harness, stand up the SRE baseline (SLOs, on-call, dashboards), and automate vulnerability scanning and patching. - Months 7–9: Drive a modernization initiative end to end; harden the release gates and enforce them across teams. - Months 10–12: Measurable delivery-speed gains from agentic adoption, production reliability against SLOs, and engineering standards adopted org-wide. Must Have - 6+ years of software engineering experience, including 2+ years at a Senior or Staff level building and operating production systems. - Hands-on agentic / LLM development (AWS Bedrock, MCP, LangChain/LangGraph) — or a strong, demonstrated track record of getting up to speed fast on modern frameworks. - Strong hands-on experience with AWS serverless (Lambda, EventBridge, DynamoDB, Step Functions) and traditional service architectures (ECS, RDS, API Gateway). - Security-minded engineering — dependency and vulnerability management, timely patching/upgrades, and secure-by-default practices. - SRE practices — observability, SLOs/error budgets, on-call, and incident response. - CI/CD and release management — quality/security gates, rollback, and progressive delivery. - Comfortable in at least two of Python, Java (Spring Boot), and TypeScript/React — and able to read and safely modify PHP. - A track record of setting standards and getting them enforced, plus the writing to make a design or standard clear enough for a peer to implement. Strongly Preferred - Experience running security tooling (Snyk/Dependabot, SAST/DAST) and observability/on-call tooling (Datadog, PagerDuty) in production. - Automotive, fintech, or multi-tenant marketplace platform experience. - Experience with data pipelines (Glue/Athena) or ETL/data-lake tooling. - Familiarity with Auth0, or with browser automation (Playwright) in production integration flows. Role Specifics - Location: Denver, CO (hybrid, 2 days/week in office) or Remote (US, outside the Denver market). - Employment type: Full-time. - Reports to: SVP, Engineering. Scope & Scale - 5,000+ destination dealer tenants, each with isolated databases and per-tenant configuration. - Billions in annual GMV flowing through platform transactions. - Tens of thousands of API requests per minute across REST, SOAP, and event-driven integration surfaces. - Data pipelines spanning 6 integration domains with multi-protocol vendor connectivity. At A2Z Sync, we replace the friction of disconnected systems with the velocity of a single platform. We pride ourselves on a fun, casual, and collaborative culture, and we're committed to our employees' well-being. - Employer-Paid Health, Dental, and Vision Insurance, starting on day one. - Flexible work: Denver hybrid (2 days/week in office) or fully remote anywhere in the US. - 401(k) Retirement Plan with Company Match. - Generous Paid Time Off: Unlimited PTO and 10 paid holidays. - Short-Term and Long-Term Disability Coverage, fully employer-paid. - Life and AD&D Insurance, employer-paid. - Free Mental Health Support via BetterHelp. - Pet Insurance options. - Identity Theft Protection. - On-site gym in Denver, great co-workers, and a stocked kitchen with snacks and beverages.
Design, develop, and operate agentic AI systems and infrastructure, mentor senior staff, and lead complex system migrations. | 8+ years software engineering with 3+ years in Staff or Principal role, deep AWS serverless expertise, and experience leading large-scale migrations. | Why This Role Exists We operate a multi-tenant automotive SaaS platform serving thousands of dealer groups across the United States. Our backend — event-driven serverless on AWS (Lambda, EventBridge, DynamoDB, S3, Step Functions) — orchestrates everything from dealer onboarding to inventory management to real-time transaction processing. That platform works. Now we need to make it think. We are building agentic AI systems: autonomous, tool-using agents that observe platform state, reason over dealer context, take action through production APIs, and learn from outcomes. These are not chatbots bolted onto a dashboard. They are first-class platform services — backed by AWS Bedrock, connected to production systems via MCP servers — that make decisions, execute workflows, and close loops without human intervention unless guardrails say otherwise. This Principal Engineer owns that entire surface. You are not advising on AI strategy from a whiteboard. You are writing agent code, defining tool interfaces, building evaluation harnesses, setting cost and latency budgets, and shipping production AI workflows that touch real dealers and real money. You set the engineering patterns the team follows, you help make the build-vs-buy calls, and when an agent misbehaves at 2 AM, your architecture is what determines whether it fails safe or fails loud. Scope & Scale 5000+ destination dealer tenants, each with isolated databases and per-tenant configuration. Billions in annual Gross Merchandise Value (GMV) flowing through platform transactions. Tens of thousands of API requests per minute across REST, SOAP, and event-driven integration surfaces. Data pipelines spanning 6 integration domains with multi-protocol vendor connectivity. What You Will Own Ownership and core development of agentic AI systems — designing, building, and operating the AI agent infrastructure (AWS Bedrock, MCP servers) that powers intelligent automation across the platform. You are not advising on AI strategy; you are writing the agent code, defining the tool interfaces, building the evaluation harnesses, and shipping production AI workflows. AI agent lifecycle end to end — from prompt engineering and tool-use design through guardrails, evaluation, cost optimization, and production observability. You own the patterns the team uses to build with AI: how agents connect to production systems, how we evaluate output quality, how we manage model costs at scale, and how we roll back when an agent misbehaves. System design and technical decision-making for migration waves — from identity/tenant services through core domain extraction and frontend decomposition. The dual-write framework, API Gateway traffic-splitting, and per-tenant feature flag rollout that make every migration step reversible. Cross-cutting concerns: observability (OpenTelemetry, CloudWatch), security posture (Auth0 consolidation, IAM), and data architecture (DynamoDB single-table design, Aurora consolidation). Mentoring and force-multiplying senior ICs — establishing patterns, reviewing designs, and raising the technical bar across 5 engineering teams. Consolidate and strategize 30+ different integrations and make the future integrations easier. Technical Environment Cloud Services: High-availability AWS stack including Lambda, EventBridge, DynamoDB, S3, ECS Fargate, Aurora, API Gateway, CloudWatch, and Secrets Manager. Development Languages: Modern Python and Java (Spring Boot) alongside TypeScript/React (Next.js 16) frontends, with legacy domain coverage in PHP/Laravel. AI & Agentic Systems: Advanced agentic workflow orchestration utilizing lean AWS Bedrock AgentCore, MCP servers, or LangChain/LangGraph frameworks. Data Engineering: Complex data architectures featuring DynamoDB single-table design, MySQL/Aurora, S3 data lakes, Glue Data Catalog, Athena, and Data pipelines. Infrastructure & Security: Enterprise-grade CI/CD and observability via CloudFormation, Auth0 consolidation, OpenTelemetry, and CircleCI. Integration Surfaces: Multi-protocol connectivity spanning REST, SOAP/XML, EventBridge event-bus patterns, SES processing, and Playwright browser automation. First 12 Months Months 1–3: Immerse in the codebase. Audit the current architecture across all stacks. Publish the first Architecture Decision Record (ADR) for the next migration wave. Establish your design review cadence with the team. Months 4–6: Drive the AI/agentic integration layer — Bedrock-powered automation in at least one production workflow. Establish the patterns for how the team builds with AI going forward; both agentic insight retrieval agentic workflow automation. Months 7–9: Own and deliver the first migration wave end-to-end — from design doc through production cutover with dual-write validation. Stand up the observability baseline (OpenTelemetry instrumentation, dashboards, SLOs). Months 10–12: Second migration wave in production. Architecture runway documented for the next 12 months. The team operates at a higher technical bar because of patterns you set. You Should Have 8+ years of software engineering experience with at least 3 years in a Staff / Principal / Architect role. Build Cloud native solutions with emphasis on speed to market. Deep hands-on experience (About 40-50% in the code, 40-50% in Design) with AWS serverless (Lambda, EventBridge, DynamoDB, StepFunctions) and traditional service architectures (ECS, RDS, API Gateway). Experience spending 10-20% of your time in mentoring, cross-team alignment and operating mechanisms. Track record of leading monolith-to-services migrations — strangler fig, dual-write validation, traffic-splitting, canary rollouts. Fluency across multiple languages: you can review PHP, architect & prototype Python, architect Java services, and reason about TypeScript frontends. Experience with event-driven architectures, configuration-driven workflow engines, and DynamoDB single-table design. Strong opinions on observability and the discipline to instrument before you migrate. The ability to write a design doc that a senior IC can implement without ambiguity, and the judgment to know when to write code yourself instead. Preferred Qualifications Automotive, fintech, or multi-tenant marketplace platform experience. Experience with data pipelines (NiFi, Glue/Athena) or ETL/data lake tooling. Familiarity with Auth0 Organization model, M2M apps, and Actions for JWT enrichment. Experience with browser automation (Playwright) in production integration flows. About A2Z Sync A2Z Sync is a fast-paced and innovative automotive SaaS company seeking to make life better for our customers. We offer you a fun, casual, and collaborative culture, while fostering an environment where you work hard, see your results, and feel your impact. We are committed to our employees, and this starts with providing benefits that allow you to care for you and your family. Mission At A2Z Sync, we replace the friction of disconnected systems with the velocity of a single platform. We integrate digital insights with in-store operations to deliver transparent transactions that bring clarity to the car buyer and increased profitability to the dealer. Our Values: We Are DRIVEN Dealership Obsessed: We measure our success by the dealer's wins and the trust of their buyers, not just our own code. Relentless Ownership: No lone wolves, but no pass-backs either. We don't say "that's not my job." Invent with Purpose: We don't chase "shiny" tech. We replace guesswork with intelligence, building the "data backbone" that turns raw information into a competitive advantage. Value Every Perspective: We are Better Together. We check egos at the door. Evolve or Evaporate: Change is our constant. We stay ahead by learning faster than the competition. Now Over Next: Perfection is the enemy of progress. We prefer action over endless analysis. Here’s how we are doing it: A2Z Sync offers comprehensive medical, dental, and vision benefits. Employer provided STD/LTD and life insurance. Matching 401k plan. Unlimited paid time off, including 10 paid holidays. Real ownership of a high-stakes AI surface — your roadmap, your architecture decisions, your metrics. The expected salary range for this role is $195,000 to $220,000 annually, commensurate with experience and qualifications.
Create tailored applications specifically for A2Z Sync with our AI-powered resume builder
Get Started for Free