Find your dream job faster with JobLogr
AI-powered job search, resume help, and more.
Try for Free
JO

Jobgether

via Lever.co

All our jobs are verified from trusted employers and sources. We connect to legitimate platforms only.

Staff ML Software Engineer (L6) — Platform Systems, AIMS Engineering

Anywhere
Full-time
Posted 8/17/2026
Direct Apply
Key Skills:
Distributed systems
Python
System architecture

Compensation

Salary Range

$120K - 180K a year

Responsibilities

Design and operate platform subsystems for ML observability and automation, leading complex technical programs to improve infrastructure reliability and cost.

Requirements

Requires strong software engineering with Python and JVM languages, plus experience building and operating production AI/ML systems at scale.

Full Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff ML Software Engineer (L6) — Platform Systems, AIMS Engineering based in the United States. This is a high-impact staff-level engineering role focused on the platform foundations powering large-scale AI and machine learning systems. You will shape observability, evaluation, reliability, and cost-efficiency capabilities for next-generation ML workflows serving hundreds of millions of users. The role combines hands-on software engineering with technical leadership across complex distributed systems and production AI/ML infrastructure. You will help modernize existing platforms while defining the architecture and operational practices that support future model paradigms and AI workloads. Your work will improve how engineers understand model behavior, detect issues, optimize infrastructure, and operate critical systems at scale. You’ll collaborate across multiple teams, influence technical direction without relying on formal authority, and turn ambiguous challenges into reusable platform solutions. This environment is ideal for a technically strong engineer who thrives on high-leverage problems, autonomy, and long-term systems thinking. \n Accountabilities: Design, build, and operate platform subsystems for observability, evaluation, tooling, and operational automation supporting next-generation ML architectures. Prove new platform capabilities against current production operations, including anomaly detection, root-cause analysis, and automated operational workflows. Build observability infrastructure that provides deep visibility into model behavior, training pipeline health, serving latency, and data quality, enabling teams to detect and diagnose issues before they become incidents. Identify and drive infrastructure cost optimization across ML training and serving, developing increasingly automated frameworks and tools that make compute efficiency a core engineering priority. Architect reliability improvements across the AI/ML stack, reducing operational toil, improving on-call experiences, and establishing strong standards for production excellence. Contribute to the target architecture and migration strategy for a modernized ML platform, coordinating with partner engineering teams and managing technical dependencies. Evaluate emerging infrastructure approaches, model paradigms, and platform capabilities, translating relevant developments into a forward-looking technical roadmap. Develop reusable frameworks and common platform patterns that can be adopted across engineering teams and ML workloads. Lead complex technical programs across organizational boundaries, building alignment and consensus while influencing teams without direct authority. Continuously improve critical systems by balancing long-term platform investments with immediate operational priorities. Requirements: Significant experience designing, building, and operating production AI/ML systems at scale, including ML training pipelines and model-serving or online-inference environments. Hands-on experience building infrastructure for advanced agentic or complex model architectures, such as memory, trace, evaluation, replay, orchestration, or routing systems. Strong software engineering fundamentals with deep expertise in Python and working proficiency in at least one JVM language, such as Java or Scala. Proven experience improving the reliability, scalability, operational maturity, and cost efficiency of AI/ML infrastructure. Strong background building observability and monitoring systems for ML workloads, with an understanding of visibility requirements across training, serving, and data pipelines. Deep distributed-systems expertise, including large-scale batch processing and real-time serving infrastructure. Experience collaborating with multiple engineering and partner teams to establish technical direction, manage dependencies, and deliver complex programs. Excellent technical judgment and the ability to identify reusable patterns, determine where investment is valuable, and make pragmatic decisions about scope and priorities. Ability to operate effectively in ambiguous environments, independently defining problems, developing approaches, and adapting as new information emerges. Preferred experience with LLM evaluation, trace, replay, observability, or debugging tools and frameworks for complex model systems. Familiarity with modern ML infrastructure such as feature stores, model-serving platforms, and experimentation frameworks. Experience migrating production AI/ML systems between technology generations and improving their architecture over time. Experience in personalization domains such as recommendation systems, search, or discovery is a plus. Benefits: Annual salary range of $600,000–$1,066,000, with compensation determined based on factors such as role, location, background, skills, and experience. Annual compensation structure consisting of salary and stock options, with employees able to determine their preferred mix each year. Comprehensive health plans and mental health support. 401(k) retirement plan with employer matching. Stock option program. Health Savings Accounts and Flexible Spending Accounts. Disability programs and life and serious injury benefits. Family-forming benefits. Paid leave programs. Flexible time off for full-time salaried employees. Fully remote position within the United States. Opportunity to work on highly scalable AI/ML systems with significant technical and organizational impact. \n How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

This job posting was last updated on 8/17/2026

Ready to have AI work for you in your job search?

Sign-up for free and start using JobLogr today!

Get Started »
JobLogr badgeTinyLaunch BadgeJobLogr - AI Job Search Tools to Land Your Next Job Faster than Ever | Product Hunt