via Careerplug
$120K - 160K a year
Build and own an internal AI platform ensuring production reliability and cost-efficiency.
Requires deep expertise in distributed systems, AWS, infrastructure management, and autonomous staff-level ownership.
Benefits: Health insurance Paid time off Vision insurance 401(k) matching Bonus based on performance Dental insurance About Facility Grid Facility Grid builds commissioning and turnover software for the teams that bring large buildings online: data centers, airports, hospitals, and commercial real estate. Our platform is the system of record proving every piece of equipment was installed, tested, and accepted. Founded in 2012 and backed by Nexa Equity, we are about 30 people with a small, onshore engineering team that is growing. The role We are building an agentic product suite on top of our system of record, and we run engineering the same way: agents that qualify changes, watch production, and share context across the team. This role owns the platform that makes both possible. You take a modern production platform (ECS Fargate, Aurora, GitOps with ArgoCD and Crossplane, SigNoz) and turn it into an AI platform for running and observing agents in production. You will set the direction for how we operate software in an agentic world and you will build it. You also keep production reliable and cheap. This is a hands-on staff role with a lot of autonomy. You decide how the work gets done. You write code every day, mostly through agents you direct. What you will do Build the internal AI platform: harnesses that qualify changes, progressive rollout, and shared context for every engineer’s agents. Move operations from reactive to proactive: agents that baseline the system and flag anomalies before anyone gets paged. Own the production platform and the path to production, including the ephemeral environment every merge request gets. Own reliability, cloud cost, and compliance (SOC 2 today, FedRAMP in progress) with measured results. Evaluate new models, runtimes, and tooling as they ship and adopt what is better. We stay model-independent. What we are looking for We optimize for four things and we test for each of them in the interview loop. Ownership. You take problems end to end without waiting to be asked. You own what you ship all the way to production and you are accountable for it there. Curiosity about code. You read application code as well as infrastructure, and you want to understand how the whole system works and how it fails. Forward-looking engineering. You already work agent-first with tools like Claude Code, and you design the platform for agents to operate as well as people. You expect this to keep changing every few months and you keep up. Using ChatGPT for lookups or Copilot for autocomplete does not meet this bar. Informed skepticism about where these tools fail is welcome. Technical fundamentals. Distributed systems, networking, Linux, and AWS. You can explain how autoscaling, queues, and deploys behave under load and what a change does to a running system. What we offer Fully remote. Real autonomy over how you lead and build. The latest models and tooling, and a greenfield AI platform to design from the ground up. How we interview Hiring manager screen (30 minutes). Reading session (45 minutes). No take-home and no live coding. We share real infrastructure code and configuration with real problems and you talk us through it. Infrastructure design (60 minutes). Compensation and location Fully remote in the US. We have an office in Waltham, MA if you want one. $180,000 to $225,000 base depending on experience, plus a 10 to 20 percent annual cash bonus. Full benefits including medical and dental. Facility Grid is an equal opportunity employer. This is a remote position.
This job posting was last updated on 9/18/2026