Railway

Railway

4 open positions available

1 location
1 employment type
Actively hiring
Full-time

Latest Positions

Showing 4 most recent jobs
Railway

Senior Product Engineer, Scalability

RailwayAnywhereFull-time
View Job
Compensation$70K - 130K a year

Architect and scale high-throughput backend systems including usage metering, billing, and fraud detection pipelines. | Requires deep expertise in Postgres, Node.js internals, scaling systems, money movement, idempotency, and asynchronous workflow orchestration. | Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing. Railway now powers workloads for millions of builders, and the systems underneath — usage metering, billing and payments, fraud and abuse protection, background workers, and the data pipelines that feed them all — have to scale every week. You'll be the person who architects that stack: making it fast, correct, and trustworthy at a volume that keeps growing. If you're looking to scale the backbone of an operating system for builders, we'd love to talk with you! Want to learn about our work culture? Here is a multi-part blog series that will help you see the unique ways our team works (Parts 1, 2, 3, and 4). About the role This is a backend-leaning role focused on scaling systems. Billing and fraud will be a focus, but your remit spans every high-throughput system at Railway — workers, queues, event pipelines, and the databases underneath them. You'll own your work end-to-end, including when a feature reaches the UI. For this role, you will: Architect and scale the pipelines that turn raw usage into accurate, real-time billing — metering, aggregation, rating, and invoicing across millions of events, from ingestion in ClickHouse to the rating engine. Build payment flows that are correct under concurrency and partial failure: idempotent charges, retries, reconciliation, and clean handling of provider edge cases (Stripe and beyond). Develop fraud and abuse detection — signal collection, real-time scoring, automated mitigation — that protects platform margin without getting in legitimate users' way. Scale the systems everything else depends on: Postgres under heavy write load, Node.js services under pressure, and long-running workflows orchestrated with Temporal where exactly-once semantics and durability actually matter. Build TypeScript + GraphQL APIs where correctness and auditability are non-negotiable. Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring its success and scaling it further. Contribute to our open-source repositories (CLI, Typescript SDK, Railpack, etc.) — Rust experience, or the desire to learn it, helps here. Be oncall from time to time. Some projects this team takes on: Re-architect billing end-to-end: per-second usage metering at platform scale, idempotent payment processing that survives provider outages without double-charging, and credit, prepayment, and enterprise-invoicing models that hold up under audit. Stand up a fraud-detection service that scores signups and deployments in real time and automatically throttles abuse (crypto mining, free-tier farming, stolen cards). Scale our Temporal workloads to orchestrate workflows across millions of deployments. Build internal tooling that gives teams across Railway a trustworthy, real-time view into the systems they depend on. This is a high impact, high agency role with direct effect on company culture, trajectory, and outcome. About you An ability to autonomously lead, design, and implement backend systems where correctness, consistency, and auditability are first-class requirements. A track record of scaling systems — you've taken a pipeline, service, or database that was falling over and made it handle 10x, and you know which tools to reach for (and when polling stops being enough). Deep expertise in Postgres and relational data modeling — you reach for the right consistency guarantees, understand the cost of getting them wrong, and know how Postgres itself behaves at scale. Strong working knowledge of Node.js internals — the event loop, memory behavior, and what to do when a service degrades under load. Experience managing complex asynchronous and long-running backend jobs, ideally with a workflow engine like Temporal, for things like billing runs or payment reconciliation. Familiarity with the realities of money movement: payment providers, idempotency, retries, reconciliation, and their failure modes. Direct billing, payments, or fraud experience is a strong plus. A security and abuse-aware mindset — you instinctively think about how a system can be gamed, and you design accordingly. A desire to be a part of the entire project development process, from research gathering and planning, to implementation and monitoring. Great written and verbal communication skills for expressing ideas, designs, and potential solutions in a mostly-asynchronous manner. We value and love to work with diverse persons from all backgrounds. Things to know For better or worse, we're a startup; our team dynamics are different from companies of different sizes and stages. We're globally distributed—and getting more so. Stuff is always happening somewhere. We don't expect you to be online all the time, but you'll need to be diligent about your boundaries — your end of day will overlap with someone else's start. We're a small, high-ownership team that cares deeply about doing exceptional work. We're scaling quickly, which means we rely on leverage—systems over coordination, judgment over process. Expect ambiguity and a fast-moving environment. You'll own real outcomes. That means making decisions, not just executing—and owning the success, or failure, that comes with them. Benefits and perks At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more. For more details, check back on the main careers page. Beyond compensation, there are a few things that we believe make working at Railway truly unique: Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work. Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company. Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor "busywork" and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions. Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there. How we hire No tricks. No surprises. Here's the entire process. 1 — Talk with us about the role This is completely open ended and we're just trying to see who you are, what you want to do, and where you wanna go. 2 — Work on a small project to discuss in the interview Asynchronously design a system that scales. You choose the domain — pick whichever shows your thinking best: A usage metering and billing pipeline that meters CPU/RAM for millions of workloads and bills accurately (you may depend on third parties such as Stripe), or A stream-processing system that ingests high-cardinality observability events in real time, or Whatever you pick, come ready to defend the architecture end-to-end. We'll dig into: Polling vs. stream processing, and how you avoid losing data when streaming Correctness under concurrency and partial failure: idempotency, retries, reconciliation, what happens when a step fails halfway through How you handle cardinality, and which tools you lean on and why The scalability of the things you depend on — what happens when Postgres becomes the bottleneck Interview Structure to expect (60 Minutes): Prework (submitted before your interview): Your design 0–5 minutes: Introductions 5–35 minutes: Walking through the design and how you'd extend it — new failure modes, 10x load, a fraud signal 35–50 minutes: Noodling on technology, data modeling, and how you think about scale, money-movement, and abuse 50–60 minutes: Time for you to ask your interviewers questions You can, and SHOULD! ask us questions ahead of time. Ask away! 3 — Review your solution with the Team You'll sit down with someone on the team and go over the above. We'll poke into your solution, as well as get you acquainted with two more members of the team. Looking for: Learn about your problem solving skills. How you break down a problem and how you present a solution. 4 — Meet the Team You'll meet the Team, which will be comprised of 4 people from vastly different sections of the company. Looking for: How you work with the rest of the team and communicate. 5 — Chat with CEO Sit down with our founder and CEO for 30 minutes. This is a 1:1, open ended conversation. 6 — Offer call Finally, we will present the offers, hammer out the details about your position, tee up onboarding, and start our journey together. Final Note: The interview goes both ways. Once again, please ask us things. Many things! Hard things. That's what we're here for.

Node.js
TypeScript
API Design
System Architecture
Distributed Systems
Direct Apply
Posted about 2 months ago
Railway

Senior Infra Engineer: Baremetal Orchestration

RailwayAnywhereFull-time
View Job
Compensation$80K - 140K a year

Build and maintain host provisioning and orchestration infrastructure with tooling and CI pipelines. | Strong expertise in distributed systems, bare metal provisioning, Golang or Rust, and configuration management. | Job description Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing. Many infrastructure platforms simply focus on how you deploy your singular application, and now how these applications function in concert. Questions like “How do you build systems for zero downtime deployment”, “How do you do service-to-service communications”, etc are usually left up to the engineers to define. At Railway, our goal is to be an all encompassing solution to all these problems. As such, we take special care as we define our networking infrastructure. “But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing” - Radia Perlman About the role For this role, you will: Build and maintain our host provisioning stack: PXE boot, Ansible, and burn-in agents that bring new bare metal online quickly and confidently Continue to evolve our homegrown orchestration engine to manage clusters, containers, and VMs through a single lens Optimize the efficiency of our bin packing algorithm to maximize utilization/performance and minimize costs Own the internal tooling that Railway engineers use to interact with our fleet every day Build out internal observability and alerting so we catch fleet problems before customers feel them Design and maintain the CI pipelines that ship our infrastructure code safely Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible Build Golang/Rust GRPC services from scratch capable of supporting millions of users Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring its success The arc of this role is more internal-facing than user-facing. You're building the platform that Railway engineers run on. This is a high impact, high agency role with direct effect on company culture, trajectory, and outcome. About you A strong understanding of distributed systems and what it takes to operate them. You enjoy building fault tolerant, resilient, and scalable services, and you care about what happens when they break at 3am Hands-on experience with bare metal provisioning, configuration management, and the unglamorous-but-critical work of getting hardware production-ready Comfort building and operating internal tools. You understand that developer experience inside the company matters as much as the product outside it A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo The tact to implement your solution, create monitors for its error boundaries, and document any requirements for when you're not around A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed A great set of communication skills for getting your point across, solution implemented, and beyond We value and love to work with diverse persons from all backgrounds Things to know For better or worse, we're a startup; our team dynamics are different from companies of different sizes and stages. We're distributed ALL across the globe, and that's only going to be more and more distributed. As a result, stuff is ALWAYS happening. We do NOT expect you to work all the time, but you'll have to be diligent about your boundaries because the end of your day may overlap with the start of someone else's. We're a small team, with high ownership, who are not only passionate about what we do, but seek to be exceptional as well. At the time of writing we're 21, serving hundreds of thousands of users. There's a lot of stuff going on, and a lot of ambiguity. We want you to own it. We believe that ownership is a key to growth, and part of that growth is not only being able to make the choices, but owning the success, or failure, that comes with those choices. Benefits and perks At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more. For more details, check back on the main careers page. Beyond compensation, there are a few things that we believe that make working at Railway truly unique: Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work. Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company. Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor “busywork” and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions. Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there. How we hire No tricks. No surprises. Here's the entire process. 1) Talk with us about the role This is completely open ended and we're just trying to see who you are, what you want to do, and where you wanna go. 2) Work on a small project to discuss in the interview Asynchronously implement the following: Imagine a theoretical or actual system like Railway which can manage stateless and stateful compute workloads. Design the engine for managing orchestration Interview Structure (60 Minutes): Pre-work (before your interview): Complete your solution (advised) 0-5m: introduction 5-50m: Building (or expanding) your solution 50-60m: Questions on Railway/Tech/etc You can, and SHOULD! ask us questions ahead of time. Ask away! 3) Review your solution with the Team You'll sit down with someone on the team and go over the above. We'll poke into your solution, as well as get you acquainted with two more members of the team. Looking for: Learn about your problem solving skills. How you break down a problem and how you present a solution. 4) Meet the Team You'll meet the Team, which will be comprised of 4 people from vastly different sections of the company. Looking for: How you work with the rest of the team and communicate. 5) Chat with CEO Sit down with our founder and CEO for 30 minutes. This is a 1:1, open ended conversation. 6) Offer call Finally, we will present the offers, hammer out the details about your position, tee up onboarding, and start our journey together. Final Note: The interview goes both ways. Once again, please ask us things. Many things! Hard things. That's what we're here for.

Distributed Systems
API Design
Observability
Direct Apply
Posted 3 months ago
Railway

Senior Infra Engineer: Observability

RailwayAnywhereFull-time
View Job
Compensation$90K - 130K a year

Build scalable, fault tolerant observability infrastructure and APIs for high-throughput telemetry ingestion and alerting. | Strong understanding of distributed systems, infrastructure as code, and building scalable, resilient backend services with good communication skills. | Job description Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing. Many infrastructure platforms simply focus on how you deploy your singular application, and now how these applications function in concert. Questions like “How do you build systems for zero downtime deployment”, “How do you do service-to-service communications”, etc are usually left up to the engineers to define. At Railway, our goal is to be an all encompassing solution to all these problems. As such, we take special care as we define our networking infrastructure. Note: Networking falls under the platform engineering umbrella. If you’re specialized, we’d love to chat! That said, we’d also like it noted you’re probably going to do a lot of non-networking + platform things “But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing” - Radia Perlman About The Role For this role, you will: • Build ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry • Build scalable, fault tolerant alerting engines for notifying users, in real-time, of threshold breaches • Craft rich backend observability APIs, working with product to build amazing experiences for instantly grokking their application • Provide APIs to access realtime log/metrics streams to be consumed by the Dashboard and Product Teams • Build Golang/Rust GRPC services from scratch capable of supporting tens of thousands of users, and the million+ to come. • Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible. • Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success. • Interface with our TypeScript and GraphQL edge to expose your microservice APIs for both internal and potentially external consumption This is a high impact, high agency role with direct effect on company culture, trajectory, and outcome. About You • A strong understanding of distributed systems. You enjoy building fault tolerant, resilient, and scalable services • Interests in VictoriaMetrics, ClickHouse, and other systems for building observability stacks from the ground up • A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo. • The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around • A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup • A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed • A great set of communication skills for getting your point across, solution implemented, and beyond We value and love to work with diverse persons from all backgrounds Things to know For better or worse, we're a startup; our team dynamics are different from companies of different sizes and stages. • We're distributed ALL across the globe, and that's only going to be more and more distributed. As a result, stuff is ALWAYS happening. • We do NOT expect you to work all the time, but you'll have to be diligent about your boundaries because the end of your day may overlap with the start of someone else's. • We're a small team, with high ownership, who are not only passionate about what we do, but seek to be exceptional as well. At the time of writing we're 21, serving hundreds of thousands of users. There's a lot of stuff going on, and a lot of ambiguity. • We want you to own it. We believe that ownership is a key to growth, and part of that growth is not only being able to make the choices, but owning the success, or failure, that comes with those choices. Benefits and perks At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more. For more details, check back on the main careers page. Beyond compensation, there are a few things that we believe that make working at Railway truly unique: • Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work. • Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company. • Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor “busywork” and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions. • Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there. How we hire No tricks. No surprises. Here's the entire process: • Talk with us about the role • This is completely open ended and we're just trying to see who you are, what you want to do, and where you wanna go. • Work on a small project to discuss in the interview • Asynchronously implement the following: • Imagine a theoretical or actual system like Railway which can manage stateless and stateful compute workloads. Design the engine for managing observability • Interview Structure (60 Minutes): • Pre-work (before your interview): Complete your solution (advised) • 0-5m: introduction • 5-50m: Building (or expanding) your solution • 50-60m: Questions on Railway/Tech/etc. You can, and SHOULD! ask us questions ahead of time. Ask away! • Review your solution with the Team You'll sit down with someone on the team and go over the above. We'll poke into your solution, as well as get you acquainted with two more members of the team. Looking for: Learn about your problem solving skills. How you break down a problem and how you present a solution. • Interview Structure (60 Minutes): • Prework (submitted before your interview): Complete your solution • 0-5m: introduction • 5-50m: Building (or expanding) your solution • 50-60m: Questions on Railway/Tech/etc • Meet the Team • You'll meet the Team, which will be comprised of 4 people from vastly different sections of the company. • Looking for: How you work with the rest of the team and communicate. • Offer and Details Chat with CEO • Finally, we will go over the process, the role, and hammer out the details about your position, onboarding, and all the deets. Final Note: The interview goes both ways. Once again, please ask us things. Many things! Hard things. That's what we're here for.

API design
System architecture
Observability and monitoring
Verified Source
Posted 3 months ago
Railway

Senior Platform Engineer: Storage

RailwayAnywhereFull-time
View Job
Compensation$Not specified

Design and evolve production Ceph clusters, including hardware design and network requirements. Create APIs for live-migrations and build storage primitives for customer applications and internal services. | Experience in architecting distributed systems and production experience with block device systems like Ceph is required. A solid understanding of current and next-gen filesystems is also important. | Job description Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing. Building the infrastructure which powers the Railway engine is the most core problem at Railway. As an infrastructure engineer working on stoarge, you will be directly responsible for designing software and hardware to back performant, high reliability block storage and object storage systems backing millions of applications. The solutions you build will be instrumental in not only scaling internal operations, but scaling the company to infinity and beyond! “But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing” - Radia Perlman Curious? Here are 3 blog posts that dive into exciting projects this team has worked on: 1, 2, 3 Want to learn about our work culture? Here is a three-part blog series that will help you see the unique ways our team works (Parts 1, 2, 3, and 4). About The Role For this role, you will: Design and evolve multiple production Ceph clusters, from hardware design, to driving network requirements to configuring, tuning and operating clusters and their clients Create efficient, generalizable APIs using systems/kernel features to provide safe, as-fast-as-possible live-migrations of stateful workload between hosts Design and build API and Orchestration services to tie storage primitives to higher level primitives using Go, gRPC, ScyllaDB and Temporal Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success Design build a suite of storage primitives that can be used by customer applications, internal services and enable higher level platform features such as streaming image pulls or movable build caches About You Experience architecting and implementing distributed systems. You enjoy building fault tolerant, resilient, and scalable services Production experience with distributed block device systems (e.g Ceph) or a solid understanding of network storage cluster design from first principles Understanding and experience with current gen filesystems (Ext4, ZFS, BTRFS). Bonus points for next gen (EROFS, bcachefs) A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo. The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed A great set of communication skills for getting your point across, solution implemented, and beyond We value and love to work with diverse persons from all backgrounds Things to know For better or worse, we're a startup; our team dynamics are different from companies of different sizes and stages. We're distributed ALL across the globe, and that's only going to be more and more distributed. As a result, stuff is ALWAYS happening. We do NOT expect you to work all the time, but you'll have to be diligent about your boundaries because the end of your day may overlap with the start of someone else's. We're a small team, with high ownership, who are not only passionate about what we do, but seek to be exceptional as well. At the time of writing we're 21, serving hundreds of thousands of users. There's a lot of stuff going on, and a lot of ambiguity. We want you to own it. We believe that ownership is a key to growth, and part of that growth is not only being able to make the choices, but owning the success, or failure, that comes with those choices. Benefits and perks At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more. For more details, check back on the main careers page. Beyond compensation, there are a few things that we believe that make working at Railway truly unique: Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work. Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company. Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor “busywork” and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions. Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there. How we hire No tricks. No surprises. Here's the entire process: Talk with us about the role This is completely open ended and we're just trying to see who you are, what you want to do, and where you wanna go. Work on a small project to discuss in the interview Asynchronously implement the following: Pre-interview: Design a Storage Engine to power something like Railway's Volume You can, and SHOULD! ask us questions ahead of time. Review your solution with the Team You'll sit down with someone on the team and go over the above. We'll poke into your solution, as well as get you acquainted with two more members of the team. Looking for: Learn about your problem solving skills. How you break down a problem and how you present a solution. Interview Structure (60 Minutes): Prework (submitted before your interview): Complete your solution 0-5m: introduction 5-50m: Building (or expanding) your solution 50-60m: Questions on Railway/Tech/etc Meet the Team You'll meet the Team, which will be comprised of 4 people from vastly different sections of the company. Looking for: How you work with the rest of the team and communicate. Offer and Details Chat with CEO Finally, we will go over the process, the role, and hammer out the details about your position, onboarding, and all the deets. #Global

Distributed Systems
Fault Tolerance
Resilience
Scalability
Block Device Systems
Network Storage
Filesystems
Monitoring
Communication
Problem Solving
Ownership
Autonomy
Creative Solutions
Growth
Direct Apply
Posted about 1 year ago

Ready to join Railway?

Create tailored applications specifically for Railway with our AI-powered resume builder

Get Started for Free

Ready to have AI work for you in your job search?

Sign-up for free and start using JobLogr today!

Get Started »
JobLogr badgeTinyLaunch BadgeJobLogr - AI Job Search Tools to Land Your Next Job Faster than Ever | Product Hunt