18 open positions available
Develop and evaluate AI model training code across full stack including JavaScript and Python, collaborating with researchers and cross-functional teams. | Senior software engineer with 3+ years experience in full-stack development using JavaScript (React, Node.js) and Python, strong software architecture and communication skills. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have worked at the frontier of AI — at companies like OpenAI, NVIDIA, Databricks, Palantir, Snowflake, or similar organizations pushing the boundaries of intelligent systems. We especially welcome graduates from top computer science programs such as Stanford, MIT, Carnegie Mellon, UC Berkeley, Georgia Tech, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections across the full stack — in Python for backend and ML workflows, and JavaScript (React, Node.js) for frontend and API layers, alongside C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code across both Python and JavaScript (React, Node.js), with additional work in C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code across backend and frontend contexts to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of the code and identify error patterns across full-stack applications. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in building full-stack applications using Python and JavaScript (React, Node.js), with the ability to work across backend and frontend codebases. • Experience deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Create and evaluate AI training datasets and code solutions in multiple programming languages, focusing on systems-level code and collaborating with cross-functional teams. | Several years of software engineering experience with strong backend or systems programming skills, expertise in modern languages and tools, and excellent communication abilities. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have built production systems at companies like Google, Microsoft, Apple, Amazon, Meta, or similar high-scale engineering organizations. We especially welcome graduates from programs with strong CS foundations such as University of Washington, University of Illinois Urbana-Champaign, UT Austin, University of Michigan, Purdue, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections in Python, C/C++, Rust, Go, Java, and JavaScript (including ReactJS) — with particular emphasis on systems-level code, performance-critical applications, and infrastructure. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code in Python, C/C++, Rust, Go, Java, and JavaScript (including ReactJS). • Evaluate and refine AI-generated code with an emphasis on systems-level correctness, performance, and reliability. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of systems-level and infrastructure code and identify error patterns. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in systems programming, infrastructure, or backend development using languages like Python, C/C++, Rust, and Go. • Experience building and deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Evaluate and refine AI-generated code across full-stack applications, collaborating with cross-functional teams to enhance AI-driven coding solutions. | Several years of software engineering experience with strong expertise in Python and JavaScript (React, Node.js), and excellent communication skills. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have built production systems at companies like Google, Microsoft, Apple, Amazon, Meta, or similar high-scale engineering organizations. We especially welcome graduates from leading programs such as Harvard, Columbia, Princeton, Yale, University of Pennsylvania, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections across the full stack — in Python for backend and ML workflows, and JavaScript (React, Node.js) for frontend and API layers, alongside C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code across both Python and JavaScript (React, Node.js), with additional work in C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code across backend and frontend contexts to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of the code and identify error patterns across full-stack applications. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in building full-stack applications using Python and JavaScript (React, Node.js), with the ability to work across backend and frontend codebases. • Experience deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Develop and evaluate AI-driven full-stack software solutions, focusing on JavaScript and Python code quality and scalability. | Senior software engineer with 3+ years experience in full-stack development, strong JavaScript and Python skills, and excellent communication. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have worked at the frontier of AI — at companies like OpenAI, NVIDIA, Databricks, Palantir, Snowflake, or similar organizations pushing the boundaries of intelligent systems. We especially welcome graduates from top computer science programs such as Stanford, MIT, Carnegie Mellon, UC Berkeley, Georgia Tech, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections across the full stack — in Python for backend and ML workflows, and JavaScript (React, Node.js) for frontend and API layers, alongside C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code across both Python and JavaScript (React, Node.js), with additional work in C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code across backend and frontend contexts to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of the code and identify error patterns across full-stack applications. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in building full-stack applications using Python and JavaScript (React, Node.js), with the ability to work across backend and frontend codebases. • Experience deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Create and evaluate AI training datasets and code solutions across full-stack applications using Python and JavaScript, collaborating with researchers to improve AI-driven coding solutions. | Several years of software engineering experience with strong full-stack skills in Python and JavaScript (React, Node.js), deep understanding of software architecture and code quality, and excellent communication skills. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have worked at the frontier of AI — at companies like OpenAI, NVIDIA, Databricks, Palantir, Snowflake, or similar organizations pushing the boundaries of intelligent systems. We especially welcome graduates from top computer science programs such as Stanford, MIT, Carnegie Mellon, UC Berkeley, Georgia Tech, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections across the full stack — in Python for backend and ML workflows, and JavaScript (React, Node.js) for frontend and API layers, alongside C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code across both Python and JavaScript (React, Node.js), with additional work in C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code across backend and frontend contexts to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of the code and identify error patterns across full-stack applications. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in building full-stack applications using Python and JavaScript (React, Node.js), with the ability to work across backend and frontend codebases. • Experience deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Create and evaluate AI training datasets by coding and refining full-stack applications in Python and JavaScript, collaborating with researchers and cross-functional teams. | Several years of software engineering experience with strong Python and JavaScript skills, expertise in software architecture and code quality assessment, and excellent communication skills. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have shipped high-impact products at fast-moving companies like Stripe, Airbnb, Cloudflare, Datadog, Coinbase, or similar high-growth engineering environments. We especially welcome graduates from programs with strong CS foundations such as University of Washington, University of Illinois Urbana-Champaign, UT Austin, University of Michigan, Purdue, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections across the full stack — in Python for backend and ML workflows, and JavaScript (React, Node.js) for frontend and API layers, alongside C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code across both Python and JavaScript (React, Node.js), with additional work in C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code across backend and frontend contexts to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of the code and identify error patterns across full-stack applications. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in building full-stack applications using Python and JavaScript (React, Node.js), with the ability to work across backend and frontend codebases. • Experience deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Create and evaluate AI model training datasets and code solutions primarily in Python and JavaScript, collaborating with researchers and cross-functional teams. | Several years of software engineering experience, strong Python expertise, full-stack development skills, and excellent communication abilities. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have built production systems at companies like Google, Microsoft, Apple, Amazon, Meta, or similar high-scale engineering organizations. We especially welcome graduates from top computer science programs such as Stanford, MIT, Carnegie Mellon, UC Berkeley, Georgia Tech, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections — with a primary focus on Python across backend services, data pipelines, and ML infrastructure, alongside JavaScript (including ReactJS), C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code — primarily in Python, with additional work in JavaScript (including ReactJS), C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents and automated verification tools in Python that can verify the quality of code and identify error patterns. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in Python with deep knowledge of frameworks, tooling, and best practices for building production-grade software. • Experience building full-stack applications and deploying scalable software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Create and evaluate AI training datasets by curating and correcting code across multiple languages and collaborating with cross-functional teams to improve AI-driven coding solutions. | Several years of software engineering experience with strong full-stack skills in Python and JavaScript, expertise in software architecture and code quality assessment, and excellent communication skills. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have built production systems at companies like Google, Microsoft, Apple, Amazon, Meta, or similar high-scale engineering organizations. We especially welcome graduates from leading programs such as Harvard, Columbia, Princeton, Yale, University of Pennsylvania, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections across the full stack — in Python for backend and ML workflows, and JavaScript (React, Node.js) for frontend and API layers, alongside C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code across both Python and JavaScript (React, Node.js), with additional work in C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code across backend and frontend contexts to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of the code and identify error patterns across full-stack applications. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in building full-stack applications using Python and JavaScript (React, Node.js), with the ability to work across backend and frontend codebases. • Experience deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Evaluate and refine AI-generated code, curate code examples, and collaborate with teams to enhance AI-driven coding solutions. | 3+ years software engineering experience with strong full-stack development and communication skills. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections in Python, JavaScript (including ReactJS), C/C++, Java, Rust, and Go; evaluating and refining AI-generated code for efficiency, scalability, and reliability; and working with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Working on AI model training initiatives by curating code examples, building solutions, and correcting code in Python, JavaScript (including ReactJS), C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents that can verify the quality of the code and identify error patterns. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • 3+ years of software engineering experience. • Strong expertise in building full-stack applications and deploying scalable, production-grade software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week (partial PST overlap required) • Type: Contractor (no medical/paid leave) • Duration: 1 month (starting next week; potential extensions based on performance and fit) • Location: Candidates must be based out of US, Canada or WEU countries (UK, Netherlands, Italy, Germany, …) Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Create and evaluate AI training datasets and code solutions primarily in Python, collaborating with researchers and cross-functional teams. | Several years of software engineering experience with strong Python expertise, full-stack application development, and excellent communication skills. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have built production systems at companies like Google, Microsoft, Apple, Amazon, Meta, or similar high-scale engineering organizations. We especially welcome graduates from top computer science programs such as Stanford, MIT, Carnegie Mellon, UC Berkeley, Georgia Tech, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections — with a primary focus on Python across backend services, data pipelines, and ML infrastructure, alongside JavaScript (including ReactJS), C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code — primarily in Python, with additional work in JavaScript (including ReactJS), C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents and automated verification tools in Python that can verify the quality of code and identify error patterns. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in Python with deep knowledge of frameworks, tooling, and best practices for building production-grade software. • Experience building full-stack applications and deploying scalable software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Design, build, and maintain AI agents and large-scale knowledge graphs with advanced AI and data integration. | 5+ years software engineering with 2+ years GenAI experience, expertise in graph databases, LLM frameworks, and cloud deployment. | About Turing Based in San Francisco, California, Turing is the world's leading research accelerator for frontier AI labs and a trusted partner for global enterprises looking to deploy advanced AI systems. Turing accelerates frontier research with high-quality data, specialized talent, and training pipelines that advance thinking, reasoning, coding, multimodality, and STEM. For enterprises, Turing builds proprietary intelligence systems that integrate AI into mission-critical workflows, unlock transformative outcomes, and drive lasting competitive advantage. Recognized by Forbes, The Information, and Fast Company among the world's top innovators, Turing's leadership team includes AI technologists from Meta, Google, Microsoft, Apple, Amazon, McKinsey, Bain, Stanford, Caltech, and MIT. Learn more at www.turing.com Location Remote / Hybrid (HQ visits as needed) Experience 5 + years in software engineering; 2 + years in GenAI Engagement Full-Time, Permanent About the Role We are looking for a talented Sr. GenAI Engineer who sits at the intersection of knowledge engineering, agentic AI, and data intelligence. In this role you will design and operate AI agents that traverse, reason over, and enrich large-scale knowledge graphs — then extend that context dynamically using live data sources such as the web, enterprise APIs, and structured databases. The ideal candidate is deeply comfortable with graph data models, LLM orchestration frameworks, and retrieval-augmented pipelines. Bonus points if you have experience working in trade-craft or intelligence-adjacent environments where provenance, precision, and adversarial robustness are non-negotiable. Key Responsibilities Knowledge Graph Engineering • Design, build and maintain large-scale property graphs and RDF triplestores (Neo4j, Amazon Neptune, Stardog, or equivalent). • Develop and govern ontologies, taxonomies, and entity-relationship schemas that reflect real-world domain semantics. • Implement graph ingestion pipelines that extract, transform, and link entities from structured, semi-structured, and unstructured data. • Optimise graph traversal queries (Cypher, SPARQL, Gremlin) for sub-second response at production scale. • Train and deploy graph neural networks (GNNs) for node classification, link prediction, and subgraph retrieval - Maintain model retraining workflows triggered by graph drift or coverage degradation. Agentic AI Systems • Architect and implement autonomous agents that plan multi-step reasoning chains over knowledge graph data using LLMs (GPT-4o, Claude, Gemini, or open-source equivalents). • Build graph-aware Retrieval-Augmented Generation (RAG) pipelines that blend structured graph context with unstructured document retrieval. • Design tool-use and function-calling layers so agents can query live data sources — web search, REST/GraphQL APIs, relational databases — to extend or verify graph knowledge. • Implement agent memory, reflection, and self-correction loops to improve reliability over multi-hop tasks. Context Enrichment & Data Fusion • Integrate web scraping, news feeds, and open-source intelligence (OSINT) sources to keep the knowledge graph current. • Build entity resolution and deduplication components that merge data from heterogeneous sources into a consistent graph. • Develop confidence-scoring and provenance-tracking mechanisms so downstream consumers understand the reliability of any piece of context. MLOps & Production Readiness • Package agents as scalable microservices; instruments with observability tooling (tracing, latency, token cost). • Collaborate with platform engineers to deploy workloads on cloud-native infrastructure (AWS / GCP / Azure). • Maintain evaluation harnesses that measure agent accuracy, hallucination rate, and graph coverage over time. Required Skills & Experience • 5 + years of professional software engineering with strong Python (or Java / Kotlin) proficiency. • Hands-on production experience with at least one major graph database — Neo4j, Amazon Neptune, TigerGraph, or comparable. • Demonstrated knowledge of graph query languages like Cypher, SPARQL, or Gremlin — at production query complexity. • Direct experience building LLM-powered agents or pipelines using frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI, AutoGen, or Semantic Kernel. • Solid understanding of RAG architectures: chunking strategies, vector stores (Pinecone, Weaviate, pgvector), hybrid retrieval, and re-ranking. • Familiarity with prompt engineering, few-shot learning, and LLM evaluation techniques. • Experience integrating external data sources via APIs, web scraping (Playwright / Scrapy), or streaming pipelines (Kafka / Kinesis). • Working knowledge of containerisation (Docker, Kubernetes) and CI/CD pipelines. • Familiarity with graph export formats - at least one GraphML, RDF/OWL, or JSON-LD. • Experience integrating GNN-derived features into vector stores or RAG pipelines Preferred Qualifications • Advanced degree (MS / PhD) in Computer Science, Information Science, Computational Linguistics, or a related field. • Experience in intelligence, defence, or trade-craft environments — working with OSINT, link analysis, entity disambiguation, or signals intelligence data. • Understanding of access-control models for sensitive graph data (need-to-know, compartmentalisation, provenance labelling). • Familiarity with knowledge representation standards like OWL, SHACL, RDF-star, JSON-LD, W3C PROV. • Experience with fine-tuning or instruction-tuning open-source LLMs (Llama, Mistral, Falcon) for domain-specific tasks. • Background in network-analysis algorithms: centrality, community detection, path-finding, anomaly detection on graphs. • Contributions to open-source graph or GenAI projects; published research or technical blog presence. • Active or adjudicatable security clearance (Secret or above) — strongly preferred for trade-craft assignments. ★ Trade-Craft Experience — A Significant Plus Candidates with backgrounds in intelligence analysis, signals intelligence, law enforcement data fusion, or related trade-craft disciplines are strongly encouraged to apply. Understanding of link analysis, entity disambiguation under adversarial conditions, handling classified or compartmentalized data, and mission-driven product constraints will set you apart. Salary: $200,000 - $250,000 Values • We are client first: We put our clients at the center of everything we do, because their success is the ultimate measure of our value. • We work at Start-Up Speed: We move fast, stay agile and favor action because momentum is the foundation of perfection • We are AI forward: We help our clients build the future of Al and implement it in our own roles and workflow to amplify productivity. Advantages of joining Turing • Amazing work culture (Super collaborative & supportive work environment; 5 days a week) • Awesome colleagues (Surround yourself with top talent from Meta, Google, LinkedIn etc. as well as people with deep startup experience) • Competitive compensation • Flexible working hours Don't meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they meet every single qualification. Turing is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, disability, protected veteran status, or any other legally protected characteristics. At Turing we are dedicated to building a diverse, inclusive and authentic workplace and celebrate authenticity, so if you're excited about this role but your past experience doesn't align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles. For applicants from the European Union, please review Turing's GDPR notice here.
Design and optimize FastAPI services for reinforcement learning projects while collaborating with researchers. | Requires 5+ years Python experience, strong FastAPI skills, and software engineering best practices. | A leading AI solutions firm is seeking a Python Developer to work on a Reinforcement Learning Gym project. You will design and optimize FastAPI services, collaborate with researchers, and ensure high-quality API delivery. Ideal candidates have 5+ years in Python, strong FastAPI experience, and familiarity with software engineering best practices. This role offers fully remote work and a flexible schedule. #J-18808-Ljbffr
Review and evaluate model-generated code and provide feedback to improve AI-assisted development tools. | Several years of software engineering experience, particularly full-stack, with ability to assess code quality and work flexible part-time hours. | A leading AI company is seeking a software engineer to review and evaluate model-generated code. This contract role requires several years of software engineering experience, particularly as a full-stack engineer at notable tech firms. You will assess code quality and provide feedback, working 10–20 hours a week with flexible hours. Competitive compensation of $50–$150 per hour based on experience. Apply if you are ready to shape the future of AI-assisted development tools. #J-18808-Ljbffr
Evaluate and refine AI-generated code and build datasets and tools for AI model training. | Several years of software engineering experience with strong Python skills and full-stack development experience. | About Us: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L. Ideal Background: This role is ideal for engineers who have built production systems at companies like Google, Microsoft, Apple, Amazon, Meta, or similar high-scale engineering organizations. We especially welcome graduates from top computer science programs such as Stanford, MIT, Carnegie Mellon, UC Berkeley, Georgia Tech, and comparable institutions — though exceptional experience and skill always take precedence over pedigree. Project Overview: As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections — with a primary focus on Python across backend services, data pipelines, and ML infrastructure, alongside JavaScript (including ReactJS), C/C++, Java, Rust, and Go. You will evaluate and refine AI-generated code for efficiency, scalability, and reliability, and work with cross-functional teams to enhance enterprise-level AI-driven coding solutions. What Does a Typical Day Look Like? • Work on AI model training initiatives by curating code examples, building solutions, and correcting code — primarily in Python, with additional work in JavaScript (including ReactJS), C/C++, Java, Rust, and Go. • Evaluate and refine AI-generated code to ensure that it is efficient, scalable, and reliable. • Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks. • Build agents and automated verification tools in Python that can verify the quality of code and identify error patterns. • Hypothesize on steps in the software engineering cycle (prototyping, architecture design, API design, production implementation, launch, experiments, monitoring, operational maintenance) and evaluate model capabilities on them. • Design verification mechanisms that can automatically verify a solution to a software engineering task. Required Skills: • Several years of software engineering experience (3 years or more) • Strong expertise in Python with deep knowledge of frameworks, tooling, and best practices for building production-grade software. • Experience building full-stack applications and deploying scalable software using modern languages and tools. • Deep understanding of software architecture, design, development, debugging, and code quality/review assessment. • Excellent oral and written communication skills for clear, structured evaluation rationales. Engagement Details: • Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week • Type: Contractor (no medical/paid leave) • Duration: 1 month (potential extensions based on performance and fit) • Location: Candidates must be based in the United States Evaluation Process: • The application process takes 15–30 minutes. • Completion of an AI video interview is required. Note: As part of assessments you will go through an AI video interview. After applying, you will receive an email with a login link. Please use that link to access the portal and complete your profile. Know amazing talent? Refer them at turing.com/referrals, and earn money from your network.
Drive growth and deliver technically scoped programs inside frontier AI organizations by managing commercial and technical relationships, leading negotiations, and coordinating internal teams. | Experience owning strategic customer relationships with technical or research-driven AI/ML teams, ability to manage delivery of technical programs involving ML systems, strong understanding of generative AI systems and foundation model lifecycles, excellent communication skills, and location in SF Bay Area. | About Turing Based in Palo Alto, California, Turing is one of the world's fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems. Turing helps customers in two ways: working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilingualism, STEM and frontier knowledge; and leveraging that expertise to build real-world AI systems that solve mission-critical priorities for Fortune 500 companies and government institutions. Turing has received numerous awards, including Forbes's "One of America's Best Startup Employers," #1 on The Information's annual list of "Most Promising B2B Companies," and Fast Company's annual list of the "World's Most Innovative Companies." Turing's leadership team includes AI technologists from industry giants Meta, Google, Microsoft, Apple, Amazon, Twitter, McKinsey, Bain, Stanford, Caltech, and MIT. For more information on Turing, visit www.turing.com. For information on upcoming Turing AGI Icons events, visit go.turing.com/agi-icons. About the Role This is a strategic, customer-facing role responsible for driving growth and delivering technically scoped programs inside frontier AI organizations. You will work directly with model developers, research teams, and infrastructure leads at some of the most advanced labs in the world. Your work will shape how these organizations scale large models, evaluate system behavior, train agents, and improve safety. The scope includes both commercial and technical ownership. You’ll identify opportunities for collaboration, define the structure and terms of each engagement, and ensure successful delivery across internal and external teams. Success in this role depends on the ability to translate loosely defined technical objectives into executable programs — and to grow relationships by consistently delivering value in highly complex, evolving environments. This role is designed for someone who can navigate early-stage model development contexts, build trust with deeply technical teams, and drive meaningful commercial and strategic outcomes without relying on standard sales tactics. Responsibilities • Manage commercial and technical relationships with frontier AI labs developing foundation models and related systems. Own overall account growth, strategic alignment, and delivery success. • Identify areas of opportunity within research and model development pipelines — including evaluation, safety, agentic workflows, or data generation — and structure these into scoped engagements. • Lead commercial negotiations and close new programs that align with both customer priorities and Turing’s AGI infrastructure. • Translate ambiguous or early-stage technical objectives into executable workstreams with clear scope, milestones, and delivery plans. • Coordinate internal teams across product, engineering, operations, and leadership to ensure execution quality and responsiveness to shifting technical requirements. • Maintain direct oversight of program performance across multiple accounts. Monitor delivery health, preempt blockers, and guide iteration based on real-time feedback. • Capture customer insights and organizational signals to inform how Turing evolves its offering for model developers. Contribute to internal roadmaps, service design, and engagement playbooks. • Build and maintain technical credibility with research, infrastructure, and product leaders at customer organizations. Engage fluently in conversations around model scaling, safety, benchmarking, and deployment. Requirements • Demonstrated experience owning strategic customer relationships with technical or research-driven teams in AI, ML infrastructure, or model development environments. • Proven ability to originate and expand high-value programs in complex, multi-stakeholder accounts. • Skilled in scoping and managing delivery of technical programs involving ML systems, data pipelines, model evaluation, or applied research. • Strong understanding of generative AI systems, foundation model lifecycles, and LLM-adjacent workflows such as alignment, evaluation, and data efficiency. • Able to navigate ambiguity, define structure where none exists, and drive alignment across internal and external stakeholders. • Confident communicator capable of engaging C-level sponsors, research leads, and engineering managers with clarity and precision. • Highly organized and execution-oriented. Comfortable managing concurrent workstreams with attention to detail, delivery health, and long-term growth strategy. • Location in the SF Bay Area is required. Advantages of joining Turing: • Amazing work culture (Super collaborative & supportive work environment; 5 days a week) • Awesome colleagues (Surround yourself with top talent from Meta, Google, LinkedIn etc. as well as people with deep startup experience) • Competitive compensation • Flexible working hours • Full-time remote opportunity Don’t meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they meet every single qualification. Turing is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, disability, protected veteran status, or any other legally protected characteristics. At Turing we are dedicated to building a diverse, inclusive and authentic workplace and celebrate authenticity, so if you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles. For applicants from the European Union, please review Turing's GDPR notice here.
Design and evaluate advanced math problems to improve AI language model performance through collaboration with researchers. | PhD in Mathematics or related field with strong math reasoning and clear communication skills. | Job Description Remote contract for PhDs in Mathematics, Statistics, or related fields. Work on cutting-edge projects with top AI labs while earning $50+/hour, fully remote, with flexible weekly hours. Role Overview: Help fine-tune large language models (like ChatGPT) using your math and analytical skills. You’ll design problems, check how well AI solves them, and work with researchers to build better benchmarks. Responsibilities: • Design advanced math problems to test AI performance (e.g., multi-step reasoning, abstraction, symbolic manipulation). • Develop clear, step-by-step solutions with rigorous logic. • Evaluate AI outputs for accuracy and quality of reasoning. • Collaborate with researchers to refine benchmarks across undergraduate to PhD-level math topics. Requirements: • PhD (pursuing or completed) in Mathematics, Applied Math, Statistics, or related field. • Strong mathematical reasoning and problem-solving skills across advanced domains. • Ability to communicate complex ideas clearly in writing and provide structured feedback. Perks: • Fully remote, flexible work. • Work on cutting-edge AI projects with leading LLM companies. Offer Details: • Pay rate: $50+/hour (depends on role and candidate expertise) • Assessment: Shortlisted experts complete an evaluation before selection. • Assignments: Contract roles with defined start/end dates; up to 40 hrs/week. About Turing: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.
Lead large teams of software engineers and data scientists to deliver high-quality LLM training datasets through supervised fine tuning and reinforcement learning. | Proven experience managing large technical teams, strong leadership and communication skills, and hands-on ability with Python and JavaScript. | About Turing Based in San Francisco, California, Turing is the world's leading research accelerator for frontier AI labs and a trusted partner for global enterprises looking to deploy advanced AI systems. Turing accelerates frontier research with high-quality data, specialized talent, and training pipelines that advance thinking, reasoning, coding, multimodality, and STEM. For enterprises, Turing builds proprietary intelligence systems that integrate AI into mission-critical workflows, unlock transformative outcomes, and drive lasting competitive advantage. Recognized by Forbes, The Information, and Fast Company among the world's top innovators, Turing's leadership team includes AI technologists from Meta, Google, Microsoft, Apple, Amazon, McKinsey, Bain, Stanford, Caltech, and MIT. Learn more at www.turing.com Position Overview As a Delivery Leader for our LLM Training business, you will play a pivotal role in leading our efforts to develop high-quality, foundational LLMs. This position requires managing large teams of software engineers and data scientists dedicated to performing various Supervised Fine Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) tasks. Your leadership will ensure the delivery of superior SFT and RLHF datasets, achieving optimal outcomes in terms of quality, throughput, and cost. Key Responsibilities • Lead and manage large teams of Python, JavaScript, Java (and other long tail of languages) developers to execute technical tasks effectively. • Collaborate closely with researcher clients, ensuring their satisfaction with the delivered work. • Maintain rigorous review processes to ensure the highest quality of datasets. • Oversee the transition of projects from initiation to stable states, taking full responsibility for delivery • Identify and implement measures to enhance the quality of Python/JavaScript/Java etc. code within the team. • Ensure the team's work is of high quality, addressing any issues related to the clarity of questions, methodology, result communication, or bugs. Qualifications • Proven experience in managing large technical teams in a delivery-oriented role. • At least one of the following, ideally both: • * Python • Javascript • Demonstrated ability to engage in hands-on technical work, identifying and resolving quality issues. • Excellent leadership and people management skills, with a focus on motivating teams and fostering a collaborative work environment. • Strong communication and stakeholder management abilities, capable of effectively collaborating with clients and internal partners. Why Join Us? Our company is at the cutting edge of AI and Machine Learning, offering unique opportunities to contribute to the advancement of LLMs. You will lead a team of talented individuals, collaborate with top-tier clients, and be part of a dynamic and innovative culture. We are committed to the professional growth of our employees, providing a supportive environment where you can thrive and make a significant impact in the AI industry. Values: • We are client first: We put our clients at the center of everything we do, because their success is the ultimate measure of our value. • We work at Start-Up Speed: We move fast, stay agile and favor action because momentum is the foundation of perfection • We are Al forward: We help our clients build the future of Al and implement it in our own roles and workflow to amplify productivity. Advantages of joining Turing: • Amazing work culture (Super collaborative & supportive work environment; 5 days a week) • Awesome colleagues (Surround yourself with top talent from Meta, Google, LinkedIn etc. as well as people with deep startup experience) • Competitive compensation • Flexible working hours Don't meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they meet every single qualification. Turing is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, disability, protected veteran status, or any other legally protected characteristics. At Turing we are dedicated to building a diverse, inclusive and authentic workplace and celebrate authenticity, so if you're excited about this role but your past experience doesn't align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles. For applicants from the European Union, please review Turing's GDPR notice here.
Design and evaluate advanced physics problems to test AI performance and collaborate with researchers to refine AI benchmarks. | PhD in Physics or related field with strong physics reasoning and problem-solving skills and ability to communicate complex ideas clearly. | Remote contract for PhDs in Physics, Applied Physics, or related fields. Work on cutting-edge projects with top AI labs while earning $50+/hour, fully remote, with flexible weekly hours. Role Overview: Help fine-tune large language models (like ChatGPT) using your physics skills. You’ll design problems, check how well AI solves them, and work with researchers to build better benchmarks. Responsibilities: • Design advanced physics problems to test AI performance (e.g., mechanics, electromagnetism, thermodynamics). • Develop clear, step-by-step solutions with rigorous logic. • Evaluate AI outputs for accuracy and quality of reasoning. • Collaborate with researchers to refine benchmarks across undergraduate to PhD-level physics topics. Requirements: • PhD (pursuing or completed) in Physics, Applied Physics, or a related field. • Strong physics reasoning and problem-solving skills across advanced domains. • Ability to communicate complex ideas clearly in writing and provide structured feedback. Perks: • Fully remote, flexible work. • Work on cutting-edge AI projects with leading LLM companies. Offer Details: • Pay rate: $50+/hour (depends on role and candidate expertise) • Assessment: Shortlisted experts complete an evaluation before selection. • Assignments: Contract roles with defined start/end dates; up to 40 hrs/week. About Turing: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.
Create tailored applications specifically for Turing with our AI-powered resume builder
Get Started for Free