745,032open jobs
44,747companies
107,631added this week
Browse all
Salary
$250k – $400k per year
Location
In office (Palo Alto)
Seniority
Staff

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026. Turing scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Turing is an AI company headquartered in San Francisco that supplies frontier AI labs with training datasets, reinforcement learning environments, research benchmarks and specialised human expertise in coding, reasoning, STEM and multimodal work. Founded in 2018, it describes itself as the largest and longest-running data provider for software-engineering model training and also builds agentic AI systems for enterprises in financial services, life sciences, healthcare, retail and automotive. It recruits research and AI engineers, strategic project leads, DevOps and IT engineers, product and marketing managers, recruiters and strategy and operations staff, many of them remote.

About Turing

Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com. 

The Role

Turing is seeking exceptional Staff Research Scientists to join our STEM research organization and develop new ways to evaluate, train, and improve frontier AI systems.

This is a research-first role focused on problems where the right benchmark, dataset, or methodology often does not yet exist. You will identify important gaps in the literature, propose ambitious new research directions, and take projects from initial hypothesis through experimentation, benchmark construction, and publication.

Our research is deliberately focused on frontier STEM evaluation, synthetic data, hallucination and reliability, and agentic science. We are looking for scientists who can recognize important problems early, formulate them precisely, and design rigorous research programs to answer them.

What You'll Do 

Frontier benchmarks and evaluation

  • Identify high-impact gaps in existing benchmark and evaluation literature.
  • Design novel benchmarks in and across STEM fields and on general model functionality.
  • Develop evaluations for emerging model capabilities that are poorly captured by traditional static benchmarks.
  • Design rigorous task-generation, grading, contamination-control, difficulty-calibration, and validation methodologies.
  • Build benchmarks that can become both valuable research contributions and meaningful standards for evaluating frontier models.

Synthetic data and post-training

  • Develop methods for generating high-quality synthetic STEM training data.
  • Study how task selection, difficulty, diversity, verification, filtering, and data quality affect downstream performance.
  • Explore methods for generating useful training signal in domains where expert human data is scarce or expensive.
  • Design experiments that determine when synthetic data genuinely improves capabilities rather than simply increasing training volume.

Hallucination, reliability, and verification

  • Study hallucination, uncertainty, calibration, and epistemic failure in technical domains.
  • Develop evaluations and methods for improving factual reliability, self-correction, verification, citation, and appropriate abstention.
  • Investigate when models should reason internally, invoke tools, seek external evidence, or recognize that they do not know.

Agentic science

  • Research AI systems capable of performing extended scientific and technical work.
  • Develop workflows involving literature search, coding, simulation, tool use, experimentation, verification, and iterative reasoning.
  • Evaluate long-horizon scientific agents and identify the bottlenecks preventing them from reliably performing real research.
  • Explore new approaches to human-AI and multi-agent scientific collaboration.

New research directions

The areas above are our core focus, not an exhaustive list. Researchers will also have significant latitude to propose new programs in areas such as reasoning, model evaluation, AI-for-science, data generation, and emerging capabilities.

What We’re Looking For

  • PhD or equivalent research experience in machine learning, computer science, mathematics, physics, chemistry, biology, engineering, statistics, or another highly technical field.
  • Demonstrated ability to formulate and execute original research.
  • Strong understanding of modern LLMs and the frontier AI research landscape.
  • Excellent experimental design, quantitative reasoning, and scientific judgment.
  • Ability to rapidly understand unfamiliar technical literature and develop expertise in new areas.
  • Strong Python skills and the ability to independently build research prototypes and evaluation pipelines.
  • Excellent technical writing and communication.
  • Comfort working in a fast-moving environment where the research agenda evolves with the frontier.

A strong publication record is valuable, but we care most about whether you can identify important questions, design rigorous ways to answer them, and execute quickly enough for the results to matter.

What Success Looks Like

You might:

  • Identify a major capability that existing benchmarks fail to measure and create the benchmark that becomes the standard for evaluating it.
  • Discover a failure mode in current synthetic-data pipelines and develop a method that materially improves post-training.
  • Build a new evaluation that changes how frontier labs understand hallucination, reasoning, or scientific capability.
  • Develop an agentic workflow that substantially advances performance on complex scientific research tasks.
  • Launch an entirely new research direction that grows into a major program within Turing.

Why Turing 

Frontier models are improving faster than the benchmarks, datasets, and research methodologies used to understand them.

The STEM research team at Turing works on that gap directly. Our goal is not simply to apply existing evaluation methods, but to invent the benchmarks, data-generation methods, and research frameworks needed for the next generation of AI systems.

If you want to define how frontier AI is evaluated and improved across science and technical reasoning, we’d like to hear from you.

This role is required to be in office five days a week, based in any of Turing's offices in San Francisco, Palo Alto, or Seattle.

Compensation: $250,000 to $400,000 OTE + Equity

Values

  • We are client first: We put our clients at the center of everything we do, because their success is the ultimate measure of our value.
  • We work at Start-Up Speed: We move fast, stay agile and favor action because momentum is the foundation of perfection
  • We are AI forward: We help our clients build the future of Al and implement it in our own roles and workflow to amplify productivity.

Advantages of joining Turing

  • Work at the frontier of AI, helping the world’s leading AI labs improve their most advanced models by building expert datasets, RL environments, and first-of-a-kind benchmarks.
  • Contribute to leading-edge AI research and showcase your work at top conferences such as ICLR, ICML, and NeurIPS.
  • Bring frontier AI innovation to the enterprise, applying lessons learned from leading AI labs to solve real-world business challenges.
  • Collaborate with and learn from exceptional colleagues with deep AI experience from Google, Meta, Amazon, and other leading technology companies.
  • Move at the pace of AI innovation, with the speed, ownership, and impact of a startup.

Turing is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, disability, protected veteran status, or any other legally protected characteristics. At Turing we are dedicated to building a diverse, inclusive and authentic workplace  and celebrate authenticity, so if you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles.

For applicants from the European Union, please review  Turing's GDPR notice here.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
745,032 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Palo Alto
$200k – $240k per year • Remote (United States) • 5+ years exp
AI/ML
Claude
Model Context Protocol
AI Agents
LLM
OpenAI Agents SDK
DevOps
Docker
Linux
Management
Slack
Apply
≈ $145k – $253k per year (Estimated) • Equity • Remote (United States) • Full-Time • 5+ years exp • United States
Python
Go
AI/ML
AI Agents
LLM
Knowledge Graph
Agentic Workflows
DevOps
Kubernetes
Platform Engineering
Cybersecurity
Crowdstrike
Analytics
A/B Testing
Apply
≈ $137k – $250k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • PhD • United States
Python
Java
Java
Spring Boot
AI/ML
Machine Learning
DevOps
GCP
GitHub Actions
Azure
CI/CD
Jenkins
AWS
Twelve-Factor App
Apply
$124k – $211k per year • In office • Full-Time • 7+ years exp • High School Diploma • Home
AI/ML
Copilot
AI Agents
Machine Learning
DevOps
Azure
Apply
CXO AI Engineer 3 hours ago
≈ $152k – $265k per year (Estimated) • Equity • Remote (United States) • 5+ years exp
Python
SQL
Databases
Amazon Redshift
Trino
AI/ML
Model Context Protocol
dbt
Function Calling
AI Agents
LLM
RAG
Context Engineering
Tool Use
DevOps
SLI/SLO/SLA
Management
Slack
Apply
≈ $73k – $138k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Toronto
Python
Go
JavaScript
Rust
TypeScript
PowerShell
Databases
Snowflake
Databricks
Delta Lake
Apache Kafka
AI/ML
LangGraph
AutoGen
LangChain
Spark
Embeddings
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Semantic Search
LLMOps
Context Engineering
Semantic Search
LLM Guardrails
Agentic Workflows
Tool Use
Machine Learning
DevOps
GCP
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Spinnaker
Cybersecurity
SIEM
Apply
Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
Python
PowerShell
DevOps
Rest API
Windows Server
Cybersecurity
ISO 27001
PCI DSS
Active Directory
PKI
Management
ServiceNow
Apply
≈ $33k – $73k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • High School Diploma • Bengaluru
Python
PowerShell
DevOps
Rest API
Splunk
Terraform
GCP
Azure
CI/CD
AWS
IAM
Cybersecurity
Okta
Crowdstrike
CyberArk
Qualys Cloud Platform
Zero Trust
Microsoft Entra ID
Active Directory
LDAP
Management
ServiceNow
Apply
≈ $21k – $41k per year (Estimated) • In office • Internship • São Paulo
Python
JavaScript
SQL
C#
Databases
MySQL
Apply
≈ $15k – $42k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Pune • Bengaluru
Python
SQL
AI/ML
Machine Learning
Apply
$250k – $400k per year • In office • Master's Degree • Palo Alto
AI/ML
Model Context Protocol
Reinforcement Learning
AI Agents
Post-training
Edge AI
Computer Use
Machine Learning
Cybersecurity
GDPR
Management
Agile
Apply
$250k – $350k per year • In office • Master's Degree • Palo Alto
AI/ML
Model Context Protocol
Reinforcement Learning
AI Agents
Post-training
Edge AI
Computer Use
Machine Learning
Cybersecurity
GDPR
Management
Agile
Apply
$250k – $350k per year • In office • PhD • Palo Alto
Python
AI/ML
Reinforcement Learning
Function Calling
AI Agents
Hallucination
Synthetic Data
Post-training
Edge AI
Agentic Workflows
Tool Use
Machine Learning
Cybersecurity
GDPR
Management
Agile
Apply
$150k – $300k per year • In office • PhD • Palo Alto
Python
AI/ML
Reinforcement Learning
Function Calling
AI Agents
Hallucination
Synthetic Data
Post-training
Edge AI
Agentic Workflows
Tool Use
Machine Learning
Cybersecurity
GDPR
Management
Agile
Apply
$250k – $350k per year • In office • Master's Degree • New York
AI/ML
Model Context Protocol
Reinforcement Learning
AI Agents
Post-training
Edge AI
Computer Use
Machine Learning
Cybersecurity
GDPR
Management
Agile
Apply
Senior ML Engineer 10 hours ago
$188k – $200k per year • Hybrid • 5+ years exp • Master's Degree • Palo Alto
Python
C++
AI/ML
Fine-tuning
AI Agents
Amazon SageMaker
Machine Learning
DevOps
GCP
AWS
Apply
$440k per year • In office • Palo Alto
Python
Rust
C++
AI/ML
Reinforcement Learning
LLM
Apply
$440k per year • In office • Palo Alto
Python
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
SGLang
PyTorch
LLM
CUDA
Apply
$600k per year • In office • Palo Alto
AI/ML
RLHF
Reinforcement Learning
DPO
Post-training
Reward Modeling
Apply
$440k per year • In office • Palo Alto
Python
Rust
C++
C++
PyTorch C++
AI/ML
Spark
Fine-tuning
JAX
Multimodal AI
Function Calling
AI Agents
PyTorch
Ray
Tokenization
SFT
Post-training
Pre-training
TPU
XLA
Tool Use
Reward Modeling
DevOps
Kubernetes
Apply
See all jobs
This is one of many
745,032 more open roles from verified company boards, updated every day.