690,898open jobs
40,529companies
98,039added this week
Browse all
Salary
$150k – $300k per year
Location
Remote/Hybrid (United States)
Seniority
Senior · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Centific is an American company spun out of the localisation group Pactera EDGE that supplies the data, evaluation and engineering services artificial intelligence systems need in production. Its work covers multilingual training data collection and annotation, human evaluation of model outputs, red teaming, and the engineering to deploy models into enterprise and retail environments, drawing on a contributor network spanning many languages and markets. Headquartered in Redmond, Washington, it works with large technology companies and enterprises, and its heritage in localisation is what gave it the global human capacity this work requires.

About Centific

Centific is a frontier AI data foundry that curates diverse, high-quality data, using our purpose-built technology platforms to empower the Magnificent Seven and our enterprise clients with safe, scalable AI deployment. Our team includes more than 150 PhDs and data scientists, along with more than 4,000 AI practitioners and engineers. We harness the power of an integrated solution ecosystem-comprising industry-leading partnerships and 1.8 million vertical domain experts in more than 230 markets-to create contextual, multilingual, pre-trained datasets; fine-tuned, industry-specific LLMs; and RAG pipelines supported by vector databases. Our zero-distance innovation™ solutions for GenAI can reduce GenAI costs by up to 80% and bring solutions to market 50% faster.

Our mission is to bridge the gap between AI creators and industry leaders by bringing best practices in GenAI to unicorn innovators and enterprise customers. We aim to help these organizations unlock significant business value by deploying GenAI at scale, helping to ensure they stay at the forefront of technological advancement and maintain a competitive edge in their respective markets.

About Job

Role: Senior Applied Reinforcement Learning Engineer

Location: Palo Alto, CA or Seattle, WA (Hybrid/Remote)

About the Team

Centific AI Research advances foundational AI models and applications through reinforcement learning, alignment, and human-centered intelligence. Our mission is to transform data, signals, and human insight into next-generation intelligent systems that redefine enterprise intelligence.

We're building a governed RL environment platform that enables enterprises to safely iterate and improve AI agent workflows through simulation-based learning, bridging human-labeled signal creation with automated RL training for high-stakes operations.

Role Overview

As an Applied RL Engineer, you will design and build RL environments that simulate complex enterprise workflows and train intelligent agents within them. You'll work at the intersection of RL research and production systems, translating customer requirements into bespoke simulation environments and post-training pipelines that deliver measurable improvements to AI agent performance.

This role requires deep expertise in both classical RL methodologies and modern LLM-based agent architectures. You'll shape our product direction and help make RL accessible to enterprise customers who need safe, compliant ways to improve their AI systems.

Core RL Competencies

Foundational RL

  • MDPs & value methods: State/action spaces, Q-learning, DQN, Double DQN, Dueling DQN
  • Policy gradient methods: REINFORCE, Actor-Critic, A2C/A3C, variance reduction
  • Advanced optimization: PPO, TRPO, SAC, trust regions, entropy regularization
  • TD learning: TD(0), TD(λ), eligibility traces, bootstrapping methods

LLM Alignment & Post-Training

  • RLHF pipelines: Reward model training, preference learning, human feedback integration
  • Direct optimization: DPO, IPO, KTO, offline preference optimization
  • Group-based methods: GRPO, RLOO, sample-efficient policy improvement
  • Reward modeling: Bradley-Terry models, reward hacking mitigation, KL constraints

Environment Design

  • Gymnasium/OpenAI Gym: Custom environments, observation/action spaces, wrapper patterns
  • Reward engineering: Sparse vs. dense rewards, potential-based shaping, intrinsic motivation
  • Verifier design: Programmatic reward functions, outcome verification, ground-truth evaluation
  • Simulation: Sim-to-real transfer, domain randomization, multi-agent dynamics

Advanced Techniques

  • Offline RL: CQL, BCQ, IQL for learning from fixed datasets without environment interaction
  • Model-based RL: World models, Dreamer, MuZero, learned dynamics
  • Hierarchical RL: Options framework, goal-conditioned policies, temporal abstraction
  • Imitation & exploration: Behavioral cloning, GAIL, curiosity-driven exploration, UCB

Key Responsibilities

  • Design and build custom RL environments (digital twins) simulating enterprise workflows: document processing, compliance, onboarding, support automation
  • Post-train LLM-based agents on domain-specific tasks using PPO, GRPO, DPO, and RLHF
  • Build end-to-end pipelines converting human-labeled traces into RL training data
  • Architect multi-step reasoning agents with tool-calling and closed learning loops
  • Design reward functions, verifiers, and validation frameworks for pre-deployment testing
  • Translate cutting-edge RL research into production systems; contribute to publications

Required Qualifications

  • Deep RL expertise: 3+ years hands-on experience with environment design, reward engineering, policy optimization
  • LLM post-training: Experience fine-tuning LLMs using RLHF, DPO, PPO, or similar
  • Production skills: Software engineering beyond research with scalable pipelines and training infrastructure
  • Agentic AI: Experience with LLM-based agents, tool use, multi-step reasoning
  • Technical stack: Strong Python; Gymnasium, RLlib, Stable Baselines; PyTorch/JAX/TensorFlow
  • Education: MS/PhD in CS, ML, or related field (or equivalent experience)

Preferred Qualifications

  • Publications at NeurIPS, ICML, ICLR, ACL, or similar venues
  • Enterprise workflow experience in healthcare, finance, logistics, or compliance
  • Open-source contributions to CleanRL, TRL, veRL, or agent frameworks
  • Experience with world models, synthetic data generation, and simulation
  • Distributed training and large-scale RL experimentation

Why Join Centific

  • Lead the frontier: Shape a new discipline at the intersection of RL, simulation, and enterprise AI
  • Ship your science: See your research power real systems across healthcare, finance, and safety
  • Collaborate with leaders: Work alongside NVIDIA, Microsoft, and the global AI community
  • Build what matters: Create governed, compliant AI systems enterprises can trust.

Salary: $150K - $300K Annually

Centific is an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy), gender identity or expression, sexual orientation, marital status, familial status, veteran status, or any other characteristic protected by applicable law. We consider qualified applicants regardless of criminal histories, consistent with legal requirements.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
690,898 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
United States
$21k – $48k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Python
Java
C#
C#
.NET
Databases
Weaviate
Pinecone
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
Prompt Engineering
Function Calling
Chain-of-Thought
AI Agents
Semantic Kernel
CrewAI
Gemini
LLM
RAG
Semantic Search
GPT-4
Structured Outputs
Semantic Search
LLM Guardrails
Agentic Workflows
Tool Use
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
CI/CD
GitOps
Git
AWS
Docker
Kubernetes
Vector
Management
ServiceNow
Apply
$16k – $38k per year (Estimated) • In office • Full-Time • 1+ year exp • Noida
Python
Java
C#
C#
.NET
Databases
Weaviate
Pinecone
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
Prompt Engineering
Function Calling
Chain-of-Thought
AI Agents
Semantic Kernel
CrewAI
Gemini
LLM
RAG
Semantic Search
GPT-4
Structured Outputs
Semantic Search
LLM Guardrails
Agentic Workflows
Tool Use
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
CI/CD
GitOps
Git
AWS
Docker
Kubernetes
Vector
Management
ServiceNow
Apply
In office
Python
SQL
Python
Flask
AI/ML
YOLO
Fine-tuning
Reinforcement Learning
Scikit-learn
Computer Vision
NLP
Transfer Learning
ONNX
Transformers
TensorFlow
Keras
PyTorch
LLM
Semantic Search
CNN
Explainable AI
Sentiment Analysis
Hugging Face
Semantic Search
Recommender Systems
Machine Learning
DevOps
Jenkins
GitHub
Robotics
Path Planning
Analytics
Tableau
ETL/ELT
Apply
In office • Bachelor's Degree
Python
Java
AI/ML
Airflow
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
AWS Bedrock
LLM
Tokenization
Semantic Search
OCR
Semantic Search
Knowledge Graph
Tool Use
DevOps
AWS
Amazon EC2
Vector
Amazon S3
IAM
Windows
Analytics
ETL/ELT
Apache NiFi
Apply
$93k – $205k per year (Estimated) • In office • Full-Time • 3+ years exp • Associate's Degree • Chicago • Arlington • San Francisco • Dallas • Denver
AI/ML
AI Agents
Apply
$140k – $175k per year • Remote/Hybrid • Full-Time • Redmond • Palo Alto
AI/ML
Reinforcement Learning
AI Agents
RAG
Vision-Language-Action
World Models
Physical AI
Apply
Remote • Full-Time • India
AI/ML
RAG
Physical AI
Apply
$100k per year • Remote • Full-Time • Bachelor's Degree • United States
Python
Bash
AI/ML
RAG
DevOps
Terraform
Ansible
GitLab CI
CI/CD
Incident Management
Linux
TCP/IP
DNS
VLAN
Apply
$50k per year • Remote • Part-Time • 2+ years exp • United States
AI/ML
RAG
Apply
$46k – $117k per year (Estimated) • Remote • Full-Time • 10+ years exp • Serbia
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Databases
Azure Cosmos DB
Azure SQL Database
AI/ML
RAG
DevOps
Terraform
Azure DevOps
Azure
CI/CD
Git
Docker
Kubernetes
Bicep
Azure AKS
Vector
Cybersecurity
HIPAA
Microsoft Entra ID
Apply
$210k – $331k per year • Remote • Full-Time • 3+ years exp • PhD • Arlington • Austin • Dallas • Houston • San Antonio
DevOps
GCP
Cybersecurity
HIPAA
Management
Microsoft Office
Apply
$210k – $331k per year • Remote • Full-Time • 3+ years exp • PhD • Arlington • Charlotte • Columbia • Baltimore
DevOps
GCP
Cybersecurity
HIPAA
Management
Microsoft Office
Apply
$97k – $195k per year • In office • Full-Time • 7+ years exp • Charlotte • Plano • Chicago
Cybersecurity
NIST CSF
Apply
$87k per year • Remote/Hybrid • Full-Time • 4+ years exp • High School Diploma • Los Angeles • San Francisco • Pleasanton • Charlotte • New York
Apply
$75k – $80k per year • Equity • Remote • Full-Time • 3+ years exp • Bachelor's Degree • United States
C#
C#
.NET
Apply
See all jobs
This is one of many
690,898 more open roles from verified company boards, updated every day.