368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$150k – $300k per year
Location
Remote/Hybrid (United States)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Centific is a global digital engineering and artificial intelligence services company headquartered in Redmond, Washington, and established in its current form in 2020. The company provides data-centric AI solutions, including data collection, curation, reinforcement learning from human feedback, and multilingual data services through its OneForma platform. It operates globally with a network of over one million experts and serves enterprise clients in sectors such as technology, healthcare, and retail to build and scale production-grade AI models.

About Centific

Centific is a frontier AI data foundry that curates diverse, high-quality data, using our purpose-built technology platforms to empower the Magnificent Seven and our enterprise clients with safe, scalable AI deployment. Our team includes more than 150 PhDs and data scientists, along with more than 4,000 AI practitioners and engineers. We harness the power of an integrated solution ecosystem-comprising industry-leading partnerships and 1.8 million vertical domain experts in more than 230 markets-to create contextual, multilingual, pre-trained datasets; fine-tuned, industry-specific LLMs; and RAG pipelines supported by vector databases. Our zero-distance innovation™ solutions for GenAI can reduce GenAI costs by up to 80% and bring solutions to market 50% faster.

Our mission is to bridge the gap between AI creators and industry leaders by bringing best practices in GenAI to unicorn innovators and enterprise customers. We aim to help these organizations unlock significant business value by deploying GenAI at scale, helping to ensure they stay at the forefront of technological advancement and maintain a competitive edge in their respective markets.

About Job

Role: Applied Reinforcement Learning Engineer

Location: Palo Alto, CA or Seattle, WA (Hybrid/Remote)

About the Team

Centific AI Research advances foundational AI models and applications through reinforcement learning, alignment, and human-centered intelligence. Our mission is to transform data, signals, and human insight into next-generation intelligent systems that redefine enterprise intelligence.

We're building a governed RL environment platform that enables enterprises to safely iterate and improve AI agent workflows through simulation-based learning, bridging human-labeled signal creation with automated RL training for high-stakes operations.

Role Overview

As an Applied RL Engineer, you will design and build RL environments that simulate complex enterprise workflows and train intelligent agents within them. You'll work at the intersection of RL research and production systems, translating customer requirements into bespoke simulation environments and post-training pipelines that deliver measurable improvements to AI agent performance.

This role requires deep expertise in both classical RL methodologies and modern LLM-based agent architectures. You'll shape our product direction and help make RL accessible to enterprise customers who need safe, compliant ways to improve their AI systems.

Core RL Competencies

Foundational RL

  • MDPs & value methods: State/action spaces, Q-learning, DQN, Double DQN, Dueling DQN
  • Policy gradient methods: REINFORCE, Actor-Critic, A2C/A3C, variance reduction
  • Advanced optimization: PPO, TRPO, SAC, trust regions, entropy regularization
  • TD learning: TD(0), TD(λ), eligibility traces, bootstrapping methods

LLM Alignment & Post-Training

  • RLHF pipelines: Reward model training, preference learning, human feedback integration
  • Direct optimization: DPO, IPO, KTO, offline preference optimization
  • Group-based methods: GRPO, RLOO, sample-efficient policy improvement
  • Reward modeling: Bradley-Terry models, reward hacking mitigation, KL constraints

Environment Design

  • Gymnasium/OpenAI Gym: Custom environments, observation/action spaces, wrapper patterns
  • Reward engineering: Sparse vs. dense rewards, potential-based shaping, intrinsic motivation
  • Verifier design: Programmatic reward functions, outcome verification, ground-truth evaluation
  • Simulation: Sim-to-real transfer, domain randomization, multi-agent dynamics

Advanced Techniques

  • Offline RL: CQL, BCQ, IQL for learning from fixed datasets without environment interaction
  • Model-based RL: World models, Dreamer, MuZero, learned dynamics
  • Hierarchical RL: Options framework, goal-conditioned policies, temporal abstraction
  • Imitation & exploration: Behavioral cloning, GAIL, curiosity-driven exploration, UCB

Key Responsibilities

  • Design and build custom RL environments (digital twins) simulating enterprise workflows: document processing, compliance, onboarding, support automation
  • Post-train LLM-based agents on domain-specific tasks using PPO, GRPO, DPO, and RLHF
  • Build end-to-end pipelines converting human-labeled traces into RL training data
  • Architect multi-step reasoning agents with tool-calling and closed learning loops
  • Design reward functions, verifiers, and validation frameworks for pre-deployment testing
  • Translate cutting-edge RL research into production systems; contribute to publications

Required Qualifications

  • Deep RL expertise: 3+ years hands-on experience with environment design, reward engineering, policy optimization
  • LLM post-training: Experience fine-tuning LLMs using RLHF, DPO, PPO, or similar
  • Production skills: Software engineering beyond research with scalable pipelines and training infrastructure
  • Agentic AI: Experience with LLM-based agents, tool use, multi-step reasoning
  • Technical stack: Strong Python; Gymnasium, RLlib, Stable Baselines; PyTorch/JAX/TensorFlow
  • Education: MS/PhD in CS, ML, or related field (or equivalent experience)

Preferred Qualifications

  • Publications at NeurIPS, ICML, ICLR, ACL, or similar venues
  • Enterprise workflow experience in healthcare, finance, logistics, or compliance
  • Open-source contributions to CleanRL, TRL, veRL, or agent frameworks
  • Experience with world models, synthetic data generation, and simulation
  • Distributed training and large-scale RL experimentation

Why Join Centific

  • Lead the frontier: Shape a new discipline at the intersection of RL, simulation, and enterprise AI
  • Ship your science: See your research power real systems across healthcare, finance, and safety
  • Collaborate with leaders: Work alongside NVIDIA, Microsoft, and the global AI community
  • Build what matters: Create governed, compliant AI systems enterprises can trust.

Salary: $150K - $300K Annually

Centific is an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy), gender identity or expression, sexual orientation, marital status, familial status, veteran status, or any other characteristic protected by applicable law. We consider qualified applicants regardless of criminal histories, consistent with legal requirements.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
United States
LLM Model Developer 2 hours ago
$28k – $77k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad
AI/ML
AI Agents
Fine-tuning
LLM
Apply
$33k – $78k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp • Bachelor's Degree • India
Apex
JavaScript
Python
TypeScript
Apex
Copado
Lightning Web Components
AI/ML
AutoGen
CrewAI
Fine-tuning
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
RAG
Semantic Kernel
Semantic Search
Synthetic Data
Vertex AI
Agentforce
AWS Bedrock AgentCore
Semantic Search
AI Agents
Model Context Protocol
DevOps
AWS
CI/CD
GitHub Actions
Jenkins
Vector
GitHub
Cybersecurity
Crowdstrike
Management
Slack
Marketing
Salesforce
Apply
$63k – $137k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Madrid • Barcelona
AI/ML
AI Agents
Apply
$39k – $84k per year (Estimated) • In office • Bachelor's Degree • Noida
JavaScript
C#
TypeScript
C#
.NET
AI/ML
AI Agents
Frontend
Angular
React.js
DevOps
AWS
GCP
Apply
$24k – $30k per year (net) • Remote • Moscow
JavaScript
Node JS
PHP
TypeScript
Node JS
Mongoose
PHP
Laravel
AI/ML
AI Agents
Model Context Protocol
Frontend
GraphQL
React.js
Sass
DevOps
AWS
CI/CD
Git
Design
Figma
Apply
$50k per year • Remote • Part-Time • 2+ years exp • United States
AI/ML
RAG
Apply
$250k – $300k per year • In office • Full-Time • 7+ years exp • Master's Degree • Palo Alto
Python
AI/ML
Fine-tuning
Gymnasium
LLM
RAG
RLHF
Synthetic Data
TRL
Transformers
DPO
GRPO
Post-training
PPO
Function Calling
Robotics
Digital Twin
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Malaysia
C#
C++
PowerShell
SQL
Databases
MS SQL
AI/ML
RAG
DevOps
Vector
CI/CD
Configuration Management
Windows Server
Apply
$64k per year • In office • Full-Time • San Antonio
AI/ML
RAG
Apply
Software QA Tester 7 days ago
$9k – $37k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Hyderabad
AI/ML
RAG
Apply
$78k – $130k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Cary • San Jose
Analytics
Power BI
Apply
$91k – $182k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • United States
Design
SolidWorks
Apply
$45k – $60k per year • Remote • Full-Time • 2+ years exp • PhD • United States
Cybersecurity
HIPAA
Apply
$117k – $258k per year (Estimated) • Equity • Remote • Full-Time • United States
C++
Java
Python
Cybersecurity
Crowdstrike
Apply
$90k – $184k per year (Estimated) • Equity • Remote • Full-Time • 3+ years exp • United States
AI/ML
Red Teaming
Cybersecurity
Burp Suite
Cobalt Strike
Crowdstrike
Metasploit
MITRE ATT&CK
Nessus
Nmap
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.