654,521open jobs
38,054companies
91,731added this week
Browse all
Salary
$147k – $305k per year (Estimated)
Location
In office (Brooklyn)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Deterministic RL environments and trajectory datasets for CUA training. Pixel-perfect multi-layer clones, temporal integrity, automated verifiable task generation.

About Us

Chakra Labs' mission is to encode human taste into intelligence. We build high-fidelity environments, evals, and datasets for frontier AI research, working with several of the top labs.

Our work sits at the frontier of post-training, agent environments, data quality, and research infrastructure. We care about building systems that make models better in ways that are measurable, useful, and hard to fake.

What You’d Work On

  • Post-training loops. You’d help design and run model improvement workflows across supervised fine-tuning, preference optimization, and reinforcement learning approaches like GRPO. The work is not just launching training jobs; it’s figuring out what signal matters, how to collect it, and whether the model actually improved.

  • Environment and task design. We build environments that feel real and scenarios that push agents past static benchmark behavior. You’d design tasks, tools, validators, reward signals, and evaluation harnesses that test meaningful capabilities instead of whatever is easiest to measure.

  • High-fidelity trajectories. You’d create, inspect, and improve the data that teaches models how to behave. That means caring about taste, correctness, edge cases, and whether a trajectory would actually help a frontier model learn.

  • Reward and evaluation systems. You’d work on reward functions, rubrics, validators, and analysis tools that turn messy model behavior into useful training signal. You should be interested in where evals lie, where rewards get hacked, and how to make measurements more robust.

  • Training and research infrastructure. You’d run experiments across distributed GPU clusters, work with PyTorch and FSDP, and build the infrastructure needed to support model training, evaluation, and data generation at scale.

  • Customer research problems. You’d work with frontier AI labs to translate ambiguous research goals into concrete environments, datasets, experiments, and deliverables.

About You

  • Machine learning fundamentals. You have Masters / PhD-level knowledge of machine learning fundamentals. You understand linear algebra, optimization, stochastic gradient descent, and can reason from first principles when a model or training run behaves unexpectedly.

  • Hands-on post-training experience. You have worked with or deeply understand LLM fine-tuning, preference data, reward modeling, or reinforcement learning for language models. You know the difference between reproducing a recipe and understanding why it works.

  • Environment-building instinct. You are excited by agent environments, multi-turn tool use, sandboxed tasks, eval harnesses, and the question of how to test capabilities that do not fit neatly into a benchmark.

  • Python and PyTorch fluency. You are strong in Python and comfortable building with PyTorch, FastAPI, and modern ML infrastructure. You can work at a high level, but you are also comfortable dropping into lower-level primitives when the abstraction leaks.

  • Strong engineering judgment. You can move between research ambiguity and production constraints. You know when to iterate quickly, when to be rigorous, and when a result is too fragile to trust.

  • Experience. No hard rule. Roughly 3-5 years is what we imagine, but more or less experience works if the expectations above resonate with you.

What Makes This Different

  • The work is concrete. You are not just “touching the latest AI stack.” You are building the environments, trajectories, evals, rewards, and training loops that frontier labs use to improve models.

  • It is research-facing, but production-minded. Our customers are AI researchers and labs pushing the edge of what agents can do. The systems you build need to support real experiments, real users, and real deadlines.

  • Ownership, not theater. You own whole problems, not isolated tickets. One week you might be designing a new environment; the next you might be debugging a training run, improving a reward function, or scaling an eval pipeline.

  • Ambiguity is part of the job. There is no fixed playbook for this work. Data, post-training, and agent evaluation are changing quickly. If you need every problem to be fully specified before you begin, this role will be challenging.

  • The team. Our team is ex-Stripe, Snap, AWS, Microsoft, Airtable - you'll work with a small team who has years of shipping high-impact products over the last decade.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
654,521 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Brooklyn
$33k – $65k per year (Estimated) • Remote • Full-Time • 4+ years exp • Moscow
Python
SQL
Python
FastAPI
Databases
ClickHouse
Milvus
FAISS
Qdrant
Google BigQuery
BigQuery
AI/ML
LoRA
Model Context Protocol
Fine-tuning
Prompt Engineering
Computer Vision
AI Agents
NLP
PEFT
QLoRA
Llama
Mistral
PyTorch
LLM
RAG
Feature Store
DevOps
gRPC
CI/CD
Docker
Kubernetes
GitHub
Management
n8n
Apply
up to $22k per year (net) • In office • Full-Time • Saint Petersburg
Python
SQL
Python
SQLAlchemy
FastAPI
Ruff
AI/ML
Claude
ChatGPT
Model Context Protocol
AI Agents
LLM
Apply
Data Engineer 1 day ago
$21k – $57k per year (Estimated) • In office • 1+ year exp • Bengaluru
Python
SQL
Databases
Snowflake
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Spark
dbt
DevOps
CI/CD
AWS
Analytics
ETL/ELT
AWS Glue
Apply
$17k – $40k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Mumbai • Bengaluru
Python
SQL
Analytics
Tableau
Matplotlib
Looker
Apply
$22k – $54k per year (Estimated) • Remote • Full-Time • Moscow
Python
C++
Bash
Python
bandit
C++
CMake
DevOps
GitLab CI
CI/CD
Jenkins
Cybersecurity
SonarQube
Semgrep
SBOM
Apply
$146k – $301k per year (Estimated) • In office • Full-Time • 3+ years exp • Brooklyn
Python
TypeScript
AI/ML
LLM
Post-training
DevOps
AWS
Management
Airtable
Stripe
Apply
$148k – $306k per year (Estimated) • In office • Full-Time • 3+ years exp • Brooklyn
Databases
Apache Kafka
AI/ML
AI Agents
LLM
Post-training
Prompt Caching
DevOps
AWS
Kubernetes
Management
Airtable
Stripe
Apply
$35k – $84k per year (Estimated) • In office • Full-Time • High School Diploma • Brooklyn
Apply
$90k – $100k per year • In office • Full-Time • Brooklyn
Apply
$63k – $169k per year (Estimated) • In office • Part-Time • Brooklyn
Apply
$75k – $151k per year (Estimated) • In office • Full-Time • 8+ years exp • High School Diploma • Brooklyn
DevOps
Azure DevOps
Azure
CI/CD
Apply
$54k – $126k per year (Estimated) • In office • Full-Time • Brooklyn
Apply
See all jobs
This is one of many
654,521 more open roles from verified company boards, updated every day.