368,657open jobs
9,442companies
50,883added this week
Browse all
Location
Remote (AMER)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Nous Research is a leader in the American open source AI movement, training world-class open source language models.

The Role

You'll work across the lab on agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together. This is a high-growth, high-ownership role on a small team, and you'll ship evaluation infrastructure that researchers depend on from day one.

Responsibilities

  • Run the full eval pipeline end to end and reproduce known results during onboarding, pairing with a senior engineer on your first task

  • Build a judge calibration protocol: sample human-labeled decisions, measure agreement (κ, per-class P/R), identify drift zones, and document it so anyone can re-run it

  • Extend an existing benchmark (GAIA, τ-Bench, SWE-bench slice, etc.) with new tasks targeting known capability gaps, including the prompt, environment, rubric, automated grader, and QA

  • Run failure analysis on model outputs: categorize failure modes, quantify prevalence, and write up findings with recommendations for training data, judge prompts, or benchmark changes

  • Own a recurring eval workflow (weekly regression suite, judge drift dashboard, red-team evaluation for a new capability) and ship tooling researchers actually use

Qualifications

  • 3+ years in software engineering, ML engineering, data science, or a research-adjacent role, with concrete evaluation experience from coursework, an internship, a side project, open source work, or a job

  • Experience with at least one LLM evaluation framework (Harbor, Nemo Evaluator, etc.), with real opinions on what it does well and where it falls short

  • Hands-on experience with LLMs: prompting, few-shot design, and ideally fine-tuning or RAG; regular use of coding agents

  • Solid Python. You write clean, tested, version-controlled code that a colleague could run without you babysitting it

  • Comfort with Git, CI/CD basics, Docker, and the Linux command line (SSH, tmux, debugging a remote job)

  • Understanding of basic eval statistics: why accuracy misleads on imbalanced judges, what Cohen's κ measures, how to think about confidence intervals on a metric

  • At least 3 of the following: you can explain why LLM-as-judge needs calibration; you've done failure analysis and can tell model bugs apart from prompt, grader, or retrieval issues; you know at least two agent benchmarks (GAIA, AgentBench, τ-Bench, MINT, SWE-bench, WebShop, ALFWorld) and a limitation of each; you've designed or extended an eval dataset with happy paths, edge cases, and adversarial examples; you've thought about non-determinism in eval, how you sample, how many runs, how you report variance

  • You communicate clearly to both researchers and engineers, in the right language for each

  • You're comfortable with ambiguity, can turn a half-formed request into a plan, and know when to ask for help

Preferred

  • RLVR / RLHF pipeline experience

  • Training data curation experience

  • Distributed eval orchestration experience

  • Benchmark design from scratch

  • Red teaming and adversarial eval experience

  • Familiarity with psychometrics or measurement theory

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
DevOps Engineer 6 hours ago
$49k – $88k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Warsaw
Bash
Python
DevOps
AWS
Azure
Azure AKS
CI/CD
Docker
GCP
Git
GitHub
GitHub Actions
GitOps
Grafana
Kubernetes
Loki
Prometheus
Terraform
Thanos
Apply
$20k – $46k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
AI/ML
AI Agents
Amazon SageMaker
Anomaly Detection
AWS Bedrock
Computer Vision
Embeddings
LLM
Multimodal AI
Prompt Engineering
RAG
Time Series Forecasting
DevOps
Amazon S3
AWS
AWS Lambda
CI/CD
Git
Vector
Analytics
Power BI
Tableau
Apply
$27k – $71k per year (Estimated) • In office • Full-Time • 7+ years exp • Pune
C#
C++
Java
Python
DevOps
Azure
CI/CD
Git
QA
Pytest
Apply
QA Engineer II 6 hours ago
$13k – $41k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Hyderabad
Python
AI/ML
AI Agents
DevOps
CI/CD
Git
QA
Playwright
Pytest
Selenium
Apply
$164k – $219k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Chantilly
JavaScript
Node JS
SQL
TypeScript
Java
Java
Hibernate
Spring Boot
Databases
MySQL
Oracle
PostgreSQL
Frontend
Less
React.js
DevOps
Amazon EC2
Amazon ECS
Amazon S3
AWS Lambda
CI/CD
CloudFormation
Docker
Jenkins
Kubernetes
AWS
Apply
Remote • Full-Time • 5+ years exp
Kotlin
Node JS
Python
Swift
TypeScript
JavaScript
AI/ML
AI Agents
LLM
Frontend
React.js
Mobile
Offline-First
React Native
Apply
$145k – $262k per year (Estimated) • Remote • Full-Time
Python
Rust
TypeScript
AI/ML
LLM
PyTorch
OpenAI
Apply
Security Engineer 18 days ago
Remote • Full-Time • 8+ years exp
AI/ML
AI Agents
DevOps
AWS
Azure
GCP
Kubernetes
Vercel
IAM
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Apply
UI/UX Designer 18 days ago
Remote • Full-Time • 3+ years exp
JavaScript
AI/ML
AI Agents
LLM
Human-in-the-Loop
Frontend
React.js
Design
Figma
Apply
Remote • Full-Time • 3+ years exp
Go
Node JS
Python
Rust
TypeScript
JavaScript
AI/ML
LLM
Edge AI
OpenAI
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.