368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$33k – $65k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
Emergent is an AI-driven tech platform that streamlines software development through its autonomous coding platform where AI agents manage coding, testing, and deployment, allowing users to focus on creativity.

We're looking for a Research Engineer to characterize, measure, and advance the capabilities of our coding agents. You will turn ambiguous notions of "agent quality" into clear, defensible metrics that the team, leadership, and the field can rely on, and you will use those metrics to drive both incremental wins and moonshots in agent performance.

This is a deep-work role at the intersection of agent behavior, evaluation research, and applied training. You will define what good looks like for long-horizon coding agents, build the evaluation dataset and methodology that produces those signals, mine production data for failure modes most teams never see, and run targeted training, fine-tuning, RL, memory, and prompt-optimization experiments that translate research advances into shipped improvements. You will operate with strong independence, make hard calls in inherently subjective and probabilistic systems, and own outcomes end-to-end. If you treat models as objects of study rather than black boxes, take pride in moving benchmark numbers with rigor, and want to apply the frontier of agent research at the scale of millions of real applications, this is your role.

Responsibilities:

  • Architect the next version of the Emergent agent. Shape the core architecture and make the foundational design choices that define how the agent thinks, learns, and improves over time.
  • Characterize agent behavior at depth. Develop a deep, evidence-grounded understanding of how the agent succeeds and fails across the full range of real-world usage, and convert that understanding into rigorous, quantitative measurement.
  • Design and ship evaluations across reasoning, planning, tool use, code correctness, long-horizon execution, security, and agent reliability. Define the metric, build the dataset, validate against known signals, and ship dashboards that make regressions impossible to miss.
  • Drive step-function gains. Take on the ambitious bets that meaningfully advance the state of the art, the 10-point leaps on hard capabilities, not incremental polish. Pick the problems where the upside is large and the path is uncertain.
  • Climb public benchmarks. Move the needle on SWE-bench Pro, Terminal-Bench, and other industry-standard benchmarks the field uses to grade coding agents.
  • Run training and post-training experiments, supervised fine-tuning, RLHF/RLAIF, DPO, distillation, reward modeling, prompt optimization, and judge-model calibration against production-grounded objectives.
  • Own end-to-end. Carry work from hypothesis through experiment design, execution, analysis, decision, rollout, and post-launch measurement. Read research papers deeply, get inspired ideas, and turn them into shipped products.
  • Make hard calls in subjective systems. Decide when a regression is real, when a win is noise, when a benchmark is overfit, when to ship despite mixed signals, and when to kill a promising direction. Communicate the reasoning crisply.

Requirements:

  • 5-8 years of AI experience, with meaningful time spent either training and fine-tuning models or designing rigorous evaluations and measurement systems for them. Both paths are equally valued for this role.
  • Hands-on with the modern AI stack and fluent in Python (Go is a plus) for research workflows: training pipelines, eval harnesses, data processing, and statistical analysis. Comfortable with transformers, RLHF/DPO/RL for agents, eval frameworks (Inspect, lm-eval-harness, or equivalent), prompt optimization, judge models, and agent frameworks. You pick up new tooling in days.
  • Take pride in numbers that move. You measure first, opine second. You can defend why a benchmark is the right benchmark, why a metric isn't gameable, and why a result is statistically real.
  • Comfortable in subjective, probabilistic systems. You reason about noise floors, confounds, distribution shifts, judge bias, and selection effects without flinching. You know when to trust a number and when to suspect it.
  • Enjoy going deep into the long tail. Sifting through large volumes of agent behavior to find the rare, hidden failure mode energizes you, not drains you.
  • Understand models like friends. You have intuitions about how a model will behave on a new task before running it, and you update those intuitions when reality disagrees. You know what came out last week, why it matters, and which paper from two years ago is suddenly relevant again.
  • Independent operator with leadership presence. You scope your own work, push back on weak ideas (including your manager's), and bring others along through clarity and conviction rather than consensus-seeking.
  • Ship fast without compromising rigor. You know which corners are safe to cut and which are load-bearing. Bias toward velocity, but never at the cost of honest measurement.
  • Bonus signal: prior publications, strong showings on coding/reasoning benchmarks, contributions to open-source agent or eval frameworks, experience with long-horizon agents, RL training infrastructure, or production data flywheels.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$120k – $250k per year • Equity 0.2–1% • In office • Full-Time • Master's Degree • Seattle
Python
AI/ML
Diffusion Models
Fine-tuning
PyTorch
Self-Supervised Learning
Apply
$230k – $260k per year • Equity • Remote • Internship • Bachelor's Degree
Python
DevOps
Amazon EKS
AWS
Azure
CI/CD
GCP
Helm
Kubernetes
Terraform
Cybersecurity
FedRAMP
Orca Security
Apply
$20k – $56k per year (Estimated) • Equity 0–0.1% • In office • Full-Time • 3+ years exp • Master's Degree • Bengaluru
MATLAB
Python
Apply
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
$23k – $53k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Moscow
Python
DevOps
Git
QA
Pytest
Apply
$33k – $83k per year (Estimated) • In office • 5+ years exp • Bengaluru
Python
AI/ML
Fine-tuning
Knowledge Distillation
lm-eval-harness
RLHF
DPO
LLM Evaluation
Post-training
SFT
Function Calling
Apply
$33k – $87k per year (Estimated) • In office • 6+ years exp • Bengaluru
JavaScript
TypeScript
Frontend
GraphQL
Material UI
Next.js
React Query
React.js
Recoil
Redux
Storybook
Turborepo
Vite
Webpack
Zustand
Mobile
State Management
DevOps
CI/CD
Design
Figma
QA
Cypress
Jest
Apply
$31k – $70k per year (Estimated) • In office • 6+ years exp • Bengaluru
DevOps
AWS
CI/CD
GCP
Kubernetes
IAM
Cybersecurity
Zero Trust
Apply
Frontend Engineer 26 days ago
$20k – $66k per year (Estimated) • In office • 6+ years exp • Bengaluru
JavaScript
TypeScript
Frontend
GraphQL
Material UI
Next.js
React Query
React.js
Recoil
Redux
Storybook
Turborepo
Vite
Webpack
Zustand
Mobile
State Management
DevOps
CI/CD
Design
Figma
QA
Cypress
Jest
Apply
$47k – $102k per year (Estimated) • In office • Bengaluru
C++
Go
Java
Python
Rust
Databases
Databricks
Snowflake
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.