369,097open jobs
9,477companies
48,181added this week
Browse all
Location
Remote (Germany, United Kingdom, Cyprus, Czech Republic, Netherlands, Poland, Serbia, Spain, Armenia)
Seniority
Middle · 3+ years exp
Overview
Company
Impact
Profile match
JetBrains is a global software vendor specializing in intelligent, productivity-enhancing developer tools and integrated development environments (IDEs). Founded in 2000 and headquartered in Prague, Czech Republic, the company is best known for creating IntelliJ IDEA, PyCharm, WebStorm, Rider, ReSharper, and the Kotlin programming language (which serves as Google's preferred language for Android development).

At JetBrains, code is our passion. Ever since we started, back in 2000, we’ve been striving to make the strongest, most effective developer tools on earth. Today, AI-powered assistance and agents are becoming a core part of how developers work in our IDEs.

We’re building multi-step coding agents that can understand large codebases, plan changes, call tools, and iterate with the user. As a Research Engineer in the Agentic Models team, you’ll be responsible for the models, training loops, and evaluation pipelines that power these agents.

You’ll work at the intersection of SFT and RL-style post-training, and product-driven evaluation, using our distributed GPU and MapReduce clusters to ship models into JetBrains products.

As part of our team, you will:

  • Design, implement, and maintain SFT and RL post-training pipelines for multi-step coding agents.
  • Train and adapt LLMs for agent workflows, including planning, tool use, and multi-step interactions inside JetBrains IDEs.
  • Build and develop evaluation and simulation environments where coding agents can act, be measured, and compared on realistic developer tasks.
  • Design evaluation frameworks and metrics for agent behavior, analyze traces and logs, and close the loop from evaluation back into training, data, and reward design.
  • Analyze training and evaluation results to propose and implement improvements to model architectures, training recipes, and datasets.
  • Work with large-scale infrastructure, including distributed training on GPU clusters and large MapReduce-style data processing for pre-training and fine-tuning datasets.
  • Collaborate closely with research, product, and infrastructure teams to turn high-level product visions into concrete models, experiments, and shipped features. 

We’ll be happy to bring you on board if you have:

  • Extensive hands-on experience training LLMs (pre-training, fine-tuning, or post-training) in a research or production setting.
  • Deep expertise in modern deep learning frameworks such as PyTorch, and specialized LLM training stacks (e.g. Megatron, NeMo, verl, or similar).
  • Strong theoretical and practical understanding of LLM fundamentals: architectures, tokenization, data pipelines, batching, mixed precision, distributed training, and debugging unstable runs.
  • The ability to own projects end to end, starting from a high-level problem or product pain point and overseeing it through the design, experimentation, implementation, and iteration phases.
  • A product-aware mindset - you care about how developers actually use agents and can translate product needs and failure modes into modeling and evaluation work.
  • At least 3 years of Python experience writing clean, maintainable code in modern ML codebases.

Our ideal candidate would have experience with:

  • ML orchestrators and workflow tools such as Kubeflow, Dagster, Airflow, ZenML, and/or job schedulers like Kubernetes or SLURM.
  • Large-scale data and training pipelines, e.g. MapReduce-style clusters, multi-node GPU training, or workloads on the order of 1M+ CPU/GPU hours.
  • Designing and maintaining evaluation pipelines for LLMs or agents, including metrics, dashboards, experiment tracking, and automated regression checks.
  • AI agent development, such as tool-using agents, planners, or multi-step coding workflows, and familiarity with agentic frameworks or patterns.
  • Experiment tracking and observability using tools like Weights & Biases, MLflow, Langfuse, or similar.
  • Inference optimization and serving optimized models in production.

We are an equal opportunity employer

We know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace - one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation.

We process the data provided in your job application in accordance with the Recruitment Privacy Policy.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
369,097 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Amsterdam
$77k – $194k per year (Estimated) • In office • Contractor • 5+ years exp • Singapore
Java
SQL
DevOps
CI/CD
Docker
Grafana
Incident Management
Kubernetes
OpenShift
SRE
Apply
$131k – $281k per year (Estimated) • In office • Contractor • 15+ years exp • Singapore
Java
SQL
Java
Maven
Databases
MS SQL
Oracle
DevOps
AWS
CI/CD
Incident Management
Kubernetes
OpenShift
OpenStack
Apply
$75k – $212k per year (Estimated) • In office • Contractor • Singapore
Java
SQL
Databases
Oracle
DevOps
AWS
Docker
Git
Kubernetes
SLI/SLO/SLA
QA
Selenium
Apply
$23k – $58k per year (Estimated) • In office • Full-Time • 7+ years exp • Gurgaon
Python
Databases
Apache Kafka
ElasticSearch
Redis
DevOps
Error Budget
Grafana
Incident Management
Kubernetes
OpenShift
Platform Engineering
Prometheus
Self-Healing
Cybersecurity
HashiCorp Vault
Cryptography
Vault
Apply
Remote • Full-Time
JavaScript
Python
TypeScript
AI/ML
AI Agents
Function Calling
LLM
Prompt Engineering
Structured Outputs
Management
n8n
Zapier
Apply
$63k – $144k per year (Estimated) • In office • 5+ years exp • Amsterdam
AI/ML
AI Agents
Apply
Remote/Hybrid • Amsterdam
AI/ML
AI Agents
Management
YouTrack
Apply
In office • Limassol
Kotlin
Design
Figma
Apply
Remote • Amsterdam
Apply
Remote • Full-Time • 5+ years exp • Amsterdam
Apply
$92k – $223k per year (Estimated) • In office • Contractor • Amsterdam
Databases
PostgreSQL
Redis
AI/ML
AI Agents
Gemini
Google ADK
LangChain
LlamaIndex
Prompt Engineering
Vertex AI
DevOps
GCP
Kubernetes
Management
Slack
Apply
$105k – $252k per year (Estimated) • In office • 8+ years exp • Amsterdam
SQL
Apply
$91k – $216k per year (Estimated) • In office • Full-Time • 10+ years exp • Amsterdam
Java
JavaScript
Java
Gradle
Frontend
npm
pnpm
DevOps
ArgoCD
AWS
Bazel
CI/CD
Datadog
GitHub Actions
GitOps
Helm
JFrog Artifactory
Kubernetes
SRE
Terraform
GitHub
IAM
Cybersecurity
GDPR
Apply
$101k – $242k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Amsterdam
Python
Databases
Databricks
AI/ML
Spark
Apply
$67k – $169k per year (Estimated) • In office • 4+ years exp • Amsterdam
Python
SQL
Python
pySpark
AI/ML
Airflow
Spark
Marketing
Salesforce
Apply
See all jobs
This is one of many
369,097 more open roles from verified company boards, updated every day.