368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$124k – $250k per year (Estimated)
Location
Remote (United States)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Porque te puede pasar, dejanos protegerte.. Juntos podemos hacerlo... Vive seguro.

Company Overview:

We are building Protege to solve the biggest unmet need in AI - getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.

Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI - and in tech.

We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.

About the Role

We're hiring a Forward Deployed Machine Learning Engineer in our Benchmarks and Evaluations vertical. You'll be the first MLE dedicated to this vertical and will work directly with the GM and our researchers to scale Protege’s position as a renowned leader in the space.

At Protege, we believe that real world data is one of the largest bottlenecks to AI progress. Our data and data expertise position us to be neutral arbiters for the market, helping model builders understand the current performance of their models, identify what data will improve performance, and show that improvement over time. Benchmarks and evaluations power that cycle. As an early engineer in the Benchmarks and Evaluations vertical, this role is an opportunity to help build the technical foundation for a critical area that greatly benefits current and future customers.

What You'll Do

Work on the eval foundation

  • Partner with the GM and early customers to define what constitutes strong evals in different domains
  • Work with Protege researchers to design and build benchmarks
  • Build the standards on how different modalities should be processed

Own infrastructure

  • Build the backend the vertical runs on which includes data pipelines, execution environments, storage, and orchestration
  • Stand up sandboxed environments for agentic evals, where models need tools, code execution, or multi-step tasks

Go from fast iteration to product

  • Find repeatable eval patterns, infrastructure gaps, and product opportunities from live engagements
  • Partner with DataLab (our research team) on domain-specific data and research questions

What Success Looks Like

In the first 90 days, we expect the following:

  • Build an understanding of the evals landscape, the GM's strategy, and customer demand
  • Build an understanding of what our platform and data partners can support today, and where the gap is for eval building
  • Identify the largest technical bets and ship multiple iterations of the eval infrastructure
  • Own the engineering portion of customer engagements end to end

What You Bring

Must Haves

  • 4+ years of engineering experience
  • Hands-on ML work evaluating models
  • Have previously owned backend and infrastructure
  • High ambiguity tolerance and bias to action
  • Comfort working with urgency to meet the pace and volume of the market demands
  • Strong written communication

Nice to Haves

  • Prior experience building benchmarks, evals, or human data pipelines for LLMs
  • Time at a frontier lab, an eval-focused team, or a research org
  • Founding or early engineer experience at a fast-moving startup
  • Familiarity with agentic systems, RL environments, code-execution sandboxes, TEE/TREs

Protege's Values

Pass the Loved Ones' Test

We act with integrity and do the right thing - especially when it's hard and no one is watching.

Always Find a Way

We are resourceful, resilient builders who solve hard problems and push through obstacles.

Go Fast and Grow Fast

Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.

Practice Kindness and Candor

We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.

Deliver Together

We win as one team. Collaboration, accountability, and shared ownership drive our success.

Own the Outcome. Hone the Craft.

We take pride in our work, sweat the details, and continuously raise the bar for excellence.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$75k – $214k per year (Estimated) • Remote • 5+ years exp • Singapore
Go
Ruby
Ruby
Ruby on Rails
AI/ML
AI Agents
Model Context Protocol
DevOps
CI/CD
Git
Rest API
Apply
$78k – $220k per year (Estimated) • Remote • 5+ years exp • Singapore
Go
Ruby
Databases
PostgreSQL
AI/ML
AI Agents
LLM
Model Context Protocol
Anthropic
Function Calling
OpenAI
DevOps
CI/CD
Kubernetes
Apply
$55k – $157k per year (Estimated) • Remote • Full-Time • Sydney
C++
Go
Lua
Python
C++
CMake
Databases
ActiveMQ
Aerospike
Apache Kafka
Cassandra
DevOps
Docker
gRPC
Kubernetes
Apply
$61k – $143k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Cambridge
Go
Node JS
Python
TypeScript
JavaScript
Databases
ElasticSearch
MySQL
Frontend
GraphQL
Next.js
React.js
DevOps
CI/CD
Kubernetes
Apply
$29k – $62k per year (Estimated) • Remote • Full-Time • 4+ years exp • Moscow
Go
PHP
Databases
Apache Kafka
PostgreSQL
RabbitMQ
Redis
Frontend
GraphQL
DevOps
CI/CD
Git
gRPC
WebSockets
Management
Confluence
Draw.io
Jira
Apply
$109k – $199k per year (Estimated) • Remote • Full-Time • 3+ years exp
Python
SQL
C
C
FFmpeg
Databases
Databricks
AI/ML
Dagster
Embeddings
LLM Evaluation
DevOps
AWS
Platform Engineering
Vercel
Apply
$112k – $204k per year (Estimated) • Remote • Full-Time • 3+ years exp
Python
Databases
Databricks
AI/ML
Dagster
Embeddings
DevOps
AWS
Platform Engineering
Vercel
Apply
$96k – $178k per year (Estimated) • Remote • Full-Time • 2+ years exp
TypeScript
JavaScript
Databases
PostgreSQL
AI/ML
Multimodal AI
Frontend
Next.js
React.js
DevOps
AWS
Cybersecurity
HIPAA
Apply
$143k – $237k per year (Estimated) • Remote • Full-Time • 5+ years exp
Go
TypeScript
JavaScript
Databases
PostgreSQL
AI/ML
Embeddings
Multimodal AI
Frontend
Next.js
React.js
DevOps
AWS
Cybersecurity
HIPAA
Apply
$132k – $257k per year (Estimated) • Remote • Full-Time • 4+ years exp
SQL
AI/ML
dbt
Great Expectations
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.