974,315open jobs
58,640companies
159,915added this week
Browse all
Salary
$250k – $290k per year
Location
Hybrid (San Francisco, United States)
Seniority
Middle · 4+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 30, 2026. First seen by Alion on Sep 29, 2026.

Overview
Company
Impact
Profile match
Firecrawl builds a web data API that lets developers and AI agents search, scrape, crawl and interact with websites and receive the content as clean markdown or structured JSON. The service is backed by an open-source crawler with tens of thousands of GitHub stars and, since September 2026, by Alexandria, a layer that combines official data providers, custom connectors and its own indexes with live web access behind one interface. The San Francisco company came out of Y Combinator, counts more than 1.5 million users, and raised a $14.5 million Series A in 2025 and a $75 million Series B led by Smash Capital in 2026.

Research Engineer - Benchmarks

Every week, someone asks which data provider is actually best: for company records, for finding the right person, for fresh job listings. Right now the answers come from the vendors themselves. You'll build the independent version. You'll run rigorous, automated, public benchmarks of the providers in Alexandria and beyond, and publish them weekly. The same results will feed straight back into Alexandria so it learns which provider to call for which job.

This isn't our internal evals role. You're measuring the market, in public, where every number will be challenged by the vendors it ranks. You'll own it end to end: the datasets, the ground truth, the scoring, the harness, the weekly release, and the loop into Alexandria's provider selection. No one hands you a methodology. You write it, defend it, and ship it every week.

Salary Range: $250,000-$290,000 USD/year (SF) / $210,000-$224,000 CAD/year (Toronto)

Equity Range: Competitive equity. Details shared during the process.

Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). Hybrid, onsite 3+ days a week.

Equity Range: Competitive equity. Details shared during the process.

Location: San Francisco, CA (SF HQ). On-site, five days a week.

Job Type: Full-Time

Experience: 4+ years in ML, research engineering, or data engineering, with evaluation or benchmark work you've shipped

Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub.

About Firecrawl

Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved.

In September 2026, we raised a $75M Series B led by Smash Capital, and we're spending it building the largest repository of knowledge in the world. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 187k+ GitHub stars, putting us in the top 40 repositories of all time, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started.

We're a small team punching far above our weight. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount.

This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. That library is called Alexandria, and it starts now.

What You'll Do

  • Design and run head-to-head benchmarks of data providers against verified ground truth. For example:

    • Apollo vs. FullEnrich vs. DataLegion on company details.

    • FullEnrich vs. DataLegion on finding the right person and their current employer and title.

    • Built In vs. ZipRecruiter on relevant, fresh, non-duplicate job listings.

  • Build and maintain the test datasets and ground truth, and keep them from going stale or leaking.

  • Own the automated harness and the weekly benchmark release. Every run has to be reproducible, versioned, and defensible.

  • Measure what buyers actually care about: accuracy, coverage, freshness, speed, and cost.

  • Work with marketing to ship public leaderboard pages that are useful and hold up under scrutiny.

  • Close the loop into Alexandria so benchmark results change which provider gets called for what.

  • Keep expanding the categories we benchmark, inside Alexandria and beyond it.

What We're Looking For

  • You've shipped evals or benchmarks, and you can explain exactly why your results were trustworthy.

  • You've done this somewhere that matters: a frontier lab, a data company like Scale, Surge, micro1 or Mercor, a third-party benchmark org, or a public benchmark project. OSS contributors very welcome.

  • You're strong in Python and API integrations, and comfortable with messy vendor APIs, rate limits, and inconsistent schemas.

  • You know how to build test sets and scoring methods: sampling, labeling, inter-rater agreement, and when to trust an LLM judge and when not to.

  • You use AI heavily and keep upgrading your own workflow. You ship without waiting for instructions.

  • You write findings clearly for engineers, marketers, and the vendors on the other side of the leaderboard.

What We're Not Looking For

  • Someone who wants to run benchmarks someone else designed.

  • Someone who picks the metric that makes the story look good. Our numbers have to survive the vendors who lose.

  • A pure researcher who won't build the harness, or a pure engineer who won't think hard about methodology.

  • Someone who needs a fully specced ticket, or a quarter, to ship the first result.

A Note On Pace

We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings, but this role probably isn't for you.

Benefits & Perks

Available to all employees

  • Salary that makes sense: $250,000-$290,000 USD/year, / $210,000-$224,000 CAD/year (Toronto), based on impact, not tenure

  • Own a piece: Gain competitive equity in what you're helping build

  • Generous PTO: 15 days mandatory, anything after 24 days, just ask (holidays excluded). Take the time you need to recharge

  • Parental leave: 12 weeks fully paid, for all parents

  • Wellness stipend: $100 USD/month for the gym, therapy, massages, or whatever keeps you human

  • Learning & Development: Expense up to $1,000 USD/year toward anything that helps you grow professionally

  • Team offsites: A change of scenery, minus the trust falls

  • Sabbatical: 3 paid months off after 4 years, do something fun and new

Available to US-based full-time employees

  • Full coverage, no red tape: Medical, dental, and vision (100% for employees, 50% for partner and kids). No weird loopholes, just care that works

  • Life & Disability insurance: Employer-paid basic life and AD&D, short-term disability, and long-term disability. Coverage for life's curveballs

  • Virtual care and a health guide: Teladoc for the couch doctor visit, plus Rightway to answer coverage questions and fight billing errors for you

  • Mental health: Talkspace, therapy and psychiatry on your schedule

  • Fertility and family building: Carrot, covering you and your partner

  • EAP: Free confidential counseling, legal and financial consults, and online will prep through Guardian

  • 401(k) plan: Retirement might be a ways off, but future-you will thank you

  • Pre-tax benefits: HSA, FSA, and commuter benefits to help your wallet out a bit

  • Supplemental options: Extra life and AD&D, accident, critical illness, hospital indemnity, plus pet, legal, and identity protection through MetLife

Available to Canada-based full-time employees

  • Full coverage, no red tape: Extended health, dental, and vision through Manulife (Diamond, the top tier), 100% employer-paid for you, your partner, and your kids

  • Life & Disability insurance: Employer-paid life, AD&D, short-term disability, and long-term disability. Coverage for life's curveballs

  • Virtual care: Dialogue Premium, so you can see a doctor or nurse from your couch, any hour

  • Mental health: Talkspace Elite, therapy and psychiatry on your schedule

  • Fertility and family building: Carrot, covering you and your partner

  • Retirement: Group RRSP through Wealthsimple, so future-you can thank you

Available to SF-based employees

  • SF HQ perks: Snacks, drinks, team lunches, intense ping pong, and peak startup energy

  • E-Bike transportation: A loaner electric bike to get you around the city, on us

Available to Toronto-based employees

  • Toronto Hub perks: Snacks, drinks, team lunches, glass-walled views down University Avenue, and a home base steps from Union Station

  • Transit, covered: A PRESTO card loaded for GO Transit, subway, and streetcar, plus station parking if you drive to the train. Winter-proof, on us

Interview Process

  • Application Review: Send us your work. We want the benchmark, leaderboard, eval, or dataset you built, and how you knew it was right. We care about what you've shipped, not where you went to school.

  • Intro Chat (~25 min): A quick conversation to get to know each other. We'll cover what you've been working on, what drew you to Firecrawl, and what you want next. Time for your questions too.

  • Technical Chat (~45 min): A real problem from our world: design a benchmark that decides whether FullEnrich or DataLegion finds the right person's current title. We'll cover where the ground truth comes from and how you'd defend the result to the vendor who loses. Come ready to think out loud.

  • Workflow Chat (~30 min): Show us how you actually work: your AI tools, your setup, and a recent thing you shipped faster than you would have a year ago.

  • Founder Chat (~25 min): Culture, pace, ownership, and how you like to work. Time for your questions too.

  • Paid Work Trial (~40 hours): Ship a small benchmark end to end on a real provider category, paid at a contractor rate. It's the truest signal for both sides. Remote-friendly, and we'll flex around your current commitments.

  • Decision: We move fast after the trial.

If you want to be the person the whole market checks before picking a data provider, you should join us.

Apply now.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
974,315 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$119k – $164k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Evanston
AI/ML
Computer Vision
NLP
Apply
≈ $130k – $247k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Master's Degree • Malvern • Charlotte • Dallas • Philadelphia • Scottsdale
Python
AI/ML
Function Calling
AI Agents
NLP
LLM
RAG
LLM Guardrails
Tool Use
Machine Learning
Apply
$151k – $280k per year • Hybrid • Full-Time • 4+ years exp • Master's Degree • Bellevue • San Francisco • New York
Python
AI/ML
Fine-tuning
Embeddings
Multimodal AI
Computer Vision
TensorFlow
PyTorch
Machine Learning
Apply
AI Engineer II 3 hours ago
$135k – $200k per year • Equity • Hybrid • Full-Time • 2+ years exp • San Francisco
Python
Go
JavaScript
Kotlin
TypeScript
Databases
PostgreSQL
Apache Kafka
AI/ML
Machine Learning
DevOps
CI/CD
AWS
Docker
Kubernetes
AWS Step Functions
Apply
$165k – $247k per year • Equity • Hybrid • Full-Time • 3+ years exp • PhD • San Francisco
Python
Go
JavaScript
Kotlin
TypeScript
Databases
PostgreSQL
Databricks
Apache Kafka
AI/ML
MLFlow
Recommender Systems
DevOps
AWS
Docker
Kubernetes
AWS Step Functions
Cybersecurity
HIPAA
Apply
≈ $20k – $47k per year (Estimated) • In office • 5+ years exp • Hyderabad
Python
Java
AI/ML
Embeddings
AI Agents
LLM
RAG
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Management
ServiceNow
Agile
Apply
≈ $21k – $42k per year (Estimated) • In office • 5+ years exp • Delhi • Hyderabad • Bengaluru
Python
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
LangChain
Spark
LlamaIndex
MLFlow
Fine-tuning
Embeddings
Scikit-learn
Prompt Engineering
AI Agents
NLP
Semantic Kernel
Transformers
TensorFlow
PyTorch
LLM
RAG
OpenAI
Hugging Face
LLM Guardrails
Agentic Workflows
Machine Learning
DevOps
Azure DevOps
GitHub Actions
Azure
CI/CD
Docker
Kubernetes
Azure AKS
Apply
≈ $7k – $18k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • India
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
AWS
Platform Engineering
Incident Management
GitHub
Apply
≈ $12k – $24k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • India
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
AWS
Platform Engineering
Incident Management
GitHub
Apply
$150k – $200k per year • Equity • In office • San Francisco
AI/ML
AI Agents
Apply
$250k – $290k per year • Hybrid • Full-Time • 3+ years exp • San Francisco
Python
AI/ML
Spark
MLFlow
AI Agents
LLM
Firecrawl
Feature Store
Machine Learning
DevOps
Kubernetes
GitHub
Analytics
A/B Testing
Apply
$250k – $290k per year • Hybrid • Full-Time • 4+ years exp • San Francisco
AI/ML
AI Agents
LLM
Firecrawl
DevOps
GitHub
Apply
Support Engineer 2 days ago
$166k – $218k per year • Hybrid • Full-Time • 3+ years exp • San Francisco
Python
Go
TypeScript
AI/ML
Cursor
Claude
Model Context Protocol
AI Agents
LLM
Firecrawl
DevOps
GitHub
Management
Slack
Apply
$140k – $160k per year • Hybrid • Full-Time • United States
AI/ML
AI Agents
LLM
Firecrawl
DevOps
GitHub
Apply
$235k – $260k per year • Hybrid • Full-Time • 5+ years exp • San Francisco
AI/ML
AI Agents
LLM
Firecrawl
DevOps
GitHub
Apply
≈ $205k – $370k per year (Estimated) • In office • 12+ years exp • Bachelor's Degree • San Francisco
AI/ML
RLHF
Multimodal AI
AI Agents
Gemini
Post-training
EU AI Act
Cybersecurity
GDPR
Apply
$50k – $66k per year • In office • Bachelor's Degree • San Francisco
Management
Microsoft Office
Apply
$126k – $210k per year • Remote (United States) • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco • Miami • Austin • Denver • Raleigh
Marketing
Salesforce
Apply
Front Desk Agent 8 hours ago
$66k per year • In office • High School Diploma • San Francisco
Apply
$66k – $68k per year • In office • 1+ year exp • High School Diploma • San Francisco
Ruby
Management
Microsoft Office
Apply
See all jobs
This is one of many
974,315 more open roles from verified company boards, updated every day.