1,433,388open jobs
83,596companies
217,057added this week
Browse all
Salary
≈ $128k – $267k per year (Estimated)
Location
In office (San Francisco)
Seniority
Staff
Visa
H-1B filings in 12 months: 2
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Oct 2, 2026. Vals AI scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Vals AI runs private, domain-specific benchmarks that measure how language models perform on legal, tax and finance work. Its public leaderboards compare frontier models on tasks professionals actually do. Enterprises use the private evaluations before deploying a model in regulated workflows.

About the Role

We’re hiring a Technical Content Lead to shape how Vals explains some of the most important questions in AI: which models actually work, how enterprises should evaluate them, and how the industry should measure progress as capabilities improve.

You’ll write across three parts of Vals:

  • Vals Smith, our product for evaluating models and agents on the actual work companies do.

  • Our public benchmarks and research, where we independently measure frontier models across increasingly complex domains.

  • Emerging policy and standards conversation around AI evaluation, where rigorous measurement is becoming increasingly important to how labs, enterprises, and policymakers understand model capabilities.

You’ll work closely with our research, engineering, product, and customer teams to turn technical work into sharp, accessible writing: benchmark reports, product launches, customer stories, model analyses, technical explainers, and occasional policy pieces.

We want someone who can understand the underlying work and develop a distinctive point of view on what matters. The goal is to make Vals one of the most trusted voices on how AI is actually performing - in benchmarks, inside companies, and as the technology becomes important enough to require better standards around how we measure it.

What You’ll Do

  • Interview customers, researchers, engineers, and technical leaders and turn those conversations into strong long-form pieces.

  • Help establish Vals Smith as the leading way for enterprises to evaluate models and agents on their own workflows through product launches, customer stories, technical case studies, and original analysis.

  • Ghostwrite and edit for Vals leadership and technical team members when their perspective should be part of the conversation.

  • Help decide what Vals should be writing about in the first place - not just execute against a content calendar.

What We're Looking For

  • A portfolio of long-form technical writing - published essays, research blog posts, lab posts, or industry reporting.

  • Strong technical literacy. You don't need to be able to train a model, but you need to be able to read a paper, sit in on a research review, and come out with something accurate and sharp.

  • An editorial point of view. You should have opinions about what's interesting, push back on weak takes, and not just transcribe what people say.

  • Speed. We ship content the day a major model releases. You should be able to turn a publishable post in hours, not weeks.

  • Ability to work in-person, in San Francisco. We will support your relocation as needed.

Nice to Haves:

  • Prior research communications experience at an AI lab or research-led startup

  • A technical background (CS, ML, math, sciences) that you've since translated into writing - e.g., engineer or PM turned writer.

  • Prior writing for a top-tier technical publication or engineering blog

What we offer:

  • Highly competitive salary and meaningful ownership. Excellence is well rewarded.

  • Relocation and transportation support

  • Full health, dental, and vision insurance coverage

  • Lunch and dinner provided, free snacks/coffee/drinks

  • 401K plan

  • Unlimited PTO

  • $1,500 housing stipend (within one-mile radius)

About us:

Founding team: The core methodology behind this platform comes from NLP evaluation research we had done at Stanford. We work with all the major foundation model labs, some of the largest financial institutions, and hospital systems in the world. Our team has prior work experience at NVIDIA, Meta, Microsoft, Palantir and HRT. Collectively, we have over 300 citations in our published work. Our early team includes Stanford PhDs, ex-Jane Street quants, and the first designer at Snorkel.

We recently announced our $40M Series A at a $400M valuation, led by Andreessen Horowitz, with participation from existing investors 8VC, Pear VC, and Bloomberg and new investors Hudson River Trading and NextLadder Ventures.

What We’re Looking For

  • Learning velocity: The role encompasses a wide variety of tasks. Rather than expecting you to be an expert on Day 1, we are looking for someone who can learn new skills and technologies quickly.

  • Ownership: Working in a small, talent-dense team, we expect everyone to show initiative to build where it's needed, not where it's asked. We strive for autonomy over consensus.

  • Intensity: The LLM landscape is constantly changing. Foundation model labs are continuously pushing the frontier. The unicorn companies that will emerge from this technology shift are being built now. Those that win will have an incredibly high speed of execution.

  • Solution-oriented mindset: We're looking for people who see opportunities to craft solutions at each juncture, not those who pass hard problems to others or admit defeat.

Vals in the Media:

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,433,388 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Same role
Similar stack
Same company
San Francisco
≈ $61k – $132k per year (Estimated) • In office • Full-Time • Singapore
Python
C
C#
C++
C
FFmpeg
C++
TensorFlow C++
PyTorch C++
Databases
Apache Kafka
AI/ML
Computer Vision
NLP
VLM
TensorFlow
PyTorch
LLM
DevOps
Docker
Kubernetes
Robotics
GStreamer
Apply
≈ $107k – $190k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp • Washington • New York • Columbia
Databases
Apache Iceberg
Apache Kafka
Trino
StarRocks
AI/ML
Copilot
Spark
Airflow
Claude Code
Model Context Protocol
AI Agents
Flink
LLM
Human-in-the-Loop
LLM Guardrails
Tool Use
DevOps
CI/CD
Apply
≈ $31k – $78k per year (Estimated) • Remote (South Africa, Egypt, Turkey, Poland, Ukraine) • Full-Time
AI/ML
Prompt Engineering
Function Calling
AI Agents
RAG
OpenAI
Anthropic
Structured Outputs
Multi-Agent Systems
Tool Use
Mobile
Clean Architecture
DevOps
CI/CD
Apply
Data Scientist 12 hours ago
$78k – $164k per year • In office • TS/SCI • Full-Time • 3+ years exp • Bachelor's Degree • Tampa
Python
AI/ML
NLP
Sentiment Analysis
Machine Learning
Analytics
Tableau
Power BI
Plotly
SpaceTech
QGIS
Apply
System Engineer 12 hours ago
In office • Full-Time • Belgrade
Java
SQL
Bash
Java
Spring Boot
Databases
MS SQL
ElasticSearch
AI/ML
LLM
DevOps
Docker Compose
GitHub Actions
Kibana
Logstash
GitLab CI
CI/CD
Jenkins
Git
Docker
Kubernetes
Linux
TCP/IP
DNS
Apply
$150k – $200k per year • In office • Full-Time • 3+ years exp • San Francisco
AI/ML
DeepSeek
NLP
LLM
OpenAI
Snorkel
Apply
$140k – $185k per year • In office • Full-Time • Master's Degree • San Francisco
Python
JavaScript
Python
Django
AI/ML
DeepSeek
NLP
LLM
OpenAI
Snorkel
Machine Learning
Frontend
React.js
DevOps
AWS
Apply
$140k – $185k per year • In office • Full-Time • San Francisco
Python
JavaScript
TypeScript
Python
FastAPI
Django
AI/ML
DeepSeek
NLP
LLM
Tokenization
OpenAI
Snorkel
Frontend
React.js
DevOps
Git
AWS
Apply
Head of Research 4 months ago
$225k – $275k per year • In office • Full-Time • PhD • San Francisco
Python
JavaScript
Python
Django
AI/ML
DeepSeek
NLP
LLM
OpenAI
Anthropic
Snorkel
Human-in-the-Loop
Frontend
React.js
DevOps
AWS
Apply
Evaluations Engineer 4 months ago
$140k – $185k per year • In office • Full-Time • San Francisco
Python
JavaScript
Python
Django
AI/ML
DeepSeek
AI Agents
NLP
LLM
OpenAI
Snorkel
Machine Learning
Frontend
React.js
DevOps
Git
AWS
Apply
$160k – $180k per year • Equity • In office • 8+ years exp • PhD • San Francisco
AI/ML
Computer Vision
Edge AI
Marketing
Salesforce
Apply
Perception Intern 9 hours ago
≈ $71k – $116k per year (Estimated) • In office • Internship • San Francisco
Python
AI/ML
Computer Vision
PyTorch
Machine Learning
Apply
$96k – $145k per year • In office • Full-Time • 10+ years exp • PhD • San Francisco • Chicago • New York • Atlanta
AI/ML
AI Agents
Agentforce
Analytics
Tableau
Management
Slack
Google Workspace
Gmail
Apply
$149k – $224k per year • In office • Full-Time • 5+ years exp • PhD • San Francisco • Chicago • New York • Atlanta • Washington
AI/ML
AI Agents
Agentforce
Design
Figma
UserTesting
Management
Slack
Google Workspace
Apply
$180k – $220k per year • In office • Contractor • San Francisco
Apply
See all jobs
This is one of many
1,433,388 more open roles from verified company boards, updated every day.