368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$180k – $450k per year
Location
In office (San Jose)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Hark was an online digital entertainment platform best known for its extensive library of short audio soundbites, video clips, and pop culture quotes. Launched in 2007, the website allowed users to browse, create, and share playable soundboards featuring memorable lines from movies, television shows, and political figures. While it grew into a popular destination for viral sound clips during the late 2000s and early 2010s, the platform has since ceased its original operations.

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Team

We are looking for a Member of Technical Staff, Frontier RL Environments to build the north-star agent environments that drive progress toward truly personal, proactive intelligence. You'll decide which skills and behaviors our agents need next, build the evaluations and RL environments that reveal whether they have them, and help set the research agenda behind our highest-stakes training runs.

This role sits at the frontier of a rapidly emerging discipline, where environment design, evaluation science, and large-scale model training converge to produce agents that reason, plan, remember, and act on a person's behalf with proactivity across apps, devices, and the physical world through our hardware. You'll partner daily with research, engineering, product, infrastructure, and safety teams to decide what belongs in the next major training run, confirm whether it actually worked, and get the resulting improvements into the hands of people who rely on Hark every day.

Responsibilities

  • Design and build RL environments and tasks, spanning long-horizon planning, tool and computer use, multimodal interaction, and our hardware, that push models toward the capabilities a proactive personal assistant actually needs.
  • Architect the reward functions, graders, and task curricula that determine what a given environment actually teaches or reveals.
  • Scale environment infrastructure so thousands of tasks and rollouts can run in parallel across simulated settings and real agentic/hardware sessions.
  • Shape decisions on our biggest training runs and get an early look at what Hark's intelligence can do next.
  • Build self-improvement loops where the model helps generate, grade, and refine its own training environments, cutting the time from idea to validated result.

Requirements

  • Strong background in machine learning, with hands-on experience training or fine-tuning large models - LLMs, multimodal, or equivalent systems.
  • Hands-on experience building RL environments, simulators, or task suites, such as Gym-style interfaces, game engines, browser or OS-level automation sandboxes, or robotics/hardware simulators, and using them to train or evaluate models.
  • Direct, hands-on work with LLMs and modern post-training techniques: RL, RLHF/RLAIF, reward modeling and graders, synthetic data pipelines, or building agentic/tool-using systems.
  • A track record on fuzzy problems, where the objective is loosely defined, the data is noisy, and getting to a good answer takes both judgment and hands-on engineering.
  • A strong point of view on what makes a personal assistant genuinely useful, trustworthy, and pleasant to rely on daily, not just a leaderboard number moving up.
  • Skill at turning a vague behavioral concern into a testable experiment: form the hypothesis, stand up the pipeline, run it, read the results, and decide the next move.
  • Ease operating across research, product, infrastructure, data, hardware, and safety teams, and translating clearly between each.

Bonus Qualifications

  • Experience with RL algorithms applied to language, code, or agentic settings: RLHF, DPO, GRPO, PPO, or similar paradigms.
  • Familiarity with agent benchmarks and evaluation environments (e.g., OSWorld 1.0/2.0, Toolathlon, GDPval etc).
  • Research- or publication-level work on reward model design, e.g., comparing outcome-based vs. process-based rewards, learned reward signals, or reward-hacking mitigations.
  • Experience with trajectory-based training, imitation learning, or data distillation from stronger models or human demonstrations.
  • Prior work on computer use, GUI agents, or multimodal tool-using systems.
  • Experience training or scaling models at 100B+ parameters, with attention to efficiency, stability, and GPU utilization.
  • Contributions to open-source ML projects or publications at top venues (NeurIPS, ICML, ICLR, EMNLP, COLM, etc.).

Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components and benefits depending on the specific role. This information will be shared if an employment offer is extended.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
Founding Engineer 1 day ago
$93k – $139k per year • In office • Full-Time • 3+ years exp • Munich
Python
Python
FastAPI
AI/ML
Fine-tuning
LLM
VLM
Apply
Founding Engineer 1 day ago
$180k – $250k per year • In office • Full-Time • 3+ years exp • New York
Node JS
Python
TypeScript
JavaScript
Node JS
BullMQ
Databases
Redis
AI/ML
AI Agents
LangGraph
LLM
Multimodal AI
LangChain
Anthropic
Deepgram
LiveKit
LLM Guardrails
OpenAI
Text-to-Speech
Frontend
Next.js
React.js
DevOps
Vercel
Apply
$67k – $160k per year (Estimated) • In office • Full-Time • France
C++
C++
PyTorch C++
TensorFlow C++
AI/ML
Computer Vision
Multimodal AI
OpenCV
PyTorch
TensorFlow
Edge AI
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
RLHF
Synthetic Data
DPO
GRPO
Post-training
PPO
AI Agents
Function Calling
Robotics
Imitation Learning
Reinforcement Learning
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
Synthetic Data
Post-training
Pre-training
AI Agents
Function Calling
Robotics
Reinforcement Learning
Apply
Interface Designer 5 days ago
$150k – $300k per year • In office • Full-Time • 5+ years exp • San Jose
AI/ML
Multimodal AI
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
AI/ML
DeepSpeed
LLM
Multimodal AI
Synthetic Data
Megatron-LM
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
AI/ML
DeepSpeed
Fine-tuning
Multimodal AI
Reinforcement Learning
RLHF
DPO
FSDP
GRPO
Megatron-LM
Post-training
PPO
SFT
AI Agents
Apply
$147k – $265k per year (Estimated) • In office • Full-Time • Folsom • San Jose
Apply
$117k – $255k per year (Estimated) • In office • Full-Time • 8+ years exp • Richardson • Boise • Folsom • San Jose
Verilog
AI/ML
Claude
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$116k – $253k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise • San Jose
Apply
$68k – $85k per year • In office • Full-Time • Master's Degree • San Jose
Python
AI/ML
AI Agents
DevOps
Amazon EC2
AWS
AWS Lambda
Bitbucket
CI/CD
CloudFormation
Docker
Git
Kubernetes
Terraform
Amazon S3
IAM
HPC
Cybersecurity
Least Privilege
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.