368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$180k – $450k per year
Location
In office (San Jose)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Hark was an online digital entertainment platform best known for its extensive library of short audio soundbites, video clips, and pop culture quotes. Launched in 2007, the website allowed users to browse, create, and share playable soundboards featuring memorable lines from movies, television shows, and political figures. While it grew into a popular destination for viral sound clips during the late 2000s and early 2010s, the platform has since ceased its original operations.

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

We are looking for a Member of Technical Staff, Post-Training to lead the development of post-training strategies that define how our models acquire coding, computer use, and agentic capabilities at scale.

This role sits at the frontier of a rapidly emerging discipline - one where reinforcement learning, simulation, and large-scale model training converge to produce agents that can reason, plan, and act over long horizons. There is no established playbook here. We're looking for researchers and engineers who can bring rigor and creativity from adjacent fields - RL, robotics, game-playing systems, compiler tooling, formal verification, or program synthesis - and apply them to the next generation of coding and agentic AI.

Responsibilities

  • Design and implement post-training strategies,  primarily RL-based,  to develop strong coding agents capable of multi-step reasoning, tool use, and long-horizon task completion.
  • Build and scale simulation and scaffolding environments for agentic RL: code execution sandboxes, computer use environments, tool-calling harnesses, and verifiable reward signals.
  • Develop reward modeling pipelines - including outcome-based, execution-based, and process-based reward signals - and iterate on them based on training dynamics.
  • Scale synthetic data generation and trajectory distillation pipelines that feed RL training and improve sample efficiency.
  • Design and run rigorous ablations to understand how algorithm choice, data mixture, reward shaping, and scale interact in the agentic setting.
  • Build evaluation frameworks grounded in real agent tasks - code correctness, execution success, multi-step tool use - to measure progress and guide iteration.
  • Collaborate with mid-training, infrastructure, and product teams to translate research insights into durable improvements on the model.

Requirements

  • Strong background in machine learning, with hands-on experience training or fine-tuning large models - LLMs, multimodal, or equivalent systems.
  • Deep understanding of reinforcement learning: policy optimization, reward design, exploration, and the interplay between environment design and agent behavior.
  • Experience building or working within simulation or execution environments (e.g., code interpreters, sandboxed execution, game environments, robotics simulators).
  • Proven ability to design and execute rigorous experiments, with strong intuition for diagnosing training failures and scaling bottlenecks.
  • Proficiency in Python and PyTorch; comfort working across research and systems code.
  • Ability to work in a fast-moving, research-forward environment where the right approach is often unknown at the outset.

We expect strong candidates to come from a range of backgrounds - RL research, robotics, competitive programming systems, compilers, formal methods, or large-scale ML - rather than post-training specifically. The field is new enough that directly relevant experience is rare; what matters is depth, rigor, and transferability.

Bonus Qualifications

  • Experience with RL algorithms applied to language or code: RLHF, DPO, GRPO, PPO, or similar paradigms in the LLM setting.
  • Familiarity with coding agent benchmarks and evaluation environments (e.g., SWE-bench, HumanEval, LiveCodeBench, competitive programming judges).
  • Background in reward modeling - outcome-based, process-based, or learned reward signals.
  • Experience with trajectory-based training, imitation learning, or data distillation from stronger models or human demonstrations.
  • Prior work on computer use, GUI agents, or tool-using LLMs (e.g., OSWorld, WebArena-style tasks).
  • Experience training or scaling models at 10B+ parameters, with attention to efficiency, stability, and GPU utilization.
  • Contributions to open-source ML projects or publications at top venues (NeurIPS, ICML, ICLR, EMNLP, COLM, etc.).

Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components and benefits depending on the specific role. This information will be shared if an employment offer is extended.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
Staff Data Engineer 6 hours ago
$160k – $200k per year • In office • Full-Time • 6+ years exp • Chicago
Python
SQL
Databases
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
Arize Phoenix
AutoGen
AWS Bedrock AgentCore
CrewAI
dbt
Fine-tuning
Function Calling
LangChain
LangGraph
LangSmith
LLM
LLM Evaluation
LLM Guardrails
Model Context Protocol
Prefect
Prompt Engineering
RAG
Semantic Kernel
Semantic Search
Semantic Search
Weights & Biases
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
GitHub Actions
Kubernetes
Vector
Analytics
ETL/ELT
Apply
$118k – $206k per year • In office • Full-Time • Bachelor's Degree • Chicago
JavaScript
Python
SQL
TypeScript
Databases
Databricks
AI/ML
LLM
LLMOps
DevOps
AIOps
Azure
Platform Engineering
Analytics
Power BI
Marketing
Amplitude
Apply
$47k – $103k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune • Bengaluru
Python
AI/ML
AI Agents
LLM
RAG
Vertex AI
DevOps
AWS
Azure
GCP
Apply
Founding Engineer 6 hours ago
$120k – $200k per year • Equity 0.5–1% • In office • Full-Time • 3+ years exp • San Francisco
TypeScript
AI/ML
AI Agents
Claude
LLM
DevOps
GitHub
Management
Linear
Slack
Apply
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
Synthetic Data
Post-training
Pre-training
AI Agents
Function Calling
Robotics
Reinforcement Learning
Apply
Interface Designer 5 days ago
$150k – $300k per year • In office • Full-Time • 5+ years exp • San Jose
AI/ML
Multimodal AI
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
AI/ML
DeepSpeed
LLM
Multimodal AI
Synthetic Data
Megatron-LM
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
AI/ML
DeepSpeed
Fine-tuning
Multimodal AI
Reinforcement Learning
RLHF
DPO
FSDP
GRPO
Megatron-LM
Post-training
PPO
SFT
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
AI/ML
Multimodal AI
Speech Recognition
Synthetic Data
Text-to-Speech
AI Agents
Apply
$78k – $130k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Cary • San Jose
Analytics
Power BI
Apply
$147k – $265k per year (Estimated) • In office • Full-Time • Folsom • San Jose
Apply
$117k – $255k per year (Estimated) • In office • Full-Time • 8+ years exp • Richardson • Boise • Folsom • San Jose
Verilog
AI/ML
Claude
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$116k – $253k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise • San Jose
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.