368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$120k – $300k per year
Location
In office (San Jose)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Hark was an online digital entertainment platform best known for its extensive library of short audio soundbites, video clips, and pop culture quotes. Launched in 2007, the website allowed users to browse, create, and share playable soundboards featuring memorable lines from movies, television shows, and political figures. While it grew into a popular destination for viral sound clips during the late 2000s and early 2010s, the platform has since ceased its original operations.

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role 

We are looking for an On-Device Research Engineer to compress large audio and multimodal models into student models that meet the size, latency, and power budgets of our shipping hardware. This role sits between training and production. You will take teacher models from our research pipeline and produce student models that run on DSP, NPU, and microcontroller targets across our product line. You will own distillation, quantization, and architecture-aware compression as a first-class work-stream.

Responsibilities

  • Design and execute distillation strategies (response, feature, and self-distillation) to compress teacher models into deployable students
  • Apply quantization (PTQ and QAT), pruning, and architecture search to hit per-product size, latency, and power budgets
  • Build a reusable distillation and compression toolchain that the broader audio ML team can adopt across model families
  • Partner with the broader audio ML team on training pipelines and with the runtime team on deployment targets
  • Define accuracy retention and resource KPIs per product and track them through the release cycle
  • Profile compressed models on target hardware and iterate with DSP and runtime engineers on bottlenecks

Requirements

  • 3+ years of professional experience in model compression, distillation, quantization, or efficient deep learning
  • Strong fluency in PyTorch or TensorFlow and modern compression libraries
  • Hands-on experience taking models from full precision to fixed-point or int8 with controlled accuracy loss
  • Comfort working close to hardware and reasoning about compute, memory bandwidth, and power as design constraints
  • Track record of producing models that have shipped to constrained devices
  • Solid foundation in audio or sequence model architectures (CNNs, transformers, RNN-T, conformers)

Bonus Qualifications

  • Experience with Hexagon DSP, NPUs, Ambiq class MCUs, or similar
  • Experience with knowledge distillation at scale, including teacher-ensemble or multi-stage distillation
  • Familiarity with neural architecture search and hardware-aware NAS
  • Background shipping voice-first or far-field audio products
  • Contributions to open-source compression toolchains (TFLite, ONNX Runtime, AIMET, and similar)

Compensation

The US base salary range for this full-time position is between $120,000 - $300,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$58k – $171k per year (Estimated) • In office • Full-Time • 1+ year exp • Munich
Python
AI/ML
Computer Vision
ONNX
PyTorch
DevOps
Docker
SLURM
Apply
$67k – $160k per year (Estimated) • In office • Full-Time • France
C++
C++
PyTorch C++
TensorFlow C++
AI/ML
Computer Vision
Multimodal AI
OpenCV
PyTorch
TensorFlow
Edge AI
Apply
Founding Engineer 1 day ago
$180k – $250k per year • In office • Full-Time • 3+ years exp • New York
Node JS
Python
TypeScript
JavaScript
Node JS
BullMQ
Databases
Redis
AI/ML
AI Agents
LangGraph
LLM
Multimodal AI
LangChain
Anthropic
Deepgram
LiveKit
LLM Guardrails
OpenAI
Text-to-Speech
Frontend
Next.js
React.js
DevOps
Vercel
Apply
$120k – $250k per year • Equity 0.2–1% • In office • Full-Time • Master's Degree • Seattle
Python
AI/ML
Diffusion Models
Fine-tuning
PyTorch
Self-Supervised Learning
Apply
$121k – $282k per year (Estimated) • Equity 0.5–1% • In office • Full-Time • 1+ year exp • San Francisco
AI/ML
LLM
PyTorch
Transformers
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
RLHF
Synthetic Data
DPO
GRPO
Post-training
PPO
AI Agents
Function Calling
Robotics
Imitation Learning
Reinforcement Learning
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
Synthetic Data
Post-training
Pre-training
AI Agents
Function Calling
Robotics
Reinforcement Learning
Apply
Interface Designer 5 days ago
$150k – $300k per year • In office • Full-Time • 5+ years exp • San Jose
AI/ML
Multimodal AI
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
AI/ML
DeepSpeed
LLM
Multimodal AI
Synthetic Data
Megatron-LM
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • San Jose
AI/ML
DeepSpeed
Fine-tuning
Multimodal AI
Reinforcement Learning
RLHF
DPO
FSDP
GRPO
Megatron-LM
Post-training
PPO
SFT
AI Agents
Apply
$78k – $130k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Cary • San Jose
Analytics
Power BI
Apply
$147k – $265k per year (Estimated) • In office • Full-Time • Folsom • San Jose
Apply
$117k – $255k per year (Estimated) • In office • Full-Time • 8+ years exp • Richardson • Boise • Folsom • San Jose
Verilog
AI/ML
Claude
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$116k – $253k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise • San Jose
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.