376,406open jobs
9,791companies
48,125added this week
Browse all
Salary
$300k – $500k per year
Location
In office (San Jose)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Hark was an online digital entertainment platform best known for its extensive library of short audio soundbites, video clips, and pop culture quotes. Launched in 2007, the website allowed users to browse, create, and share playable soundboards featuring memorable lines from movies, television shows, and political figures. While it grew into a popular destination for viral sound clips during the late 2000s and early 2010s, the platform has since ceased its original operations.

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role 

You'll own how Hark's models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You'll build and lead the team that does it. The ceiling on what our hardware can do is set here.

Responsibilities

  • Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device deployment, and own the recommendation hardware decisions are made against.
  • Work with the foundation model and audio ML teams to shape architectures that meet deployment constraints before training locks them in.
  • Build the low-level execution layer, custom kernels, runtime systems, and compiler paths that transformer workloads run through on target hardware.
  • Partner with silicon vendors and internal hardware teams to bring up new accelerators and get efficient transformer execution on them early.
  • Hire and lead a team of engineers on performance-critical software, and set the technical bar for the inference stack.

Requirements

  • 8-12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators.
  • Deep understanding of attention, KV-cache behavior, quantization effects, and memory bandwidth limits.
  • You've designed or optimized inference engines, distributed runtimes, or ML compilers, and you write the kernels yourself when it matters.
  • Experience leading teams on performance-critical software. You've set direction on a stack, not just contributed to one.
  • You've taken a model from a research checkpoint to running on constrained hardware in a product people use.

Bonus Qualifications

  • Hands-on experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Experience with speech, audio, or streaming multimodal inference where latency is perceptible to the user.
  • Contributions to open-source inference or compiler toolchains (TensorRT, ONNX Runtime, TVM, MLIR, and similar).

Compensation

The US base salary range for this full-time position is between $300,000 - $500,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
376,406 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
In office • Full-Time • 2+ years exp • Bachelor's Degree • Suzhou
C++
Python
AI/ML
AI Agents
Computer Vision
LangChain
LlamaIndex
LLM
Multimodal AI
Time Series Forecasting
Apply
Engineering Manager 6 hours ago
$37k – $82k per year (Estimated) • In office • Noida
Java
Python
AI/ML
AI Agents
LLM
Model Context Protocol
Apply
$41k – $102k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree
Bash
Go
Python
AI/ML
AI Agents
Claude
Claude Code
Cursor
LLM Guardrails
Model Context Protocol
OpenAI Codex
DevOps
AWS
Azure
CI/CD
Cloudflare
FinOps
GCP
GitHub
GitHub Actions
IAM
Jenkins
Kubernetes
Platform Engineering
Terraform
Terragrunt
Cybersecurity
Least Privilege
SOC 2
Zero Trust
Apply
$41k – $102k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Go
Python
AI/ML
AI Agents
Claude
Claude Code
Cursor
Model Context Protocol
OpenAI Codex
DevOps
AWS
Azure
CI/CD
Cloudflare
FinOps
GCP
GitHub
GitHub Actions
IAM
Jenkins
Kubernetes
Platform Engineering
Terraform
Terragrunt
Cybersecurity
Least Privilege
SOC 2
Zero Trust
Apply
$35k – $86k per year (Estimated) • Remote • Full-Time
C#
Java
C#
.NET
Java
Testcontainers
Databases
Apache Kafka
RabbitMQ
AI/ML
AI Agents
DevOps
Azure
Azure DevOps
CI/CD
Git
GitHub
GitHub Actions
GitLab
GitLab CI
Jenkins
QA
Pact
Playwright
Apply
$200k – $450k per year • In office • Full-Time • 8+ years exp • San Jose
C++
AI/ML
AI Agents
Multimodal AI
ONNX
TensorRT
Apply
$120k – $300k per year • In office • Full-Time • Bachelor's Degree • San Jose
MATLAB
AI/ML
AI Agents
Fine-tuning
Multimodal AI
Apply
$120k – $300k per year • In office • Full-Time • San Jose
C++
Python
AI/ML
AI Agents
Multimodal AI
DevOps
RTOS
Management
Jira
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
RLHF
Synthetic Data
DPO
GRPO
Post-training
PPO
AI Agents
Function Calling
Robotics
Imitation Learning
Reinforcement Learning
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
Synthetic Data
Post-training
Pre-training
AI Agents
Function Calling
Robotics
Reinforcement Learning
Apply
$122k – $195k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
Apply
$91k – $183k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Longmont • San Jose
Python
DevOps
CI/CD
Apply
$180k – $225k per year • In office • Full-Time • 15+ years exp • Bachelor's Degree • San Jose
Apply
$216k – $270k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Jose
AI/ML
AI Agents
Apply
$145k – $180k per year • In office • Full-Time • 5+ years exp • San Jose
Verilog
Apply
See all jobs
This is one of many
376,406 more open roles from verified company boards, updated every day.