375,679open jobs
9,756companies
47,886added this week
Browse all
Salary
$200k – $450k per year
Location
In office (San Jose)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Hark was an online digital entertainment platform best known for its extensive library of short audio soundbites, video clips, and pop culture quotes. Launched in 2007, the website allowed users to browse, create, and share playable soundboards featuring memorable lines from movies, television shows, and political figures. While it grew into a popular destination for viral sound clips during the late 2000s and early 2010s, the platform has since ceased its original operations.

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role 

You'll make Hark's models run fast on the hardware we ship. That means writing the kernels, building the runtime paths, and profiling transformer workloads on DSPs, NPUs, and other constrained targets until they hit the latency and power budgets our devices are built around. This is hands-on systems work close to the metal, on a small team where the code you write is what users feel as response time.

Responsibilities

  • Write and optimize the low-level kernels and runtime paths that transformer workloads execute through on target silicon.
  • Decide how multiple models share limited memory and power i.e. residency, scheduling, and swap behavior across concurrent workloads.
  • Profile models on real hardware, find the bottlenecks, and close the gap between theoretical and delivered performance.
  • Take models from full precision to INT8/INT4  and get them running within per-product size, latency, and power budgets.
  • Get transformer workloads executing efficiently on new accelerators as they come online, working alongside the hardware team.
  • Feed real deployment constraints back to the model teams so architecture decisions account for what the hardware can actually do.

Requirements

  • 4-8+  years writing performance-critical software, with hands-on optimization on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C/C++ and comfort with SIMD, custom kernels, memory layout, and the profiling tools that go with them.
  • You understand attention, KV-cache behavior, and where transformer inference actually spends its time and memory bandwidth.
  • You reason in compute, memory, and power budgets, and you've optimized against them rather than around them.
  • You've had a model you optimized run in a product on constrained hardware.

Bonus Qualifications

  • Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
  • Background in speech or audio inference, where latency is perceptible to the user.

Compensation

The US base salary range for this full-time position is between $200,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
375,679 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
In office • Full-Time • 2+ years exp • Bachelor's Degree • Suzhou
C++
Python
AI/ML
AI Agents
Computer Vision
LangChain
LlamaIndex
LLM
Multimodal AI
Time Series Forecasting
Apply
Engineering Manager 5 hours ago
$37k – $82k per year (Estimated) • In office • Noida
Java
Python
AI/ML
AI Agents
LLM
Model Context Protocol
Apply
$41k – $102k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree
Bash
Go
Python
AI/ML
AI Agents
Claude
Claude Code
Cursor
LLM Guardrails
Model Context Protocol
OpenAI Codex
DevOps
AWS
Azure
CI/CD
Cloudflare
FinOps
GCP
GitHub
GitHub Actions
IAM
Jenkins
Kubernetes
Platform Engineering
Terraform
Terragrunt
Cybersecurity
Least Privilege
SOC 2
Zero Trust
Apply
$41k – $102k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Go
Python
AI/ML
AI Agents
Claude
Claude Code
Cursor
Model Context Protocol
OpenAI Codex
DevOps
AWS
Azure
CI/CD
Cloudflare
FinOps
GCP
GitHub
GitHub Actions
IAM
Jenkins
Kubernetes
Platform Engineering
Terraform
Terragrunt
Cybersecurity
Least Privilege
SOC 2
Zero Trust
Apply
$35k – $86k per year (Estimated) • Remote • Full-Time
C#
Java
C#
.NET
Java
Testcontainers
Databases
Apache Kafka
RabbitMQ
AI/ML
AI Agents
DevOps
Azure
Azure DevOps
CI/CD
Git
GitHub
GitHub Actions
GitLab
GitLab CI
Jenkins
QA
Pact
Playwright
Apply
$300k – $500k per year • In office • Full-Time • San Jose
AI/ML
AI Agents
Multimodal AI
ONNX
Quantization
TensorRT
Apply
$120k – $300k per year • In office • Full-Time • Bachelor's Degree • San Jose
MATLAB
AI/ML
AI Agents
Fine-tuning
Multimodal AI
Apply
$120k – $300k per year • In office • Full-Time • San Jose
C++
Python
AI/ML
AI Agents
Multimodal AI
DevOps
RTOS
Management
Jira
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
RLHF
Synthetic Data
DPO
GRPO
Post-training
PPO
AI Agents
Function Calling
Robotics
Imitation Learning
Reinforcement Learning
Apply
$180k – $450k per year • In office • Full-Time • San Jose
Python
AI/ML
Fine-tuning
Knowledge Distillation
LLM
Multimodal AI
PyTorch
Reinforcement Learning
Synthetic Data
Post-training
Pre-training
AI Agents
Function Calling
Robotics
Reinforcement Learning
Apply
$122k – $195k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
Apply
$91k – $183k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Longmont • San Jose
Python
DevOps
CI/CD
Apply
$180k – $225k per year • In office • Full-Time • 15+ years exp • Bachelor's Degree • San Jose
Apply
$216k – $270k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Jose
AI/ML
AI Agents
Apply
$145k – $180k per year • In office • Full-Time • 5+ years exp • San Jose
Verilog
Apply
See all jobs
This is one of many
375,679 more open roles from verified company boards, updated every day.