441,244open jobs
15,342companies
64,215added this week
Browse all
Salary
$170k – $450k per year
Location
In office (San Jose)
Seniority
Staff · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Hark was an online digital entertainment platform best known for its extensive library of short audio soundbites, video clips, and pop culture quotes. Launched in 2007, the website allowed users to browse, create, and share playable soundboards featuring memorable lines from movies, television shows, and political figures. While it grew into a popular destination for viral sound clips during the late 2000s and early 2010s, the platform has since ceased its original operations.

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role 

You'll build the data infrastructure that turns raw signals into the training data Hark's models learn from, and the pipelines that keep it flowing at scale.

That means owning the full data engineering stack: ingestion, transformation, quality filtering, and delivery to training and evaluation systems. The models we ship are only as good as the data behind them, and this role owns that foundation.

This is a high-ownership role on a small team. You'll work directly with model researchers, data collection leads, and infrastructure engineers, and the systems you build will directly shape the quality and pace of model development.

Responsibilities

  • Design and build scalable data pipelines that ingest, process, and deliver training data across multiple modalities: text, audio, vision, and structured feedback signals.
  • Own the data infrastructure stack end-to-end: ingestion, transformation, deduplication, quality filtering, versioning, and delivery to model training and evaluation systems.
  • Collaborate closely with model researchers and data collection leads to understand data requirements and translate them into reliable, auditable pipelines.
  • Build tooling and frameworks that make it easy for the team to inspect, evaluate, and iterate on data quality. The insights surfaced should feed back into collection and curation decisions.
  • Define and enforce data quality standards. Instrument pipelines for correctness, freshness, and coverage. Catch regressions before they reach training.
  • Design data systems for reproducibility and scale. The pipelines you build need to handle growing volumes across modalities without becoming a bottleneck.
  • Identify gaps in the current stack and drive concrete improvements to throughput, quality, and reliability.

Requirements

  • Strong data engineering fundamentals. You are comfortable designing and operating large-scale batch and streaming pipelines, and you care about correctness and reliability.
  • Experience building data systems for machine learning. You understand the difference between a data pipeline for analytics and one that feeds model training, and you know what it takes to get the latter right.
  • Fluency with the modern data stack. You've worked with tools like Spark, Beam, or Flink, and you know how to make tradeoffs between them. Experience with data versioning systems (e.g., DVC, Delta Lake, Iceberg) is a strong plus.
  • Systems thinking. You reason about schema evolution, backfills, and failure modes before they become production incidents. You build for the day-2 case, not just the demo.
  • A quality instinct. You don't just move data. You understand what's in it, catch problems early, and close the feedback loop with the people who need clean data.
  • Strong communication. You can work closely with model researchers and engineers, explain data tradeoffs clearly, and make good decisions across team boundaries.
  • 5+ years of relevant data engineering experience. Experience at a fast-growing AI or research-driven company is a strong plus.

Bonus Qualifications

  • Experience building data infrastructure for large language model or multimodal model training.
  • Familiarity with multimodal data formats and processing pipelines (audio, video, image).
  • Experience with human feedback or preference data pipelines (RLHF, DPO, or similar).
  • Hands-on experience with data quality evaluation frameworks or annotation tooling.
  • Background in distributed systems, stream processing, or large-scale ETL.
  • Experience at a fast-moving AI lab or research-driven company.

Compensation

The US base salary range for this full-time position is between $170,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
441,244 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$20k – $49k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
Python
FastAPI
Databases
Databricks
AI/ML
LangGraph
AutoGen
LangChain
LlamaIndex
Vertex AI
Fine-tuning
Prompt Engineering
Multimodal AI
AI Agents
CrewAI
LLM
RAG
Edge AI
Multi-Agent Systems
DevOps
GCP
Azure
AWS
Apply
$16k – $45k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Master's Degree • Hyderabad
Python
SQL
AI/ML
XGBoost
Scikit-learn
AI Agents
TensorFlow
PyTorch
LLM
RAG
Agentic Workflows
Apply
$53k – $116k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Lisbon
JavaScript
Databases
Snowflake
AI/ML
AI Agents
Agentic Workflows
Marketing
Zendesk
GA4
Apply
$64k – $145k per year (Estimated) • Remote/Hybrid • Berlin
Go
JavaScript
Kotlin
TypeScript
AI/ML
AI Agents
Frontend
React.js
DevOps
Terraform
CI/CD
Docker
QA
Selenium
Cypress
Apply
In office • Full-Time • Bachelor's Degree • Cork
AI/ML
AI Agents
Apply
$120k – $300k per year • In office • Full-Time • Bachelor's Degree • San Jose
AI/ML
Multimodal AI
AI Agents
Apply
$300k – $500k per year • In office • Full-Time • San Jose
AI/ML
Quantization
Multimodal AI
AI Agents
ONNX
TensorRT
ONNX Runtime
MLIR
Apache TVM
KV Cache
Apply
$150k – $350k per year • In office • Full-Time • 8+ years exp • San Jose
Rust
Kotlin
C++
Dart
Swift
AI/ML
Multimodal AI
AI Agents
Mobile
Jetpack Compose
UIKit
SwiftUI
Core Animation
Flutter
Design
Figma
Rive
Apply
$200k – $450k per year • In office • Full-Time • 8+ years exp • San Jose
C++
AI/ML
Multimodal AI
AI Agents
ONNX
TensorRT
ONNX Runtime
MLIR
Apache TVM
KV Cache
Apply
$120k – $300k per year • In office • Full-Time • Bachelor's Degree • San Jose
MATLAB
AI/ML
Fine-tuning
Multimodal AI
AI Agents
Apply
$209k – $240k per year • In office • Bachelor's Degree • San Jose
Python
C++
AI/ML
InfiniBand
NVLink
Apply
$84k – $189k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Guaynabo • San Jose • Panama • Mexico City
Analytics
Power BI
Microsoft Excel
Apply
$110k – $282k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Jose
Apply
$166k – $331k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • San Jose
Python
SystemVerilog
Perl
AI/ML
NVLink
Chips/EDA
UVM
Synopsys ZeBu
Cadence Palladium
Siemens Veloce
Synopsys HAPS
Apply
$173k – $346k per year (Estimated) • In office • 10+ years exp • San Jose
AI/ML
NVLink
Chips/EDA
Ansys HFSS
Cadence Allegro
Cadence Sigrity
Keysight ADS
Ansys SIwave
Apply
See all jobs
This is one of many
441,244 more open roles from verified company boards, updated every day.