812,549open jobs
52,293companies
130,976added this week
Browse all
Salary
$170k – $280k per year
Location
In office (New York)
Seniority
Staff
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 17, 2026.

Overview
Company
Impact
Profile match
Real human ratings, in a single API call. The human evaluation layer for voice, speech, and conversational AI.

Hume AI is looking for a systems-oriented engineer to own the path from trained model to production inference. Join us in the heart of New York City and contribute to our endeavor to ensure that AI is guided by human values, the most pivotal challenge-and opportunity-of the 21st century.

About Us

Hume AI is a Series B startup dedicated to building artificial intelligence that is directly optimized for human well-being. As the first company to release speech language models, we’re focused on expanding our research to encompass audio understanding models and evaluation platforms for enterprises.

Our goal is to enable a future in which technology draws on an understanding of human emotional expression to better serve human goals. As part of our mission, we also conduct groundbreaking scientific research, publish in leading scientific journals like Nature, and support a non-profit, The Hume Initiative, that has released the first concrete ethical guidelines for empathic AI (www.thehumeinitiative.org). You can learn more about us on our website (https://hume.ai/) and read about us in WIRED, Forbes, and Venturebeat.

About the Role

As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference: graph export, engine compilation, runtime integration, serving contracts, client libraries, artifact verification, and the performance and correctness of what runs in production.

You will work closely with research scientists, machine learning engineers, backend engineers, and the Data Plane team to bring new models and inference capabilities into production.

This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy, make sound architectural decisions, and own systems throughout their lifecycle-from initial design and implementation through deployment, observability, optimization, and production support.

About the Role

As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference. This includes graph export, engine compilation, runtime integration, serving contracts, artifact verification, and production performance and correctness.

You will also own Hume’s internal inference platform, which serves many of our in-house models across products. This includes deployment, routing, load balancing, health checking, observability, capacity management, and safe model rollout.

You will work closely with research scientists, machine learning engineers, backend engineers, and product teams to turn new model capabilities into reliable production systems.

This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy and own systems from design through production.

What You’ll Do

  • Own the path from trained checkpoint to served request, including graph export, engine compilation, runtime integration, and serving configuration.

  • Build and evolve Hume’s internal inference platform for serving multiple models across products and workloads.

  • Design and operate serving infrastructure, including routing, load balancing, health checking, autoscaling, capacity management, and failure handling.

  • Build reproducible, versioned inference artifacts and tooling for validation, deployment, promotion, and rollback.

  • Build verification gates that catch numerical, behavioral, and performance regressions before production.

  • Design and maintain internal client libraries and standardized serving contracts.

  • Profile and optimize latency, throughput, memory usage, batching, scheduling, and accelerator utilization.

  • Diagnose production issues across application, runtime, container, networking, driver, and hardware boundaries.

  • Improve observability, resilience, testability, and operational safety across the inference stack.

  • Write clear technical documentation for the systems and APIs you build.

What You’ll Bring

  • Significant professional experience building server-side, infrastructure, distributed, or systems software.

  • Strong Linux fundamentals and hands-on experience troubleshooting and profiling production systems.

  • Professional experience with at least one systems-oriented language such as Rust, Go, C, or C++.

  • Experience with distributed systems concepts such as load balancing, health checking, failure recovery, observability, and capacity management.

  • A practical understanding of neural-network execution, including computation graphs, tensor shapes, data types, and accelerator execution.

  • Experience profiling and optimizing production systems.

  • Comfort working across languages and tooling, including Python for model export, validation, and integration workflows.

  • Strong ownership, independent technical judgment, and clear written and verbal communication.

  • The ability to use AI-assisted coding tools effectively while retaining the ability to explain, validate, debug, and modify the result independently.

Bonus Points

  • Experience building or operating shared model-serving or inference platforms.

  • Experience with GPU inference, CUDA, ONNX, TensorRT, PyTorch, Triton, vLLM, or similar systems.

  • Experience diagnosing numerical correctness issues such as precision loss, numerical drift, or nondeterminism.

  • Experience building high-performance client libraries, SDKs, or networked systems.

  • Experience with containers, CI/CD, or deploying software into on-premises or customer-managed environments.

  • Contributions to systems, inference, distributed-systems, or machine-learning infrastructure open-source projects.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
812,549 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
New York
≈ $162k – $278k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp • Bachelor's Degree
JavaScript
TypeScript
Node JS
Node JS
Fastify
AI/ML
Copilot
Cursor
Claude Code
Frontend
React.js
Ant Design
Zod
DevOps
AWS
Amazon ECS
Apply
Sr. Software Engineer 23 days ago
≈ $143k – $237k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp
Python
Java
AI/ML
Claude Code
Human-in-the-Loop
DevOps
AWS
Apply
$160k – $240k per year • In office • 4+ years exp • Bachelor's Degree • New York
Python
Python
FastAPI
Databases
Redis
Apache Kafka
AI/ML
Flink
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Amazon S3
Management
Slack
Microsoft Teams
Apply
≈ $123k – $240k per year (Estimated) • Hybrid • 8+ years exp • Bachelor's Degree • Houston
C#
C#
.NET
DevOps
Rest API
Azure
CI/CD
Management
Agile
Apply
$168k – $227k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
AI/ML
AI Agents
Apply
≈ $25k – $62k per year (Estimated) • In office • Noida
Python
Go
Java
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Cybersecurity
GDPR
HIPAA
Apply
≈ $116k – $215k per year (Estimated) • Remote (United States) • Full-Time • 3+ years exp
Go
TypeScript
DevOps
GCP
Kubernetes
GitHub
Linux
Unix
Apply
≈ $126k – $243k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • PhD • Cambridge
Python
AI/ML
Multimodal AI
AI Agents
Time Series Forecasting
Machine Learning
DevOps
Bitbucket
GitHub
Apply
≈ $122k – $210k per year (Estimated) • Remote (United States) • Full-Time
Python
SQL
Databases
PostgreSQL
Snowflake
Databricks
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Claude Code
dbt
Great Expectations
LLM
DevOps
Terraform
AWS
Cybersecurity
GDPR
HIPAA
Analytics
Power BI
Fivetran
Data Vault
Apply
≈ $23k – $44k per year (Estimated) • Hybrid • Full-Time • Saint Petersburg
C++
DevOps
Windows
Cybersecurity
Dallas Lock
Apply
$150k – $250k per year • Hybrid • Full-Time • New York
Python
Rust
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
ONNX
TensorRT
PyTorch
CUDA
Machine Learning
DevOps
CI/CD
Linux
Apply
$150k – $250k per year • Hybrid • Full-Time • New York
Go
Rust
C++
Databases
PostgreSQL
DynamoDB
AI/ML
Machine Learning
DevOps
Rest API
gRPC
GCP
OpenTelemetry
Datadog
Envoy
Prometheus
WebSockets
HAProxy
CI/CD
Kubernetes
Nginx
Linux
Apply
$150k – $250k per year • In office • Full-Time • New York
JavaScript
TypeScript
Frontend
Next.js
React.js
DevOps
WebRTC
WebSockets
Cybersecurity
Auth0
QA
Playwright
Apply
Business Development 12 days ago
$75k – $100k per year • In office • Full-Time • New York
Marketing
LinkedIn
Apply
Research Engineer 12 days ago
$190k – $260k per year • In office • Full-Time • 2+ years exp • PhD • New York
Python
AI/ML
Fine-tuning
Multimodal AI
Speech Recognition
PyTorch
Post-training
Text-to-Speech
Machine Learning
Apply
Python Developer 1 day ago
≈ $141k – $260k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • New York
Python
SQL
Apply
Java Developer 1 day ago
≈ $134k – $247k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • New York
Java
Java
Hibernate
Apply
$60k – $80k per year • In office • Internship • New York
AI/ML
Model Context Protocol
Design
Canva
Apply
$110k – $140k per year • Equity 0.1–0.3% • In office • Full-Time • New York
Design
Canva
Apply
$150k – $240k per year • Equity 0.1–0.3% • In office • Full-Time • 6+ years exp • New York
AI/ML
Claude
Apply
See all jobs
This is one of many
812,549 more open roles from verified company boards, updated every day.