368,611open jobs
9,439companies
50,719added this week
Browse all
Location
In office
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Inworld AI is a company founded in 2021 that builds a runtime for artificial intelligence characters and agents in games and interactive media. Its engine handles dialogue, memory, emotion and safety so studios can drop believable non-player characters into a title without building model infrastructure. The company works with game developers and consumer platforms and has partnered with Microsoft and Nvidia.

About Inworld

Inworld is a research lab of top researchers and engineers, building the world’s top-ranked realtime voice models.

Today our models are the #1 ranked realtime voice models in the world. They are used to power the largest consumer-facing AI applications available, across categories like health, fitness, learning, therapy, companions, customer experience and media; representing 100s of millions of end users. Our work spans areas like research and development of state-of-the-art models, optimizing realtime inference, and creating best-in-class APIs and products that allow developers to engage their users.

We’ve raised more than $125M from Lightspeed, Section 32, Kleiner Perkins, Microsoft’s M12 venture fund, Founders Fund, Meta and Stanford, among others. Our technology has powered experiences from companies such as NVIDIA, Microsoft Xbox, Niantic, Logitech Streamlabs, Wishroll, Little Umbrella and Bible Chat. We’ve also been recognized by CB Insights as one of the 100 most promising AI companies globally and have been named one of LinkedIn’s Top 10 Startups in the USA.

Who We're Looking For

A year ago, reliably working agentic systems and sub-second multimodal inference at scale barely existed. Nobody has a decade of experience here. So we're not screening for a resume template - we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood.

Experience We Find Useful

You don't need all of this. But you need enough to make a case.

  • Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.

  • Model Acceleration. Hands-on experience with quantization, distillation, caching strategies , continuous batching, paged attention, and speculative decoding.

  • High-Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimized Python. You know how to profile code and squeeze every ounce of performance out of NVIDIA GPUs.

  • Distributed Systems & Scaling. Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and reliably handling thousands of concurrent connections.

  • Public work. Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups.

  • Full-cycle ownership. You can take a model from the research team, containerize it, optimize its serving, and ensure it runs reliably in production.

  • Background. PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.

  • Professional fluency in English (written and spoken) is required, as you will be collaborating daily with our US-based leadership and engineering teams.

Who Thrives Here

  • You don’t need a roadmap to start walking; you’re comfortable picking a direction and building the map as you go.

  • You believe engineering isn't finished until it’s shipped and stable. You have a bias for impact over purely theoretical optimizations.

  • You don't just ship code; you obsess over the why. You’re the first to question an architecture if you think there’s a better way to solve the core latency or throughput problem.

  • You aren't satisfied with "the PM said so." You thrive on deep context and want to understand the fundamental logic behind every decision we make.

What Working Here Is Like

We hand you unclear problems and expect you to make them clear. We value engineers who say "I don't know yet" and then design the benchmark or prototype that finds out. We treat performance, latency, and reliability as first-class product features, not a box to check before launch. Impact comes before everything else, though we support sharing work and open-source contributions that move the field forward. Your work should be visible. Flat structure, fast iterations, minimal process theater.

For candidates interested in relocating to the San Francisco Bay Area in the future, full U.S. visa and relocation support may be available, subject to business needs and applicable legal and work authorization requirements.

Inworld Jobs Privacy

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$32k – $72k per year (Estimated) • In office • Full-Time • 6+ years exp • Bengaluru
C++
Java
Python
YARA
Databases
Amazon Aurora
DevOps
Azure
GCP
Kubernetes
Cybersecurity
MITRE ATT&CK
Suricata
YARA
Zeek
Apply
HMAX Senior Developer 11 hours ago
$42k – $86k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Genoa • Turin
Java
Java
Spring Boot
Databases
Apache Kafka
InfluxDB
PostgreSQL
DevOps
Azure
CI/CD
Docker
Grafana
Kubernetes
Apply
$18k – $51k per year (Estimated) • Remote • Moscow
Bash
Python
Databases
ClickHouse
PostgreSQL
AI/ML
Feature Store
Hadoop
LLM
DevOps
Ansible
CI/CD
Docker
GitLab
Kubernetes
Apply
$71k – $154k per year (Estimated) • In office • Full-Time • Dublin
Java
Databases
Apache Kafka
AI/ML
Copilot
LLM
LLM Guardrails
DevOps
AWS
CI/CD
Kubernetes
GitHub
Apply
$75k – $198k per year (Estimated) • In office • Full-Time • Dublin
Java
Kotlin
Java
Spring Boot
Databases
Amazon Aurora
Apache Kafka
PostgreSQL
DevOps
AWS
CI/CD
Dynatrace
GitLab CI
Kubernetes
Trunk-Based Development
GitLab
Apply
$170k – $250k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Mountain View
JavaScript
Python
TypeScript
AI/ML
LLM
Quantization
Text-to-Speech
DevOps
WebRTC
WebSockets
Management
n8n
Zapier
Apply
$140k – $334k per year (Estimated) • Remote • Full-Time • 10+ years exp • PhD
C++
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
Quantization
Ray
vLLM
AI Agents
DevOps
Kubernetes
Apply
$190k – $271k per year • In office • Full-Time • 10+ years exp • PhD
C++
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
Quantization
Ray
vLLM
AI Agents
DevOps
Kubernetes
Apply
$105k – $251k per year (Estimated) • In office • Full-Time • 10+ years exp • PhD
C++
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
Quantization
Ray
vLLM
AI Agents
DevOps
Kubernetes
Apply
$270k – $500k per year • In office • Full-Time • 10+ years exp • PhD • Mountain View
C++
Python
Rust
AI/ML
CUDA
CUDA Toolkit
Knowledge Distillation
LLM
Multimodal AI
Quantization
Ray
vLLM
AI Agents
DevOps
Kubernetes
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.