1,035,818open jobs
60,918companies
173,892added this week
Browse all
Salary
≈ $210k – $380k per year (Estimated)
Location
In office (Mountain View)
Seniority
Staff · 8+ years exp

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Oct 1, 2026.

Overview
Company
Impact
Profile match
Google DeepMind is the artificial intelligence research and engineering division of Alphabet, formed in 2023 by merging the London laboratory founded in 2010 with the Google Brain team. It builds frontier foundation models such as the Gemini family, generative media systems including Veo and Imagen, and scientific tools like AlphaFold, whose protein structure predictions earned a share of the 2024 Nobel Prize in Chemistry. The unit pairs long-horizon research on reinforcement learning and general intelligence with product delivery across Google Search, Workspace, Android and Google Cloud.

About the job

At DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.

As an Inference Performance Engineer, you will push the boundaries of AI model execution at scale. In this role, you will be at the forefront of making large-scale AI inference faster, cheaper, and more efficient. You will analyze the entire inference stack to identify critical bottlenecks and drive systemic improvements. By combining deep systems profiling, benchmarking, and first-principles problem solving, your work will directly maximize hardware throughput, reduce cost-to-serve, and empower our cross-functional teams to make data-driven capacity and latency tradeoffs.

Artificial intelligence will be one of humanity’s most transformative inventions. At DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities

  • Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.
  • Design and implement inference optimization techniques.
  • Investigate and resolve complex model inference performance bottlenecks across the stack.
  • Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.
  • Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.

Qualifications

Minimum qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.
  • 8 years of experience in software development.
  • Experience in Python and C++, including navigating, debugging, and modifying serving codebases.
  • Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.

Preferred qualifications:

  • Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).
  • Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.
  • Experience with observability and reliability for large distributed systems.
  • Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,035,818 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Mountain View
≈ $136k – $228k per year (Estimated) • Remote (United States) • Full-Time • 6+ years exp • Bachelor's Degree
Python
JavaScript
TypeScript
Python
Flask
FastAPI
Django
Django REST Framework
Databases
PostgreSQL
DynamoDB
Frontend
Webpack
Next.js
React.js
Vite
Astro
Turbopack
DevOps
CI/CD
Git
AWS
Docker
AWS Fargate
AWS Lambda
Amazon ECS
API Gateway
Apply
$113k – $138k per year • In office • TS/SCI • Full-Time • 5+ years exp • Bachelor's Degree • Seal Beach
Python
Java
TypeScript
SQL
Python
FastAPI
Databases
PostgreSQL
DevOps
CI/CD
Kubernetes
GitLab
Management
Agile
Apply
$165k – $265k per year • Equity • In office • 5+ years exp • Bachelor's Degree • Hawthorne
C++
C++
CMake
AI/ML
CUDA Toolkit
CUDA
DevOps
CI/CD
Linux
Apply
$231k – $323k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Denver
DevOps
Nomad
Docker
Kubernetes
Service Mesh
Apply
$169k – $216k per year • Equity • In office • 5+ years exp • Bachelor's Degree • El Segundo
Python
C++
AI/ML
Time Series Forecasting
DevOps
Ansible
CI/CD
Docker
Kubernetes
Apply
≈ $13k – $35k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Gurgaon
Python
Java
PowerShell
AI/ML
AI Agents
DevOps
GCP
Azure
AWS
Kubernetes
Apply
≈ $43k – $113k per year (Estimated) • In office
Python
AI/ML
Machine Learning
DevOps
Azure DevOps
Azure
CI/CD
GitOps
Git
Docker
Kubernetes
Platform Engineering
Management
Miro
Confluence
Jira
Microsoft Teams
Agile
Scrum
Kanban
Apply
≈ $24k – $43k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Pune
Python
Java
SQL
Scala
Python
pySpark
Databases
Databricks
Delta Lake
MS SQL
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Hadoop
Spark
Airflow
MLFlow
Pandas
DevOps
GCP
CI/CD
Git
Incident Management
SLI/SLO/SLA
Analytics
ETL/ELT
Management
Agile
Apply
≈ $8k – $21k per year (Estimated) • Remote (India) • Full-Time • Bachelor's Degree • Chennai
Python
SQL
DevOps
Unix
Analytics
ETL/ELT
Management
Microsoft Office
Apply
≈ $20k – $44k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Chennai • Pune
Python
SQL
Management
Agile
Apply
≈ $158k – $292k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Cambridge
AI/ML
AI Agents
Edge AI
Machine Learning
Apply
≈ $217k – $392k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Mountain View
AI/ML
TPU
Edge AI
Apply
≈ $176k – $326k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Mountain View
AI/ML
AI Agents
Edge AI
Machine Learning
Apply
≈ $142k – $283k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Mountain View
Python
C++
AI/ML
Knowledge Distillation
AI Agents
Gemini
Edge AI
Model Distillation
Machine Learning
Apply
≈ $168k – $311k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Mountain View
Python
Go
JavaScript
Kotlin
TypeScript
C++
Swift
AI/ML
AI Agents
Google AI Studio
Red Teaming
Apply
≈ $197k – $355k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Mountain View
C++
AI/ML
AI Agents
Gemini
Machine Learning
Apply
≈ $128k – $255k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Mountain View
Java
Kotlin
Swift
Databases
Google Cloud Spanner
AI/ML
Gemini
Time Series Forecasting
DevOps
Rest API
Apply
≈ $208k – $375k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Mountain View
AI/ML
Fine-tuning
Reinforcement Learning
DevOps
GCP
Apply
$58k – $68k per year • In office • Mountain View
Apply
$58k – $68k per year • In office • Mountain View
Apply
See all jobs
This is one of many
1,035,818 more open roles from verified company boards, updated every day.