{"id":1580080,"url":"https://alion.io/job/google-deepmind-staff-software-engineer-inference-performance-optimization-genai","title":"Staff Software Engineer, Inference Performance Optimization, GenAI","company":{"id":4,"name":"Google DeepMind","domain":"deepmind.google","url":"https://alion.io/company/deepmind","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Career site","truth_index":null},"role":"Backend","role_family":"Backend","seniority":"staff","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Mountain View, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":210000,"max_usd":380000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":976},"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"C++","optional":false},{"name":"Edge AI","optional":false},{"name":"Machine Learning","optional":false},{"name":"Python","optional":false},{"name":"LLM","optional":true},{"name":"PyTorch","optional":true},{"name":"PyTorch C++","optional":true},{"name":"Quantization","optional":true},{"name":"SGLang","optional":true},{"name":"TensorRT","optional":true},{"name":"TensorRT-LLM","optional":true},{"name":"TPU","optional":true},{"name":"vLLM","optional":true}],"status":"live","first_seen_at":"2026-10-01T10:39:36Z","employer_posted_date":"2026-10-01","last_verified_at":"2026-10-02T02:10:56Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"About the job\nAt DeepMind our mission is to build the world's first general-purpose learning agent. Central to this mission is the complex task of measuring the intelligence of our prototypes. As a Software Engineer, you will be working with the cutting edge AI agents developed by our exceptional team of Machine Learning and Neuroscience research scientists. Your responsibilities will include everything from creating systems for agent testing using 2D and 3D games to developing test problems within physics simulators. You will create graphical visualization of results, build competitive agent leaderboards and test new algorithms on robots. To succeed in this role you will need to have a strong foundation in software engineering and enjoy working on a wide range of challenging problems within a mission-driven team.\nAs an Inference Performance Engineer, you will push the boundaries of AI model execution at scale. In this role, you will be at the forefront of making large-scale AI inference faster, cheaper, and more efficient. You will analyze the entire inference stack to identify critical bottlenecks and drive systemic improvements. By combining deep systems profiling, benchmarking, and first-principles problem solving, your work will directly maximize hardware throughput, reduce cost-to-serve, and empower our cross-functional teams to make data-driven capacity and latency tradeoffs.\nArtificial intelligence will be one of humanity’s most transformative inventions. At DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.\nWe are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.\nIndividual pay is determined by factors including job-related skills, experience, and relevant education or training.US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits\nLearn more about benefits at Google.\nResponsibilities\nAnalyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency.\nDesign and implement inference optimization techniques.\nInvestigate and resolve complex model inference performance bottlenecks across the stack.\nModel the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems.\nDevelop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet.\nQualifications\nMinimum qualifications:\nBachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience.\n8 years of experience in software development.\nExperience in Python and C++, including navigating, debugging, and modifying serving codebases.\nExperience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures.\nPreferred qualifications:\nExperience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo).\nExperience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools.\nExperience with observability and reliability for large distributed systems.\nFamiliarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.","description_format":"text","description_chars":4034,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Equity"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["LLM & Generative AI","Foundation Models","AI Research Labs","AI for Science"],"lifecycle":[{"event":"open","at":"2026-10-01T12:38:39Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":53,"reasons":["conf:1","win:early"],"computed_at":"2026-10-02T03:22:20Z"},"pay":null,"html_url":"https://alion.io/job/google-deepmind-staff-software-engineer-inference-performance-optimization-genai","json_url":"https://alion.io/job/google-deepmind-staff-software-engineer-inference-performance-optimization-genai.json","meta":{"generated_at":"2026-10-02T03:22:20Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4630,"day_limit":5000,"remaining_today":370,"minute_limit":60,"resets_at":"2026-10-03T00:00:00Z"}}}