406,377open jobs
14,133companies
78,486added this week
Browse all
Salary
$137k – $254k per year
Location
In office (San Jose)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Cadence Design Systems is an American software company founded in 1988 that supplies the electronic design automation tools engineers use to design and verify integrated circuits. Its products cover digital synthesis and place-and-route, custom and analogue design, circuit simulation with the Spectre and Virtuoso families, functional verification, and increasingly system-level analysis for packaging, thermal behaviour and computational fluid dynamics. Headquartered in San Jose and listed on Nasdaq, it also licenses design intellectual property such as processor and interface blocks, and competes primarily with Synopsys and Siemens EDA.

At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology.

We are seeking a highly skilled and experiencedAI Systems Engineer to join our team. This is a hands-on, senior individual contributor role that will be pivotal in leading the development, operations, and support of our entire AI infrastructure. You will be responsible for the entire lifecycle of our AI systems, from architecting and building high-performance GPU clusters to deploying and optimizing our most advanced AI models and agentic services.

Responsibilities

  • AI Infrastructure Architecture & Strategy: Lead the design and implementation of our next-generation AI infrastructure to support our Agentic AI initiatives. You will define the technical strategy for our on-premise GPU clusters, storage solutions, and networking to ensure optimal performance, scalability, and reliability for all our AI workloads.

  • Cloud AI Service Integration: Support and secure the use of public cloud AI services, includingAzure OpenAI services and Google Cloud Platform (GCP) services likeGemini. This includes managing secure access, monitoring usage, and tracking billing to ensure cost-effectiveness. You will also have hands-on experience supporting compute, GPUs, and AI services on both GCP and Azure.

  • Hands-on GPU Cluster Management: Take a leadership role in the configuration, installation, and optimization of GPU server clusters. This includes advanced troubleshooting of hardware and software, performance tuning, and implementing best practices for cluster utilization and resource management. You will be an expert in administering job schedulers likeLSF in a production environment, including integration withDocker for containerized job submission.

  • Full-Stack AI Tech Stack Development & Operations: Architect and deploy a robust and scalable AI tech stack. You will be responsible for the end-to-end operational lifecycle, including setting up and managing deep learning frameworks (PyTorch,TensorFlow), containerization withDocker andKubernetes, and implementing CI/CD pipelines for AI model development.

  • Advanced LLM Deployment & Optimization: Lead the deployment, serving, and optimization of Large Language Models (LLMs). You will be an expert in techniques such as model quantization, distillation, and using high-performance serving frameworks (e.g.,vLLM,TGI,TensorRT-LLM) to maximize inference throughput and minimize latency.

  • Agentic AI Workflow & Service Engineering: Architect and build production-grade Agentic AI workflows and services. You will be responsible for the technical design and implementation of systems that integrate LLMs with external tools, APIs, and databases, and will mentor other engineers on building robust and scalable AI agent applications.

  • Automation & Monitoring: Develop and maintain automation scripts using languages likePython,Bash, orPerl to streamline system maintenance, deployment, and reporting. Implement and manage monitoring solutions for system health, job statuses, GPU utilization, and container performance to proactively identify and resolve issues.

  • AI Systems Support & Mentorship: Act as the final escalation point for the most complex technical issues related to our AI infrastructure. You will also serve as a technical leader and mentor to other engineers, providing guidance on best practices in AI systems engineering, performance tuning, and operational excellence.

  • Security and Compliance: Develop and implement security best practices for our AI systems and data, ensuring compliance with relevant regulations and protecting our intellectual property.

Required Skills and Qualifications

  • 10+ years of experience in a senior technical role, with at least 5 years focused on building and operating high-performance computing or AI infrastructure. Proven track record as a Principal or Senior Staff Engineer.

  • Expert-level knowledge ofNVIDIA GPU architecture and technologies likeCUDA andcuDNN. Extensive experience with multi-GPU and multi-node training and inference.

  • Proven experience with public cloud AI services, specifically managing access, usage, and billing forAzure OpenAI and Google Cloud Platform (GCP) services.

  • Extensive hands-on experience withDocker: image management, container orchestration, and troubleshooting.

  • Proficiency in scripting languages such asPython,Bash, orPerl.

  • Deep expertise in Linux system administration (RHEL preferred), including networking, storage, and performance tuning.

  • Familiarity with user authentication and integration using systems likeLDAP or Active Directory.

  • Strong problem-solving and communication skills with the ability to work in a multi-platform, cross-functional, and geographically distributed team.

Preferred/Bonus Skills

  • Understanding of AI job profiling and tuning (memory, GPU, I/O).

  • Experience administeringLSF clusters in a production or research environment. Familiarity with other job schedulers likeSlurm is a plus.

  • Experience withLSF Docker integration and job submission using container images.

  • Experience with macOS/AppleSilicon system admin tasks and troubleshooting.

The annual salary range for California is $136,500 to $253,500. You may also be eligible to receive incentive compensation: bonus, equity, and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the salary range is a guideline and compensation may vary based on factors such as qualifications, skill level, competencies and work location. Our benefits programs include: paid vacation and paid holidays, 401(k) plan with employer match, employee stock purchase plan, a variety of medical, dental and vision plan options, and more.

We’re doing work that matters. Help us solve what others can’t.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
406,377 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
Remote/Hybrid • 3+ years exp
C#
SQL
Databases
MS SQL
DevOps
Azure
Azure DevOps
CI/CD
Git
GitLab
Apply
.NET Developer 11 hours ago
In office • 3+ years exp
C#
SQL
C#
.NET
Databases
MS SQL
DevOps
AWS
Azure
Azure DevOps
CI/CD
Git
GitLab
Apply
Senior Test Analyst 11 hours ago
$73k – $162k per year (Estimated) • Remote
SQL
C#
COBOL
C#
.NET
COBOL
IBM MQ
DevOps
Azure
Azure DevOps
Management
Confluence
Jira
QA
JMeter
SoapUI
Apply
Financial Analyst II 11 hours ago
In office • 5+ years exp
DevOps
AWS
Azure
FinOps
GCP
Analytics
Power BI
Apply
In office • 10+ years exp • Bachelor's Degree
DevOps
AWS
Azure
SLI/SLO/SLA
Management
Jira
ServiceNow
Apply
In office • Full-Time • 10+ years exp • Seoul
Apply
In office • Full-Time • Hsinchu
Chips/EDA
Cadence Pegasus
Apply
In office • Full-Time • 17+ years exp • Hyderabad
DevOps
HPC
Apply
Equity • In office • Internship • Dublin
Apply
$161k – $299k per year • Equity • In office • Full-Time • 10+ years exp • Master's Degree • San Jose
AI/ML
AI Agents
Apply
$68k – $159k per year (Estimated) • In office • San Jose
Apply
$56k – $94k per year • In office • Contractor • San Jose
Apply
$64k – $74k per year • In office • Contractor • San Jose
Apply
$56k per year • In office • Contractor • San Jose
Apply
$84k – $94k per year • In office • Contractor • San Jose
Apply
See all jobs
This is one of many
406,377 more open roles from verified company boards, updated every day.