368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$137k – $254k per year
Location
In office (San Jose)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Semiconductor Engineering is a technical media and research organization based in San Jose, California, and founded in 2013. The company provides in-depth news, analysis, and technical papers focused on the design, manufacturing, testing, and integration of advanced semiconductor technologies. It operates as a global information hub for chip engineers and industry professionals, offering newsletters, webinars, and specialized research reports.

At Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology.

We are seeking a highly skilled and experiencedAI Systems Engineer to join our team. This is a hands-on, senior individual contributor role that will be pivotal in leading the development, operations, and support of our entire AI infrastructure. You will be responsible for the entire lifecycle of our AI systems, from architecting and building high-performance GPU clusters to deploying and optimizing our most advanced AI models and agentic services.

Responsibilities

  • AI Infrastructure Architecture & Strategy: Lead the design and implementation of our next-generation AI infrastructure to support our Agentic AI initiatives. You will define the technical strategy for our on-premise GPU clusters, storage solutions, and networking to ensure optimal performance, scalability, and reliability for all our AI workloads.

  • Cloud AI Service Integration: Support and secure the use of public cloud AI services, includingAzure OpenAI services and Google Cloud Platform (GCP) services likeGemini. This includes managing secure access, monitoring usage, and tracking billing to ensure cost-effectiveness. You will also have hands-on experience supporting compute, GPUs, and AI services on both GCP and Azure.

  • Hands-on GPU Cluster Management: Take a leadership role in the configuration, installation, and optimization of GPU server clusters. This includes advanced troubleshooting of hardware and software, performance tuning, and implementing best practices for cluster utilization and resource management. You will be an expert in administering job schedulers likeLSF in a production environment, including integration withDocker for containerized job submission.

  • Full-Stack AI Tech Stack Development & Operations: Architect and deploy a robust and scalable AI tech stack. You will be responsible for the end-to-end operational lifecycle, including setting up and managing deep learning frameworks (PyTorch,TensorFlow), containerization withDocker andKubernetes, and implementing CI/CD pipelines for AI model development.

  • Advanced LLM Deployment & Optimization: Lead the deployment, serving, and optimization of Large Language Models (LLMs). You will be an expert in techniques such as model quantization, distillation, and using high-performance serving frameworks (e.g.,vLLM,TGI,TensorRT-LLM) to maximize inference throughput and minimize latency.

  • Agentic AI Workflow & Service Engineering: Architect and build production-grade Agentic AI workflows and services. You will be responsible for the technical design and implementation of systems that integrate LLMs with external tools, APIs, and databases, and will mentor other engineers on building robust and scalable AI agent applications.

  • Automation & Monitoring: Develop and maintain automation scripts using languages likePython,Bash, orPerl to streamline system maintenance, deployment, and reporting. Implement and manage monitoring solutions for system health, job statuses, GPU utilization, and container performance to proactively identify and resolve issues.

  • AI Systems Support & Mentorship: Act as the final escalation point for the most complex technical issues related to our AI infrastructure. You will also serve as a technical leader and mentor to other engineers, providing guidance on best practices in AI systems engineering, performance tuning, and operational excellence.

  • Security and Compliance: Develop and implement security best practices for our AI systems and data, ensuring compliance with relevant regulations and protecting our intellectual property.

Required Skills and Qualifications

  • 10+ years of experience in a senior technical role, with at least 5 years focused on building and operating high-performance computing or AI infrastructure. Proven track record as a Principal or Senior Staff Engineer.

  • Expert-level knowledge ofNVIDIA GPU architecture and technologies likeCUDA andcuDNN. Extensive experience with multi-GPU and multi-node training and inference.

  • Proven experience with public cloud AI services, specifically managing access, usage, and billing forAzure OpenAI and Google Cloud Platform (GCP) services.

  • Extensive hands-on experience withDocker: image management, container orchestration, and troubleshooting.

  • Proficiency in scripting languages such asPython,Bash, orPerl.

  • Deep expertise in Linux system administration (RHEL preferred), including networking, storage, and performance tuning.

  • Familiarity with user authentication and integration using systems likeLDAP or Active Directory.

  • Strong problem-solving and communication skills with the ability to work in a multi-platform, cross-functional, and geographically distributed team.

Preferred/Bonus Skills

  • Understanding of AI job profiling and tuning (memory, GPU, I/O).

  • Experience administeringLSF clusters in a production or research environment. Familiarity with other job schedulers likeSlurm is a plus.

  • Experience withLSF Docker integration and job submission using container images.

  • Experience with macOS/AppleSilicon system admin tasks and troubleshooting.

The annual salary range for California is $136,500 to $253,500. You may also be eligible to receive incentive compensation: bonus, equity, and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the salary range is a guideline and compensation may vary based on factors such as qualifications, skill level, competencies and work location. Our benefits programs include: paid vacation and paid holidays, 401(k) plan with employer match, employee stock purchase plan, a variety of medical, dental and vision plan options, and more.

We’re doing work that matters. Help us solve what others can’t.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$76k – $165k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Canada
PowerShell
SQL
Databases
MS SQL
MySQL
Oracle
PostgreSQL
DevOps
AWS
Azure
CI/CD
Analytics
ETL/ELT
Apply
$25k – $60k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Chennai
Java
PowerShell
Python
TypeScript
JavaScript
Java
Spring Boot
Databases
Amazon Redshift
DynamoDB
Frontend
Angular
DevOps
Amazon EC2
Amazon EKS
Ansible
ArgoCD
AWS
AWS Lambda
Azure
Azure DevOps
Bicep
Chef
CI/CD
CloudFormation
Configuration Management
Docker
GCP
GitLab CI
Jenkins
Kubernetes
Platform Engineering
Puppet
Terraform
Amazon S3
GitLab
IAM
Apply
$45k – $114k per year (Estimated) • Remote • Full-Time • São Paulo • Vitoria-Gasteiz • Fortaleza • Rio de Janeiro • Recife
DevOps
AWS
Azure
Azure DevOps
Bicep
CI/CD
CloudFormation
GCP
GitHub
GitHub Actions
GitLab
GitLab CI
IAM
Incident Management
Jenkins
Kubernetes
Terraform
Apply
$115k – $204k per year (Estimated) • In office • Full-Time • United States
Bash
C++
Python
C++
CMake
DevOps
GitHub
GitHub Actions
Apply
Sr. UX Designer 1 hour ago
$76k – $161k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Toronto
Python
AI/ML
AI Agents
Hallucination
Human-in-the-Loop
LLM
LLM Guardrails
Design
Figma
Apply
$27k – $71k per year (Estimated) • In office • Full-Time • 7+ years exp • Pune
C#
C++
Java
Python
DevOps
Azure
CI/CD
Git
QA
Pytest
Apply
In office • Full-Time • Seoul
C++
SystemVerilog
Verilog
Apply
$28k – $74k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Noida
C++
Python
Chips/EDA
Cadence Pegasus
Apply
In office • Full-Time • Master's Degree • Beijing
C++
Chips/EDA
Cadence AMS Designer
Cadence Spectre
Cadence Xcelium
Apply
In office • Full-Time • Master's Degree • Hsinchu
C++
Apply
$147k – $265k per year (Estimated) • In office • Full-Time • Folsom • San Jose
Apply
$117k – $255k per year (Estimated) • In office • Full-Time • 8+ years exp • Richardson • Boise • Folsom • San Jose
Verilog
AI/ML
Claude
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$116k – $253k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise • San Jose
Apply
$68k – $85k per year • In office • Full-Time • Master's Degree • San Jose
Python
AI/ML
AI Agents
DevOps
Amazon EC2
AWS
AWS Lambda
Bitbucket
CI/CD
CloudFormation
Docker
Git
Kubernetes
Terraform
Amazon S3
IAM
HPC
Cybersecurity
Least Privilege
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.