Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 9, 2026.
Job Summary
KLA is seeking a highly experienced and hands-on Principal AI Platform Architect to lead the design and implementation of a next-generation AI infrastructure platform on Google Cloud Platform (GCP). The ideal candidate will possess deep expertise in architecting and deploying Google Kubernetes Engine (GKE) clusters supporting both TPU and GPU workloads for large-scale AI and Machine Learning applications. This role requires a technical leader who can directly engage with engineering teams, drive architecture discussions, define platform standards, and establish a scalable, secure, and observable AI platform capable of supporting model development, training, experimentation, and production operations. The successful candidate will have proven experience designing and implementing TPU-enabled clusters for large-scale model training while optimizing resource utilization, performance, cost, and operational efficiency
Job Description
Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or related field.
- 12+ years of IT infrastructure, cloud architecture, or platform engineering experience.
- 5+ years of hands-on experience with Google Cloud Platform (GCP).
- Strong experience designing and implementing production-scale GKE environments.
- Proven experience deploying TPU-enabled infrastructure for AI model training.
- Strong expertise in GPU-based workload orchestration and optimization.
- Experience supporting large-scale AI/ML platforms in production environments.
- Deep understanding of Kubernetes architecture and operations.
- Experience with distributed training and high-performance computing environments.
- Strong communication and stakeholder management skills
Key Responsibilities
AI Platform Architecture & Design
- Design and implement enterprise-scale AI platforms on Google Cloud Platform.
- Architect highly available and scalable GKE environments supporting: o TPU workloads for AI model training o GPU workloads for inference and accelerated computing o Multi-tenant engineering teams
- Define architecture patterns for AI/ML platform lifecycle management.
- Establish best practices for workload scheduling, resource isolation, cluster autoscaling, and infrastructure optimization.
- Drive platform design reviews and technical governance. GKE, TPU & GPU Infrastructure
- Design and deploy production-grade GKE clusters optimized for AI workloads.
- Implement TPU-enabled GKE architectures for large-scale model training.
- Configure GPU node pools supporting frameworks such as TensorFlow, PyTorch, JAX, and distributed training environments.
- Design workload placement strategies across TPU and GPU resources.
- Optimize cluster performance, utilization, availability, and cost efficiency.
- Develop infrastructure blueprints and reference architectures for AI platforms. AI Platform Operations & Engineering
- Define operational models for AI infrastructure management.
- Implement cluster lifecycle management, upgrades, patching, and maintenance strategies.
- Design automated provisioning and deployment processes using Infrastructure as Code (IaC).
- Establish platform automation using CI/CD and GitOps methodologies.
- Develop operational standards for capacity planning and infrastructure scaling. Resource Allocation & Optimization
- Design mechanisms for: o Resource allocation o Quota management o Workload scheduling o Chargeback/showback models
- Implement intelligent resource utilization and optimization strategies.
- Establish cost governance and monitoring frameworks for TPU/GPU consumption.
- Improve utilization efficiency across shared AI compute environments. Observability & Monitoring
- Define and implement platform observability frameworks.
- Establish: o Monitoring o Logging o Tracing o Performance analytics
- Build operational dashboards for TPU, GPU, and cluster health monitoring.
- Create alerting and incident management processes for critical AI workloads. Security & Compliance
- Design security controls aligned to enterprise cloud security standards.
- Implement: o Identity and Access Management (IAM) o Workload Identity o Network segmentation o Secret management o Encryption controls
- Establish secure multi-tenant AI environments.
- Support compliance and governance requirements for AI workloads. Stakeholder Engagement
- Collaborate directly with KLA engineering leadership and AI teams.
- Lead technical workshops, architecture reviews, and roadmap discussions.
- Translate business and engineering requirements into scalable platform solutions.
- Mentor engineering teams on GCP AI infrastructure best practices.
- Act as the primary SME for AI platform architecture and implementation
Skill Requirements
Google Cloud Platform
- Google Kubernetes Engine (GKE)
- Cloud TPU
- Compute Engine GPUs
- Cloud Storage
- VPC Networking
- IAM
- Cloud Monitoring
- Cloud Logging
- Cloud Operations Suite Kubernetes & Containerization
- Kubernetes Administration
- Cluster Autoscaling
- Node Pool Management
- Resource Quotas
- Workload Scheduling
- Multi-Cluster Architectures
- Service Mesh Technologies AI/ML Infrastructure
- TensorFlow
- PyTorch
- JAX
- Distributed Training Frameworks
- MLOps Platforms
- Model Training Pipelines
- AI Infrastructure Optimization DevOps & Automation
- Terraform
- Infrastructure as Code (IaC)
- GitOps
- CI/CD Pipelines
- ArgoCD
- Jenkins
- GitHub Actions Observability & Security
- Prometheus
- Grafana
- OpenTelemetry
- Cloud Monitoring
- Cloud Logging
- Security Hardening
- Identity Management
- Kubernetes Security
Other Requirements
Google Cloud Professional Cloud Architect Certification.
- Google Cloud Professional Machine Learning Engineer Certification.
- Experience with Vertex AI and enterprise MLOps platforms.
- Experience with large language model (LLM) training environments.
- Experience supporting multi-petabyte AI datasets.
- Familiarity with NVIDIA AI ecosystem and CUDA-based workloads.
- Experience implementing FinOps strategies for AI infrastructure.

