1,469,448open jobs
87,948companies
231,463added this week
Browse all
Location
In office
Seniority
Architect · 12+ years exp

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 9, 2026.

Overview
Company
Impact
Profile match
HCLTech is a major Indian multinational information technology (IT) services and consulting company headquartered in Noida, Uttar Pradesh. Spun off from the original HCL Group in 1991, it ranks as one of India's largest technology companies alongside firms like TCS, Infosys, and Wipro.

Job Summary

KLA is seeking a highly experienced and hands-on Principal AI Platform Architect to lead the design and implementation of a next-generation AI infrastructure platform on Google Cloud Platform (GCP). The ideal candidate will possess deep expertise in architecting and deploying Google Kubernetes Engine (GKE) clusters supporting both TPU and GPU workloads for large-scale AI and Machine Learning applications. This role requires a technical leader who can directly engage with engineering teams, drive architecture discussions, define platform standards, and establish a scalable, secure, and observable AI platform capable of supporting model development, training, experimentation, and production operations. The successful candidate will have proven experience designing and implementing TPU-enabled clusters for large-scale model training while optimizing resource utilization, performance, cost, and operational efficiency

Job Description

Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or related field.

  • 12+ years of IT infrastructure, cloud architecture, or platform engineering experience.
  • 5+ years of hands-on experience with Google Cloud Platform (GCP).
  • Strong experience designing and implementing production-scale GKE environments.
  • Proven experience deploying TPU-enabled infrastructure for AI model training.
  • Strong expertise in GPU-based workload orchestration and optimization.
  • Experience supporting large-scale AI/ML platforms in production environments.
  • Deep understanding of Kubernetes architecture and operations.
  • Experience with distributed training and high-performance computing environments.
  • Strong communication and stakeholder management skills

Key Responsibilities

AI Platform Architecture & Design

  • Design and implement enterprise-scale AI platforms on Google Cloud Platform.
  • Architect highly available and scalable GKE environments supporting: o TPU workloads for AI model training o GPU workloads for inference and accelerated computing o Multi-tenant engineering teams
  • Define architecture patterns for AI/ML platform lifecycle management.
  • Establish best practices for workload scheduling, resource isolation, cluster autoscaling, and infrastructure optimization.
  • Drive platform design reviews and technical governance. GKE, TPU & GPU Infrastructure
  • Design and deploy production-grade GKE clusters optimized for AI workloads.
  • Implement TPU-enabled GKE architectures for large-scale model training.
  • Configure GPU node pools supporting frameworks such as TensorFlow, PyTorch, JAX, and distributed training environments.
  • Design workload placement strategies across TPU and GPU resources.
  • Optimize cluster performance, utilization, availability, and cost efficiency.
  • Develop infrastructure blueprints and reference architectures for AI platforms. AI Platform Operations & Engineering
  • Define operational models for AI infrastructure management.
  • Implement cluster lifecycle management, upgrades, patching, and maintenance strategies.
  • Design automated provisioning and deployment processes using Infrastructure as Code (IaC).
  • Establish platform automation using CI/CD and GitOps methodologies.
  • Develop operational standards for capacity planning and infrastructure scaling. Resource Allocation & Optimization
  • Design mechanisms for: o Resource allocation o Quota management o Workload scheduling o Chargeback/showback models
  • Implement intelligent resource utilization and optimization strategies.
  • Establish cost governance and monitoring frameworks for TPU/GPU consumption.
  • Improve utilization efficiency across shared AI compute environments. Observability & Monitoring
  • Define and implement platform observability frameworks.
  • Establish: o Monitoring o Logging o Tracing o Performance analytics
  • Build operational dashboards for TPU, GPU, and cluster health monitoring.
  • Create alerting and incident management processes for critical AI workloads. Security & Compliance
  • Design security controls aligned to enterprise cloud security standards.
  • Implement: o Identity and Access Management (IAM) o Workload Identity o Network segmentation o Secret management o Encryption controls
  • Establish secure multi-tenant AI environments.
  • Support compliance and governance requirements for AI workloads. Stakeholder Engagement
  • Collaborate directly with KLA engineering leadership and AI teams.
  • Lead technical workshops, architecture reviews, and roadmap discussions.
  • Translate business and engineering requirements into scalable platform solutions.
  • Mentor engineering teams on GCP AI infrastructure best practices.
  • Act as the primary SME for AI platform architecture and implementation

Skill Requirements

Google Cloud Platform

  • Google Kubernetes Engine (GKE)
  • Cloud TPU
  • Compute Engine GPUs
  • Cloud Storage
  • VPC Networking
  • IAM
  • Cloud Monitoring
  • Cloud Logging
  • Cloud Operations Suite Kubernetes & Containerization
  • Kubernetes Administration
  • Cluster Autoscaling
  • Node Pool Management
  • Resource Quotas
  • Workload Scheduling
  • Multi-Cluster Architectures
  • Service Mesh Technologies AI/ML Infrastructure
  • TensorFlow
  • PyTorch
  • JAX
  • Distributed Training Frameworks
  • MLOps Platforms
  • Model Training Pipelines
  • AI Infrastructure Optimization DevOps & Automation
  • Terraform
  • Infrastructure as Code (IaC)
  • GitOps
  • CI/CD Pipelines
  • ArgoCD
  • Jenkins
  • GitHub Actions Observability & Security
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Cloud Monitoring
  • Cloud Logging
  • Security Hardening
  • Identity Management
  • Kubernetes Security

Other Requirements

Google Cloud Professional Cloud Architect Certification.

  • Google Cloud Professional Machine Learning Engineer Certification.
  • Experience with Vertex AI and enterprise MLOps platforms.
  • Experience with large language model (LLM) training environments.
  • Experience supporting multi-petabyte AI datasets.
  • Familiarity with NVIDIA AI ecosystem and CUDA-based workloads.
  • Experience implementing FinOps strategies for AI infrastructure.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,469,448 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Solutions
Similar stack
Same company
In your city
≈ $105k – $215k per year (Estimated) • Remote (United States) • Full-Time
Apply
Remote (United States) • Part-Time
DevOps
Azure
Windows Server
Cybersecurity
Active Directory
Management
OneDrive
SharePoint
Apply
≈ $112k – $229k per year (Estimated) • Remote (United States) • Full-Time
SQL
Databases
Snowflake
Databricks
Google BigQuery
BigQuery
AI/ML
AI Agents
Chips/EDA
PoC Library
Apply
Sales Engineer - DACH 2 hours ago
Remote (United States) • Full-Time
Python
Cybersecurity
GreyNoise
SIEM
Apply
≈ $136k – $257k per year (Estimated) • Remote (United States) • Boston
AI/ML
AI Agents
Human-in-the-Loop
Knowledge Graph
LLM Guardrails
DevOps
GCP
Azure
AWS
Platform Engineering
Apply
In office • Contractor • 4+ years exp • Mexico City
Java
SQL
DevOps
Terraform
GCP
VMWare
CI/CD
GitOps
ArgoCD
Docker
Kubernetes
Platform Engineering
Google GKE
Linux
Apply
$95k – $179k per year • In office • Full-Time • 2+ years exp • Berlin
Python
JavaScript
TypeScript
Node JS
Python
FastAPI
Node JS
Fastify
Databases
PostgreSQL
Redis
AI/ML
AI Agents
LiveKit
Frontend
React.js
DevOps
GCP
OpenTelemetry
Azure
AWS
Honeycomb
Apply
≈ $16k – $37k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • India
Python
JavaScript
TypeScript
Ruby
Node JS
Python
Django
Databases
MySQL
PostgreSQL
Frontend
Vue.js
Angular
React.js
DevOps
GCP
Azure
CI/CD
Git
AWS
Apply
≈ $123k – $226k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Sunnyvale
Python
JavaScript
Java
Node JS
Databases
Apache Kafka
Google BigQuery
BigQuery
AI/ML
AI Agents
Kubeflow
Flink
TensorFlow
Frontend
GraphQL
React.js
DevOps
GCP
AWS
Analytics
A/B Testing
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • India
Python
TypeScript
SQL
AI/ML
LangGraph
AutoGen
LangChain
Embeddings
Prompt Engineering
AI Agents
AgentOps
CrewAI
RAG
AWS Bedrock AgentCore
Machine Learning
DevOps
OpenTelemetry
AWS
Grafana
AWS Lambda
Amazon S3
Amazon CloudWatch
API Gateway
Management
Agile
Apply
In office • 4+ years exp • Bachelor's Degree
JavaScript
ABAP
ABAP
SAP Fiori
CDS Views
SAP Gateway
SAP BTP
Apply
Hybrid • 1+ year exp
ABAP
ABAP
CDS Views
Apply
In office • 5+ years exp • Bachelor's Degree
SQL
ABAP
Databases
SAP HANA
SAP BW
Analytics
ETL/ELT
Apply
In office • 5+ years exp • Bachelor's Degree
JavaScript
SQL
C#
ABAP
ABAP
SAP Fiori
SAP BTP
DevOps
Rest API
Git
SOAP
Management
Jira
ServiceNow
Apply
In office • 3+ years exp
ABAP
ABAP
SAP BTP
Apply
See all jobs
This is one of many
1,469,448 more open roles from verified company boards, updated every day.