1,434,312open jobs
83,785companies
217,826added this week
Browse all
Salary
≈ $21k – $43k per year (Estimated)
Location
In office (Hyderabad)
Seniority
Senior · 3+ years exp
Employment
Full-Time

First seen by Alion on Oct 7, 2026.

Overview
Company
Impact
Profile match
A proven valued partner to Mobile Network Operators, Telecom Vendors, Social Networking Platforms, Global Email Providers and Resellers.
ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career.

Senior Network Engineer – GPU Cluster Networking

The Role

We are seeking a Senior Network Engineer with 8 to 15 yrs to join the AMD IT Network Engineering team.

This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.

The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.

The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.

You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.

The Person

You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.

You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.

You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.

Key Responsibilities

  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.
  • Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.
  • Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.
  • Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.
  • Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.
    • Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
  • Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.
  • Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.
  • Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.
  • Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations

Preferred Experience

  • Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.
  • Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.
  • Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation
  • Strong hands-on experience with RDMA and RoCEv2 in production environments.
  • Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.
  • Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.
  • Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.
  • Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.
  • Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.
  • Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.
  • Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.
  • Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.
  • Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.

Academic Creditals


  • Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience.

LOCATION:


HYDERABAD, TELANGANA

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's Responsible AI Policy is available here.

This posting is for an existing vacancy.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,434,312 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Hyderabad
$160k – $180k per year • Equity • Hybrid • Full-Time • 5+ years exp • United States
AI/ML
Claude
ChatGPT
Web3
TRM Labs
Apply
AI Engineer 10 months ago
$62k – $91k per year • Hybrid • 3+ years exp • Leusden
AI/ML
Copilot Studio
DevOps
Azure
Management
n8n
Apply
≈ $21k – $46k per year (Estimated) • In office • 15+ years exp • Bachelor's Degree • Chennai
Python
AI/ML
AI Agents
RAG
Machine Learning
DevOps
CI/CD
Apply
$150k – $300k per year • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
LLM
LLM Guardrails
DevOps
Terraform
CI/CD
Kubernetes
Grafana
Platform Engineering
IAM
Cybersecurity
Zero Trust
Apply
$86k – $114k per year • In office • Full-Time • The Hague
AI/ML
AI Agents
DevOps
Azure
AWS
Management
Agile
Apply
≈ $67k – $164k per year (Estimated) • In office • TS/SCI • Bethesda
Python
Databases
RabbitMQ
MinIO
ElasticSearch
Apache Kafka
AI/ML
Scikit-learn
TensorFlow
Machine Learning
DevOps
Ansible
OpenShift
Helm
CloudFormation
GitLab CI
Azure
CI/CD
Jenkins
AWS
Kubernetes
SaltStack
Configuration Management
Amazon S3
Linux
Analytics
Apache NiFi
Management
Agile
Apply
In office
DevOps
Kubernetes
Management
ITSM
Apply
In office
Python
PowerShell
DevOps
Terraform
Azure DevOps
Azure
CI/CD
Git
Kubernetes
Bicep
Apply
In office
Python
Bash
AI/ML
Airflow
MLFlow
Vertex AI
Kubeflow
LLM
RAG
Feast
Amazon SageMaker
LLMOps
Feature Store
Machine Learning
DevOps
Rest API
Terraform
GCP
Azure DevOps
GitHub Actions
CloudFormation
Prometheus
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
Argo Workflows
Apply
In office
DevOps
Red Hat
Linux
Management
ITIL
Apply
≈ $25k – $54k per year (Estimated) • In office • Full-Time • Bengaluru
Python
C++
AI/ML
CUDA Toolkit
OpenCL
CUDA
ROCm
DevOps
GitHub Actions
GitLab CI
CI/CD
Jenkins
Linux
Windows
Apply
≈ $21k – $43k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Stable Diffusion
TensorFlow
PyTorch
Machine Learning
Apply
≈ $20k – $41k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
Python
C++
Apply
In office • Full-Time • Hyderabad
C++
AI/ML
NCCL
InfiniBand
DevOps
CI/CD
Git
HPC
Linux
TCP/IP
Apply
≈ $18k – $39k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
Linux
Management
Agile
Apply
Principal AI Engineer 12 hours ago
≈ $22k – $49k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Hyderabad
Python
Python
FastAPI
Asyncio
Databases
PostgreSQL
Redis
AI/ML
LangGraph
AutoGen
LangChain
Claude
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
Langfuse
LangSmith
AWS Bedrock
CrewAI
LLM
OpenAI
Anthropic
LLMOps
A2A
GPT-4
Multi-Agent Systems
Tool Use
DevOps
GCP
WebSockets
Azure
CI/CD
Git
AWS
Docker
Cybersecurity
LDAP
Apply
≈ $36k – $94k per year (Estimated) • Hybrid • Full-Time • 8+ years exp • Master's Degree • Hyderabad • Pune
Python
Databases
Databricks
AI/ML
Fine-tuning
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
Reranking
LLMOps
Human-in-the-Loop
Structured Outputs
LLM Guardrails
Agentic Workflows
Tool Use
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Management
Waterfall
Apply
≈ $12k – $28k per year (Estimated) • Remote (India) • Full-Time • 3+ years exp • Hyderabad
Apply
≈ $10k – $20k per year (Estimated) • Remote (India) • Full-Time • 1+ year exp • Bachelor's Degree • Hyderabad
DevOps
SLI/SLO/SLA
Analytics
Microsoft Excel
Apply
≈ $11k – $22k per year (Estimated) • Remote (India) • Full-Time • Hyderabad
Management
Google Sheets
Apply
See all jobs
This is one of many
1,434,312 more open roles from verified company boards, updated every day.