{"id":2038825,"url":"https://alion.io/job/amdtelecom-senior-network-engineer-gpu-cluster-networking","title":"Senior Network Engineer GPU Cluster Networking","company":{"id":2660826,"name":"Amdtelecom","domain":"amdtelecom.net","url":"https://alion.io/company/amdtelecom-2","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Schema","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Hyderabad, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":22000,"max_usd":45000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":27},"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"BGP","optional":false},{"name":"HPC","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"ROCm","optional":false},{"name":"SLURM","optional":false},{"name":"VLAN","optional":false},{"name":"Grafana","optional":true},{"name":"Prometheus","optional":true}],"status":"live","first_seen_at":"2026-10-07T12:30:46Z","employer_posted_date":null,"last_verified_at":"2026-10-07T12:30:46Z","board_verified":false,"closed_at":null,"days_open":4,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":4},"description":"ADVANCE YOUR CAREER. ADVANCE THE WORLD. \n\nAt AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.\n\nWhether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career.\n\nSenior Network Engineer – GPU Cluster Networking\n\nThe Role\n\nWe are seeking a Senior Network Engineer with 8 to 15 yrs to join the AMD IT Network Engineering team.\n\nThis role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.\n\nThe ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.\n\nThe primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.\n\nYou will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.\n\nThe Person\n\nYou are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.\n\nYou understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.\n\nYou are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.\n\nKey Responsibilities\n\nArchitect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.\nDesign network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.\nOwn the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.\nDesign and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.\nDevelop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.\nPerform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.\nConfigure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings\nDesign and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.\nOptimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.\nLead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.\nPlan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations\n\nPreferred Experience\n\nSignificant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.\nExperience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.\nDeep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation\nStrong hands-on experience with RDMA and RoCEv2 in production environments.\nDemonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.\nStrong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.\nExperience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.\nStrong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.\nExperience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.\nExperience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting\nExperience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.\nExperience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.\nStrong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting \nExperience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.\nExperience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.\n\nAcademic Creditals\n\nBachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience.\n\nLOCATION:\n\nHYDERABAD, TELANGANA\n\nBenefits offered are described: AMD benefits at a glance.\n\nAMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants needs under the respective laws throughout all stages of the recruitment and selection process.\n\nAMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's Responsible AI Policy is available here.\n\nThis posting is for an existing vacancy.","description_format":"text","description_chars":7655,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Cybersecurity","Email Hosting & Delivery"],"lifecycle":[{"event":"open","at":"2026-10-07T17:09:15Z"}],"visa":[],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":2,"expected_fill_days":45,"reasons":["seen:2","velocity","win:early"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/amdtelecom-senior-network-engineer-gpu-cluster-networking","json_url":"https://alion.io/job/amdtelecom-senior-network-engineer-gpu-cluster-networking.json","meta":{"generated_at":"2026-10-11T20:53:53Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler_verified","counted_by":"address","units_charged":1,"used_today":9356,"day_limit":null,"remaining_today":null,"minute_limit":300,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2660826},"rest":"https://alion.io/mcp/rest/get_company?id=2660826"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Famdtelecom-senior-network-engineer-gpu-cluster-networking"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Famdtelecom-senior-network-engineer-gpu-cluster-networking"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Famdtelecom-senior-network-engineer-gpu-cluster-networking"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/amdtelecom-senior-network-engineer-gpu-cluster-networking\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Famdtelecom-senior-network-engineer-gpu-cluster-networking"}]}