We are looking for an experienced Software Engineer specializing in Kubernetes and GPU infrastructure to design and operate the foundational systems that power eBay's AI platform. You will own critical layers of our infrastructure, from Kubernetes CRD-based automation and a custom AI-aware GPU scheduler to RDMA-optimized multi-NIC GPU clusters and large-scale training environments, enabling our ML teams to train and serve AI models at eBay scale. You will work on Kubernetes operator development, Gateway API and hybrid cloud networking, multi-NIC RDMA fabric design, a global GPU scheduler with cross-availability-zone dispatch, GPU pool management with provisioned throughput integration, topology-aware workload placement, and KubeRay infrastructure, partnering closely with ML Platform, AI Research, and Networking teams.
Responsibilities:
- Design and build Kubernetes Custom Resource Definitions (CRDs) and operators for ML workloads, GPU node pools, and RayService CRDs.
- Architect and operate Kubernetes networking layers, including Gateway API, Service Load Balancers, and API gateways for hybrid cloud connectivity.
- Design, deploy, and operate multi-NIC Kubernetes clusters with RDMA-enabled networking.
- Implement and optimize RDMA networking using GPUDirect RDMA, RoCE, and InfiniBand for distributed GPU workloads.
- Build and operate a custom AI-aware GPU scheduler with topology-aware placement, preemption, and GPU defragmentation.
- Design and manage GPU pool management systems across on-premise and cloud-burst GPU environments.
- Deploy and operate KubeRay infrastructure for distributed Ray clusters supporting training and inference workloads.
- Implement cloud bursting and spot or preemptible GPU scheduling to improve utilization.
- Automate GPU asset provisioning, node configuration, and cluster lifecycle management using Infrastructure-as-Code and GitOps.
- Implement networking policies, multi-tenant isolation, RBAC, and security controls across the Kubernetes environment.
- Build observability for GPU utilization, NCCL communication, scheduler decisions, and network throughput.
- Collaborate with the ML Platform, AI Research, and Networking teams to optimize infrastructure for training and online inference.

