NVIDIA is looking for Senior Customer Success/Partnership Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to redefine deep learning and data analytics, and to power data centers. We are building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent communication skills.
This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer!
Whatyou'llbedoing:
Lead the hands-on analysis, optimization, and performance tuning of complex GPU-accelerated systems and AI workloads, ensuring high availability and efficiency across customer data centers.
Serve as a senior technical authority on NVIDIA technologies, contributing to architecture reviews and guiding infrastructure decisions at scale.
Establish and refine monitoring and optimization methodologies using analytics, telemetry, and automation to detect bottlenecks and improve infrastructure resiliency.
Join post-deployment reviews, incident retrospectives, and sessions to craft the customer experience and provide insights into NVIDIA’s infrastructure strategy.
Completeandleadcomplextechnicalprojectsfrom initial designthroughimplementationandcontinuousimprovement,ensuringalignmenttoSLAs andmitigationoftechnicalrisks.
SupportbusinessgrowthbyidentifyingAIinfrastructureopportunitiesincloudandenterpriseenvironmentsanddrivingtechnicalinitiativesthatshowcaseNVIDIA’sleadershipinthisspace.
What we need to see:
10+ years of experience in large-scale data center service operations with a focus on infrastructure.
BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
Stronganalytical,solvingproblems, anddecision-makingskills,capableofidentifyingrootcauses,drivingcontinuousimprovement, anddeliveringresilienttechnicalsolutions.
Strongcommunication, timemanagement, and organizationalskills, withtheabilitytoleadcomplexprojects,guidetechnicalteams.
Preferredcertificationsindatacenter,server, ornetworkingtechnologies, and awillingnesstotravelupto25%forcustomerengagementsand teamcollaboration.
Proficiencyin system-levelaspects,encompassingOperating Systems, Linuxkerneldrivers, GPUs, NICs, andhardwarearchitecture.
Shownexpertiseincloudorchestrationsoftwareandjobschedulers,includingplatformslikeKubernetes, Docker Swarm, and HPC-specificschedulerssuchasSlurm.
Familiarity with cloud-native technologies and their integration with traditional infrastructure is essential.
Ways to stand out from the crowd:
Deep familiarity with AI infrastructure and workflows, including training/inference pipelines, MLOps/DevOps tools, containerization (Docker, Kubernetes), and large-scale system deployments.
Knowledge of data center infrastructure operations, including safety, security, environmental controls, and standard operating procedures.
Good interpersonal and collaboration skills, with the ability to lead discussions, influence outcomes, and build positive relationships with both internal and external collaborators.

