Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Aug 27, 2026. Cognizant scores B on the Alion truth index.
Role Overview
We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms. The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.
Key Responsibilities
· HPC Infrastructure Management
o Operate and maintain HPC clusters based on CentOS, RHEL, and hardware platforms like HPE and NVIDIA DGX.
o Ensure optimal performance, scalability, and reliability of compute resources.
· Storage Administration
o Manage large-scale storage systems including Dell Isilon, VAST Storage, Lustre, and GPFS.
o Implement data lifecycle management and optimize storage performance for HPC workloads.
· Networking
o Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
o Troubleshoot network performance issues and ensure secure connectivity.
· Cluster and Job Scheduling
o Administer cluster management tools such as Bright Cluster Manager, Altair Grid Manager, and IBM LSF.
o Optimize job scheduling and resource allocation for diverse workloads.
· Monitoring and Automation
o Implement monitoring solutions using Zabbix, Grafana, and ELK Stack.
o Automate provisioning and configuration using Cobbler, Chef, Ansible, and AWS ParallelCluster.
· Performance Tuning & Troubleshooting
o Conduct performance benchmarking and tuning for HPC workloads.
o Diagnose and resolve hardware/software issues across compute, storage, and network layers.
· Security & Compliance
o Ensure HPC environment adheres to security best practices and compliance standards.
Required Skills & Qualifications
· Technical Expertise
o Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
o Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
o Proficiency in InfiniBand networking and high-speed interconnects.
o Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
· Automation & Scripting
o Expertise in Ansible, Chef, Cobbler, and scripting languages (Bash, Python).
o Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
· Monitoring & Logging
o Practical experience with Zabbix, Grafana, and ELK Stack for system health and performance monitoring.
· Soft Skills
o Strong problem-solving and analytical skills.
o Ability to work in a fast-paced environment and lead technical teams.
o Excellent communication and documentation skills.
Preferred Qualifications
· Exposure to AI/ML workloads on HPC clusters.
· Experience with containerization (Docker, Singularity) in HPC environments.
· Knowledge of security hardening for HPC systems.
Education
· Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.

