609,056open jobs
32,790companies
86,625added this week
Browse all
Salary
$140k – $185k per year
Location
In office (Springfield)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environmen. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Here, your work is more than a job- it's a journey in innovation. With opportunities to work on high-impact projects, access to the latest technologies, and a culture that thrives on creativity and collaboration, INflow Federal is where your expertise can truly make a difference.

Specific Duties and Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.

Required Skills:

  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

Preferred Skills:

    • Familiarity with parallel file systems and high-throughput storage architectures.
    • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
609,056 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Springfield
In office • 3+ years exp • PhD
Python
SQL
DevOps
GCP
Docker
Apply
$98k – $244k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tunis
Python
Python
SQLAlchemy
FastAPI
Pydantic
Databases
PostgreSQL
AI/ML
Model Context Protocol
Prompt Engineering
AI Agents
LLM
DevOps
Kong
Azure
CI/CD
Docker
Kubernetes
API Gateway
Cybersecurity
Microsoft Entra ID
Apply
$56k – $164k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Tunis
Python
AI/ML
Model Context Protocol
AI Agents
DevOps
Helm
Azure DevOps
GitHub Actions
Azure
CI/CD
GitOps
ArgoCD
AWS
Docker
Kubernetes
Amazon EKS
Azure AKS
Apply
$80k – $214k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Tunis
Python
AI/ML
Model Context Protocol
AI Agents
DevOps
Helm
Azure DevOps
GitHub Actions
Azure
CI/CD
GitOps
ArgoCD
AWS
Docker
Kubernetes
Amazon EKS
Azure AKS
Apply
$68k – $165k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Tunis
Python
AI/ML
Model Context Protocol
Dagster
AI Agents
RAG
Semantic Search
Semantic Search
DevOps
Rest API
Azure
Docker
Vector
Analytics
ETL/ELT
Apply
Cloud Engineer 19 days ago
$130k – $165k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Springfield
Python
Go
JavaScript
TypeScript
Ruby
PowerShell
Perl
Databases
PostgreSQL
Frontend
Angular
DevOps
Rest API
CI/CD
Jenkins
AWS
Docker
Kubernetes
Amazon EC2
GitLab
Amazon S3
Management
Agile
Scrum
Apply
$110k – $155k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Springfield
Python
PowerShell
DevOps
Azure
Cybersecurity
Tanium
Management
Agile
Apply
Sr Network Engineer 19 days ago
$135k – $155k per year • In office • TS/SCI • Full-Time • 10+ years exp • Bachelor's Degree • Springfield
DevOps
Incident Management
Apply
$149k – $187k per year • In office • TS/SCI • 5+ years exp • Bachelor's Degree • Springfield
Python
Java
Scala
AI/ML
LLM
RAG
DevOps
OpenShift
Helm
CI/CD
Git
AWS
Docker
Kubernetes
Management
Agile
Scrum
Apply
Sr Test Engineer (T&E) 2 months ago
$140k – $220k per year • In office • TS/SCI • 8+ years exp • Bachelor's Degree • San Diego
JavaScript
SQL
DevOps
Jenkins
AWS
Docker
GitHub
GitLab
Management
Agile
QA
Postman
Apply
$50k – $60k per year • In office • Part-Time • High School Diploma • Springfield
Apply
In office • Full-Time • Springfield
Apply
In office • Full-Time • Springfield
Apply
$141k – $317k per year • In office • Springfield
JavaScript
TypeScript
AI/ML
Human-in-the-Loop
Frontend
React.js
DevOps
CI/CD
Design
Figma
Management
Agile
Apply
$144k – $324k per year • In office • Springfield
DevOps
Terraform
Azure
CI/CD
AWS
Kubernetes
FinOps
Cybersecurity
Zero Trust
Apply
See all jobs
This is one of many
609,056 more open roles from verified company boards, updated every day.