This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Infrastructure Engineer based in Canada.
As a Senior Infrastructure Engineer, you will take ownership of business-critical infrastructure supporting demanding AI and GPU workloads.
You will design, deploy, and operate large-scale OpenStack and Kubernetes environments with a strong focus on performance, reliability, and scalability.
The role combines hands-on systems engineering with automation, infrastructure-as-code, observability, security, and incident response.
You will work directly with physical servers, racks, networking, storage, and GPU infrastructure across data-centre environments.
Your work will have a visible impact on platform stability, customer experience, and the ability to scale globally.
You will collaborate closely with platform, DevOps, AI, product, and support teams in a fast-moving, technically ambitious environment.
This is an ownership-driven opportunity for an engineer who enjoys solving complex infrastructure problems and improving systems end-to-end.
Accountabilities:
- Design, deploy, operate, and continuously improve OpenStack and Kubernetes environments supporting high-performance GPU workloads.
- Take end-to-end ownership of infrastructure performance, scalability, resilience, and operational reliability.
- Build and maintain infrastructure using infrastructure-as-code, GitOps, CI/CD, and Git-based operational practices.
- Automate provisioning, deployment, configuration, and recurring infrastructure workflows to improve efficiency and consistency.
- Optimize GPU workload scheduling using Kubernetes and NVIDIA technologies to maximize platform performance and resource utilization.
- Implement and improve monitoring, logging, alerting, and observability systems to identify and resolve infrastructure issues proactively.
- Lead incident response activities, troubleshoot complex production problems, and drive post-incident improvements that strengthen overall reliability.
- Maintain robust security controls across infrastructure and container environments, including RBAC, network policies, access controls, and tenant isolation.
- Build, rack, cable, configure, and commission physical server and GPU infrastructure within data-centre environments.
- Work with networking and storage systems to ensure reliable, scalable infrastructure for compute-intensive workloads.
- Collaborate with Platform, DevOps, AI, Product, and Support teams to align infrastructure capabilities with technical and customer requirements.
- Contribute to the continuous evolution of infrastructure architecture and operational practices as global environments scale.
- Identify opportunities to improve automation, performance, reliability, and operational efficiency across the infrastructure platform.
- Extensive hands-on experience administering Linux systems, with strong technical depth across operating systems, troubleshooting, and infrastructure operations.
- Proven experience physically building, assembling, cabling, configuring, and commissioning servers and racks.
- Direct hands-on experience working in data centres, including physically installing and racking hardware on-site.
- Willingness and ability to travel to data-centre sites in Quebec as required.
- Strong understanding of networking and storage technologies and their role within large-scale infrastructure environments.
- Experience installing, racking, and configuring GPU hardware is highly advantageous, particularly with NVIDIA platforms.
- Production experience operating OpenStack and/or Kubernetes at scale is strongly preferred.
- Experience with infrastructure automation, infrastructure-as-code, CI/CD pipelines, GitOps, and Git-based workflows is an advantage.
- Exposure to high-performance computing, large-scale compute, or other demanding infrastructure environments is desirable.
- Experience with Kubernetes GPU scheduling and NVIDIA tooling is a plus.
- Strong troubleshooting and analytical abilities, with a methodical approach to diagnosing complex infrastructure issues.
- Ability to take ownership of systems from design through deployment, operation, troubleshooting, and continuous improvement.
- Comfortable working in a fast-paced environment where priorities can evolve and engineers are trusted to make decisions independently.
- Strong collaboration and communication skills, with the ability to work effectively across technical and non-technical teams.
- Contributions to open-source infrastructure or technology projects are a plus.
- Competitive salary.
- Annual discretionary bonus scheme.
- Employee wellbeing benefits.
- 25 days of annual holiday plus public holidays.
- Flexible working arrangements, with remote or hybrid options depending on role and location.
- Significant autonomy and ownership, with the freedom to take initiative and experiment.
- Opportunity to work on challenging, high-performance AI and GPU infrastructure.
- Visible impact on infrastructure reliability, scalability, and customer experience.
- Clear career progression and professional growth opportunities.
- Collaborative international environment built around trust, transparency, and ownership.
- Opportunity to contribute to the evolution of a rapidly scaling AI infrastructure platform and its engineering culture.

