Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on May 29, 2026. Hadrian scores B on the Alion truth index.
Join our team as a Site Reliability Engineer (Robotics) and take ownership of the reliability of our robotics systems. You will build interfaces for observability, write code frameworks and tools, and partner with various engineering teams to ensure reliability is baked in early. Enjoy benefits such as a relocation stipend, 100% coverage of platinum medical, dental, vision, and life insurance plans, a 401k, and a flexible vacation policy.
Missions
- Assurer la fiabilité des systèmes robotiques, des PLCs à ROS2/middleware en passant par Kubernetes.
- Construire des interfaces pour le système d'observabilité afin d'ingérer la télémétrie des systèmes de contrôle et robotiques.
- Écrire des frameworks et des outils de code pour soutenir les systèmes de contrôle et robotiques, y compris les outils de diagnostic.
Profil recherché
- Problem Solver. Solving complex puzzles excites and motivates you to find an efficient solution- Systems Thinker. Focused on understanding the relationship among various systems to design sustainable solutions, not one-time fixes
- Strong Communication. You can run a war room, write a post-mortem, and explain a reliability tradeoff to a stakeholder
- T-Shaped Skill Set. Comfortable with bare metal Kubernetes, networking, GitOps workflows, and Infrastructure as Code (IaC). Also skilled in programming in TypeScript, Python, Golang, or C++
- Ownership. Someone who has owned the reliability of a production system where downtime had physical or operational consequences (manufacturing line, autonomous vehicle, lab automation, network operations)
- You've built automated or self-healing remediation at scale. We want systems that remove humans from the loop
- Background in edge/on-prem infrastructure. You’ve run Kubernetes at the edge (k3s, k0s, k0smotron), managing on-prem clusters, time-series at the edge, or air-gapped deployments. A deep understanding of Linux operating system fundamentals such as cgroups, sockets, and system tuning, is a big plus
- Deep understanding of shipping and storing telemetry data at scale. Experience with Kafka/MQTT/RabbitMQ is a plus
- Direct robotics experience. ROS/ROS2, OPC UA, EtherCAT, motion controllers, or fleet management for autonomous systems
- An individual who is self-directed and can deliver with high velocity

