Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Jun 24, 2026.
Join our DevOps team as a Site Reliability Engineer in New York. You will be responsible for the reliability, stability, and operational support of our production systems. This role involves production ownership, monitoring, incident response, and on-call support. You will work with a modern cloud-native stack and play a key role in keeping systems highly available, secure, and performant.
Missions
- Assurer la fiabilité, la disponibilité et la performance des environnements de production, y compris la gestion des clusters Kubernetes et des environnements AWS.
- Évoluer l'infrastructure en tant que code en utilisant Terraform et Helm, et soutenir les pipelines CI/CD de GitLab.
- Diriger la réponse aux incidents de bout en bout, y compris le dépannage, l'atténuation et la résolution.
Profil recherché
- Strong troubleshooting skills across infrastructure, CI/CD, and networking- Scripting experience with Bash and Python
- Willingness to participate in on-call rotations
- Experience with monitoring, logging, and alerting systems
- Proficiency with Terraform, Helm, and GitLab CI (or similar)
- Familiarity with pub/sub systems (SQS, Kafka, or similar)
- 3+ years of hands-on DevOps / SRE experience
- Strong production experience with Docker and Kubernetes
- Solid knowledge of AWS (EKS, EC2, Organizations, RDS, S3, CloudWatch, Lambda, DynamoDB)
- GitOps workflows and advanced Git usage
- Experience supporting databases such as Postgres, Snowflake, or ClickHouse
- Experience with Redis, Airflow, Databricks, Spark/EMR

