First seen by Alion on Oct 11, 2026.
Remote work only from within Poland
Do you enjoy solving complex reliability challenges for cutting-edge technology?
Do you have a passion for automation and building systems that scale?
As a Senior SRE, responsibilities include owning reliability workstreams for the company's serverless inference platform, building automation and tooling, and contributing to architecture and operational decisions. Opportunities exist to take ownership of critical reliability problems end-to-end, partner with product engineering teams, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads.
As a Site Reliability Engineer, you will be responsible for:
Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed
Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
Integrating AI workloads into our existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems
Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation
Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases
Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure
Do what you love
To be successful in this role you will:
Demonstrate expertise in SRE, infrastructure, or platform engineering, managing large-scale distributed systems with extensive operational experience.
Demonstrate expertise in Kubernetes and large-scale containerization systems.
Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
Demonstrate proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code like Terraform.
Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads
Resolve issues independently while maintaining accountability throughout the process.
Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.
Working for you
We will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life:
Your health
Your finances
Your family
Your time at work
Your time pursuing other endeavors
Our benefit plan options are designed to meet your individual needs and budget, both today and in the future.
Tech stack
- AI
- GPU
- Kubernetes
- Python
- Linux

