Location
In office
Seniority
Senior
Overview
Company
Impact
Profile match
Link Group is a Polish technology services company founded in Warsaw in 2016 that builds and supplies engineering teams to clients across Europe. Its model combines body leasing and managed teams with delivery of complete software, cloud and cybersecurity projects, drawing on a large bench of contractors rather than a fixed permanent staff. The company works heavily in financial services, telecommunications and public sector projects, and has built a specialisation in blockchain and distributed ledger engineering alongside more conventional cloud and application development work.
We are looking for a seasoned Site Reliability Engineer to join the team responsible for the backbone of our global AI/ML services. This isn't your typical SRE role. You won't just be maintaining systems; you'll be the guardian of a massive, distributed AI compute platform that processes workloads at an incredible scale. You will ensure that our AI models and GPU-powered infrastructure are not just fast, but fundamentally reliable, observable, and built to last.
If you are passionate about building and operating large-scale systems and are excited by the unique challenges of the AI/ML world, this is the role for you.
What You Will Do (Your Impact):
- Build Bulletproof Observability: You will design and implement the "nervous system" for our AI platform. This means going beyond basic monitoring to build comprehensive observability with robust telemetry, insightful dashboards (Grafana), and intelligent alerting (Prometheus). You will define and track SLOs/SLIs to ensure our services meet their promises and drive improvements when they don't.
- Automate Everything: Your mantra is "if you have to do it twice, automate it." You will write clean, effective code in Python or Go to eliminate manual toil, create self-healing systems, and build sophisticated tooling that accelerates incident response and makes deployments safer.
- Own the Incident Response Lifecycle: When critical systems fail, you will be on the front lines. You will lead the charge in incident management, participate in a blameless on-call rotation, and conduct insightful post-mortems that lead to real, lasting improvements. You'll build the runbooks that others will rely on.
- Engineer a World-Class Deployment Pipeline: You will be a key contributor to our CI/CD ecosystem, building rock-solid integrations, automated safety checks, and seamless rollback capabilities to ensure that we can innovate at speed without sacrificing stability.
- Act as a Reliability Partner for Product Teams: You will work side-by-side with product engineers who are building the next generation of AI services. You will be their trusted advisor on reliability, helping shape their architecture and ensuring their products are operationally sound long before they hit production.
Who We're Looking For (Your Profile):
- You are a true Site Reliability, Platform, or Infrastructure Engineer at heart, with a proven track record of managing complex, large-scale distributed systems.
- Kubernetes is your natural habitat. You have deep, practical experience managing large-scale containerized environments and understand the complexities of orchestration under heavy load.
- You speak the language of observability fluently, with hands-on experience using tools like Prometheus, Grafana, and distributed tracing systems to make systems transparent and understandable.
- You are a strong programmer. You write clean, scalable automation scripts and infrastructure-as-code using Python or Go and tools like Terraform.
- You have a genuine curiosity or, ideally, direct experience with the unique challenges of AI/ML infrastructure, such as model serving pipelines, inference engines, or managing GPU-accelerated workloads.
- You are a problem-solver who takes full ownership of issues from start to finish. When you see a problem, you don't just fix it-you figure out how to prevent it from ever happening again.
- You excel at collaboration and enjoy mentoring other engineers, helping them adopt SRE principles and build more reliable software from the ground up.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,695 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
In your city
In office • Full-Time • Toulouse
Python
Databases
ElasticSearch
AI/ML
Copilot
LLM
Ollama
vLLM
DevOps
CI/CD
GitHub
GitLab
GitLab CI
Grafana
Kibana
Kubernetes
Logstash
Prometheus
Apply
Cloud Native Platform Engineer
4 hours ago
≈ $81k – $171k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bern • Zurich
Python
DevOps
Ansible
CI/CD
GitOps
Kubernetes
Apply
≈ $26k – $54k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Databases
Amazon DocumentDB
DynamoDB
PostgreSQL
DevOps
Amazon CloudWatch
AWS
Azure
CI/CD
Datadog
Docker
GCP
Git
GitHub
GitHub Actions
GitLab
Jenkins
Kubernetes
OpenShift
Splunk
Terraform
Apply
AQA Lead/ Лид автотестирования
4 hours ago
≈ $19k – $53k per year (Estimated) • Remote • 5+ years exp • Saint Petersburg
Python
TypeScript
JavaScript
Python
FastAPI
Databases
PostgreSQL
Frontend
Angular
DevOps
ArgoCD
CI/CD
Docker
Git
GitLab
gRPC
Jenkins
Kubernetes
QA
Playwright
Postman
Pytest
Selenium
Apply
Senior Software Developer - Full Stack Developer
4 hours ago
In office • 6+ years exp • Bachelor's Degree
C#
JavaScript
SQL
TypeScript
AI/ML
Copilot
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
GitHub
Kubernetes
Management
Microsoft Teams
Apply
In office • Master's Degree
Python
DevOps
Ansible
Chef
Configuration Management
KVM
Puppet
QEMU
SaltStack
Apply
In office • Master's Degree
Python
AI/ML
Anomaly Detection
LLM
DevOps
Grafana
Loki
OpenTelemetry
PagerDuty
Prometheus
Apply
Remote/Hybrid • 7+ years exp
Python
SQL
Databases
Apache Kafka
Kafka
DevOps
ZooKeeper
Apply
Senior Webscraping & Data Engineer
1 day ago
Remote/Hybrid • 6+ years exp
JavaScript
Python
SQL
Databases
Apache Kafka
Kafka
AI/ML
AI Agents
LLM
DevOps
AWS
Docker
Kubernetes
Apply
Senior Identity Infrastructure Engineer
1 day ago
Remote/Hybrid
PowerShell
DevOps
Azure
IAM
Windows Server
Cybersecurity
Microsoft Entra ID
Management
ServiceNow
Apply
This is one of many
386,695 more open roles from verified company boards, updated every day.

