Location
In office
Seniority
Senior
Overview
Company
Impact
Profile match
Link Group is a Polish technology services company founded in Warsaw in 2016 that builds and supplies engineering teams to clients across Europe. Its model combines body leasing and managed teams with delivery of complete software, cloud and cybersecurity projects, drawing on a large bench of contractors rather than a fixed permanent staff. The company works heavily in financial services, telecommunications and public sector projects, and has built a specialisation in blockchain and distributed ledger engineering alongside more conventional cloud and application development work.
We are looking for an elite Site Reliability Engineer to join the core team responsible for our global AI compute infrastructure. Your mission will be to ensure that the physical and virtualized backbone of our AI platform-the bare-metal servers, high-density GPU racks, and the advanced network that connects them-is exceptionally reliable, performant, and scalable.
This is a role for a hands-on engineer who is as comfortable writing Python automation and Infrastructure-as-Code as they are designing BGP routing strategies and collaborating with data center technicians.
What You Will Do (Your Impact):
- Become a Master of Fleet Automation: You will write sophisticated tooling and automation in Python to manage the entire lifecycle of our server fleet. Your code will handle everything from initial provisioning and configuration to ongoing maintenance and decommissioning, eliminating manual effort across thousands of machines.
- Architect Seamless Operational Workflows: You will integrate our core operational systems, connecting platforms like JIRA, Siebel, and PagerDuty through robust APIs. Your goal is to create automated workflows that dramatically reduce the time it takes to resolve hardware and network incidents.
- Build the Future of Observability: You will design and implement a world-class observability stack tailored for bare-metal and virtualized hardware. This includes building custom telemetry pipelines, creating insightful Grafana dashboards with data from Prometheus, and pioneering the use of AI-driven anomaly detection to predict failures before they happen.
- Lead in Times of Crisis: As a senior member of the team, you will be a leader during critical incidents. You'll participate in a 24/7 on-call rotation, spearhead the response to high-severity outages, and drive comprehensive, blameless post-mortems that result in concrete architectural improvements.
- Leverage AI to Build Better Systems: We believe in using our own tools. You will actively use advanced AI utilities and LLM-assisted development to enhance your own technical execution, from generating complex automation scripts to evaluating system performance.
Who We're Looking For (Your Profile):
- You have a deep background in Site Reliability or Production Engineering, built on a solid Computer Science foundation and proven experience managing large-scale, mission-critical infrastructure.
- You are an exceptional Python programmer. You don't just write scripts; you build scalable, robust operational tools and automation frameworks from the ground up.
- You are a Networking expert. You have a strong, practical understanding of advanced network topologies, high-bandwidth routing and switching, BGP, and the complexities of dual-stack IPv4/IPv6 environments.
- You live and breathe Observability. You have hands-on, expert-level experience with modern monitoring stacks like Prometheus, Grafana, OpenTelemetry, and Loki.
- You are an operational leader. You have extensive experience designing service rollout strategies, defining meaningful alerting thresholds, creating clear technical runbooks, and leading incident response "war rooms."
- You are a natural owner. You thrive on solving ambiguous, complex technical problems and have a proven ability to take a challenge from a vague idea to a production-grade, fully automated solution.
- You are a strong collaborator, able to partner effectively with external data center vendors and coordinate with on-site field technicians to ensure maximum uptime.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,695 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
In your city
In office • Full-Time • Toulouse
Python
Databases
ElasticSearch
AI/ML
Copilot
LLM
Ollama
vLLM
DevOps
CI/CD
GitHub
GitLab
GitLab CI
Grafana
Kibana
Kubernetes
Logstash
Prometheus
Apply
≈ $43k – $121k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Madrid
JavaScript
Python
Databases
Snowflake
AI/ML
LLM
Prompt Engineering
RAG
Analytics
Tableau
Management
Google Docs
ServiceNow
Apply
MLOps Engineer
4 hours ago
≈ $27k – $113k per year (Estimated) • In office • Gurgaon
AI/ML
CUDA
CUDA Toolkit
LLM
DevOps
AWS
Azure
CI/CD
GCP
Apply
≈ $64k – $152k per year (Estimated) • In office • Full-Time • France
Python
AI/ML
LangChain
LangGraph
LLM
Ollama
Qwen
RAG
vLLM
DevOps
Docker
Git
Apply
Data Analyst II, Incentive Compensation
4 hours ago
≈ $18k – $43k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Databases
Databricks
AI/ML
Anomaly Detection
DevOps
AWS
Azure
CI/CD
GCP
Analytics
ETL/ELT
Power BI
Tableau
Apply
In office • Master's Degree
Python
DevOps
Ansible
Chef
Configuration Management
KVM
Puppet
QEMU
SaltStack
Apply
Remote/Hybrid • 7+ years exp
Python
SQL
Databases
Apache Kafka
Kafka
DevOps
ZooKeeper
Apply
Senior Webscraping & Data Engineer
1 day ago
Remote/Hybrid • 6+ years exp
JavaScript
Python
SQL
Databases
Apache Kafka
Kafka
AI/ML
AI Agents
LLM
DevOps
AWS
Docker
Kubernetes
Apply
Senior Identity Infrastructure Engineer
1 day ago
Remote/Hybrid
PowerShell
DevOps
Azure
IAM
Windows Server
Cybersecurity
Microsoft Entra ID
Management
ServiceNow
Apply
In office
Python
DevOps
CI/CD
Grafana
Incident Management
Kubernetes
Prometheus
Self-Healing
Terraform
Apply
This is one of many
386,695 more open roles from verified company boards, updated every day.

