661,119open jobs
38,508companies
97,957added this week
Browse all
Salary
$48k – $117k per year (Estimated)
Location
Remote (Spain)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Spain.

This is a senior infrastructure role focused on building and operating highly reliable cloud platforms in a distributed engineering environment.

You will take ownership of production infrastructure across Kubernetes, Linux, networking, virtualization, and bare-metal environments.

The role combines deep technical expertise with automation, observability, incident management, and proactive reliability engineering.

You will help shape infrastructure architecture, improve availability and performance, and establish scalable operational practices.

Working closely with engineering and cross-functional teams, you will solve complex infrastructure challenges and optimize resource utilization.

The environment is fully remote, international, and highly collaborative, with significant autonomy and ownership.

This is an opportunity to make a direct impact on ambitious cloud infrastructure projects while working with modern technologies.

Accountabilities

    • Operate, maintain, and continuously improve Linux-based infrastructure, with a strong focus on Debian and Ubuntu environments.
    • Deploy, manage, and scale production Kubernetes clusters across bare-metal, virtualized, and on-premise environments, overseeing upgrades, node pools, networking, storage, and security hardening.
    • Design, implement, and maintain complex networking architectures covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
    • Build and maintain infrastructure automation using Ansible, Bash, Python, Git-based workflows, and GitOps practices, including automated provisioning through PXE boot, Preseed, and cloud-init.
    • Deploy and maintain observability and monitoring platforms such as Prometheus, Grafana, Loki, ELK, and Graylog, ensuring operational data generates actionable insights.
    • Lead incident response and escalation activities, troubleshoot complex infrastructure issues, and implement improvements that increase availability and reduce latency.
    • Define and implement SLOs and SLIs across physical infrastructure, networking, virtualization, and software services to establish measurable reliability standards.
    • Optimize alerting and monitoring pipelines while establishing effective on-call schedules to provide operational coverage across time zones.
    • Create and maintain Standard Operating Procedures for recurring infrastructure operations, maintenance, troubleshooting, and incident management.
    • Coordinate physical infrastructure maintenance, including hardware issues, periodic maintenance, and data-center operations.
    • Manage virtualization and orchestration layers using technologies such as OpenStack, Proxmox, and VMware.
    • Contribute to the overall architecture and evolution of infrastructure products, ensuring solutions remain scalable and reliable.
    • Plan infrastructure capacity and resources for future initiatives based on projected demand and business growth.
    • Partner with development teams to improve system quality, optimize resource utilization, and strengthen engineering practices.
    • Collaborate with cross-functional stakeholders to align infrastructure priorities with broader product and customer needs.
    • Requirements:

      • Expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, storage, security, and scaling.
      • Strong network engineering expertise is essential, particularly across VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
      • Strong Linux systems administration skills, particularly with Debian and Ubuntu.
      • Solid understanding of networking fundamentals and the ability to design and operate complex network architectures.
      • Proven experience developing infrastructure automation using Ansible, Bash and/or Python, Git-based workflows, and GitOps methodologies.
      • Practical experience with observability platforms such as Prometheus, Grafana, ELK, Loki, or Graylog.
      • Experience working with virtualization technologies including OpenStack, Proxmox, and VMware.
      • Experience with bare-metal provisioning and MAAS (Metal as a Service).
      • Strong understanding of distributed systems and container orchestration.
      • A process-oriented mindset, with the ability to create SOPs and operational procedures from the ground up.
      • Experience managing production incidents, escalation processes, and on-call rotations.
      • Ability to work independently and make sound technical decisions in a fast-paced, engineering-driven environment.
      • Strong communication and collaboration skills, combined with a high level of technical ownership and alignment with team values.
      • Fluent English is mandatory.
      • Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI implementations, is a plus.
      • Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations is advantageous.
      • Experience with GPU infrastructure, node preparation, resource scheduling, security practices such as RBAC, firewalls, and network policies is beneficial.
      • Familiarity with IT asset management or license tracking workflows is an advantage.
      • Experience working across multiple time zones and establishing SRE or reliability frameworks within growing organizations is highly valued.
      • Benefits:

        • 100% remote work within the EU time zone, with CET ±2 hours preferred.
        • Flexible working hours designed to support autonomy and effective collaboration.
        • High-impact position with significant ownership and the opportunity to influence infrastructure strategy.
        • Opportunity to work with a modern technology stack spanning Kubernetes, cloud infrastructure, networking, virtualization, automation, and observability.
        • Collaborative and international engineering environment with exposure to complex infrastructure challenges.
        • Significant autonomy to shape operational processes, reliability practices, and technical solutions.
        • Opportunity to contribute to ambitious cloud infrastructure initiatives with a strong focus on reliability, automation, and continuous improvement.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
661,119 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$68k – $163k per year (Estimated) • In office • Full-Time • Nuremberg
Java
Java
Spring Boot
DevOps
Terraform
Ansible
GCP
GitHub Actions
Loki
Prometheus
GitLab CI
Azure
CI/CD
GitOps
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
IAM
Apply
$215k – $428k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Python
Databases
PostgreSQL
Apache Kafka
AI/ML
OpenAI
Replit
DevOps
AWS
Docker
Kubernetes
Apply
$39k – $93k per year (Estimated) • Remote/Hybrid • Full-Time
Python
SQL
Python
FastAPI
Databases
PostgreSQL
Snowflake
AI/ML
dbt
Pandas
DevOps
Terraform
GitHub Actions
CloudFormation
Prometheus
GitLab CI
CI/CD
Git
AWS
Grafana
AWS Lambda
Amazon CloudWatch
Analytics
ETL/ELT
A/B Testing
Apply
Remote • Full-Time • Bachelor's Degree
Python
Java
Rust
Scala
AI/ML
Spark
DevOps
GCP
Azure
Git
AWS
Bitbucket
GitHub
Amazon S3
Management
Agile
Apply
$53k – $130k per year (Estimated) • In office • Bengaluru
Python
Java
C++
Apply
$31k – $76k per year (Estimated) • Remote • Contractor • 3+ years exp
Apply
$39k – $93k per year (Estimated) • Remote/Hybrid • Full-Time
Python
SQL
Python
FastAPI
Databases
PostgreSQL
Snowflake
AI/ML
dbt
Pandas
DevOps
Terraform
GitHub Actions
CloudFormation
Prometheus
GitLab CI
CI/CD
Git
AWS
Grafana
AWS Lambda
Amazon CloudWatch
Analytics
ETL/ELT
A/B Testing
Apply
Remote • Full-Time • 2+ years exp
DevOps
Azure DevOps
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Git
Management
Jira
Agile
QA
TestRail
Selenium
JMeter
Cypress
Playwright
Postman
Rest-Assured
Locust
Apply
Remote • Full-Time • Bachelor's Degree
Python
Java
Rust
Scala
AI/ML
Spark
DevOps
GCP
Azure
Git
AWS
Bitbucket
GitHub
Amazon S3
Management
Agile
Apply
$51k – $113k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree
SQL
C#
C#
.NET
Databases
Amazon Aurora
DevOps
Rest API
CI/CD
AWS
AWS Fargate
AWS Lambda
Amazon S3
Amazon ECS
Amazon CloudWatch
API Gateway
Apply
See all jobs
This is one of many
661,119 more open roles from verified company boards, updated every day.