368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$62k – $162k per year (Estimated)
Location
Remote (Spain)
Seniority
Middle · 3+ years exp
Overview
Company
Impact
Profile match
HostPapa is a Canadian hosting company founded in 2006 and headquartered in Burlington, Ontario. It provides domains, web hosting and cloud services to small businesses. The company serves customers in dozens of countries.

Position Summary:

With team members and customers in 39 countries around the globe, HostPapa is currently one of the fastest-growing web hosting companies with a wide range of products available. At its core, we provide individuals and small and medium-sized businesses with access to valuable tools and services critical to their online success, including a Website Builder service for making website creation an ultra-easy task for anyone. Tailored to meet every user's unique needs, our award-winning customer support, email, and cloud-based solutions keep HostPapa at the cutting edge of the web hosting industry and innovation by putting our customers first.

This role focuses on CloudBlue, a HostPapa business that powers cloud commerce for many of the world’s largest service providers, including major Telcos, distributors, and MSPs. CloudBlue enables partners to monetize and manage cloud services and subscriptions at scale, combining the agility of a high-growth business with the backing of a global organization.

As the Site Reliability Engineer, you will help ensure the reliability, scalability, and observability of CloudBlue’s multi-tenant SaaS platforms used by service providers worldwide. You will focus on improving system stability and performance through monitoring, high availability, and incident response, while working closely with DevOps, Platform, and Engineering teams to build and operate resilient production systems.

What you’ll do

  • Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services to ensure reliability and performance
  • Influence system architecture with a strong focus on reliability, scalability, and operability, designing systems for fault tolerance, graceful degradation, and self-healing
  • Reduce operational toil by identifying opportunities for automation and process improvement
  • Design and operate CloudBlue’s observability stack across metrics, logs, and traces using tools such as Datadog, Grafana, and Elastic Stack
  • Develop actionable alerting strategies and dashboards that provide clear insight into platform and business health
  • Design and maintain high-availability architectures, implementing redundancy, failover, and disaster recovery strategies across regions and availability zones
  • Conduct capacity planning, load testing, and performance optimization to ensure platform stability and scalability
  • Act as a senior responder during production incidents, leading incident coordination, communication, and service restoration
  • Own blameless postmortems and drive improvements that reduce incident frequency, MTTR, and customer impact
  • Improve reliability of Kubernetes-based platforms through health checks, autoscaling strategies, rollout safety, and resilience testing
  • Partner with engineering and DevOps teams to improve deployment safety, rollback strategies, and platform reliability
  • Maintain runbooks and operational documentation, and promote SRE best practices across engineering teams
  • Support other tasks or projects as assigned to meet team and business needs

About you

  • 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer, with strong ownership of production systems
  • Proven experience operating highly available, enterprise-grade, multi-tenant SaaS platforms
  • Hands-on experience with observability and monitoring tools such as Datadog, Grafana, and Elasticsearch/Kibana
  • Solid understanding of Linux, networking, and distributed systems fundamentals
  • Experience working with containerized environments such as Docker and Kubernetes
  • Strong scripting and automation skills using Python and/or Bash
  • Experience participating in on-call rotations and incident response in production environments
  • Strong written and spoken English
  • Experience defining SLIs/SLOs and managing error budgets at scale will be considered a plus
  • Exposure to hyperscale or service-provider-grade platforms is an advantage
  • Cloud experience, preferably with Azure; experience with AWS and/or GCP will also be valued
  • Experience working with hybrid or on-premises integrations is beneficial
  • Familiarity with chaos engineering and resilience testing will be considered an asset

What We Offer:

  • This is a remote opportunity. While we welcome applications globally, we are prioritizing candidates based in Spain
  • A competitive salary that values you and your unique skill sets
  • Career advancement & professional development opportunities to help you reach your full potential
  • Flexible work arrangements to support work/life balance

About Us:

At HostPapa, we’ve been committed to providing a complete array of enterprise-grade cloud services solutions to every business owner since 2006. These services, traditionally out of reach to smaller businesses, are offered in a one-stop shop, making it quick and easy for customers to select the services they need to grow. We back these offerings with 24/7 award-winning customer support in four languages.

Our HostPapa team values diversity and inclusion. We have a friendly company culture built on trust and respect. With the acquisition of several companies into our product portfolio, we’re growing at an incredible rate and have ample opportunities for career growth.

Come join our talented team of enthusiastic, hard-working, passionate, driven people engaged in meaningful, innovative work. We can’t wait to meet you!

HostPapa is an equal-opportunity employer committed to diversity and inclusion. As a multicultural organization, we encourage individual achievement and recognize the strength of our diverse team.

HostPapa is committed to providing accommodations for people with disabilities. If you require accommodation, please let us know, and we will work with you to meet your needs. Accommodation may be provided in all parts ofthe hiring process.

It is anticipated that this position will be performed outside of Ontario.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$58k – $136k per year (Estimated) • Remote/Hybrid • Full-Time • Guadalajara
Bash
Node JS
Python
JavaScript
Python
FastAPI
Databases
RabbitMQ
Frontend
Next.js
React.js
DevOps
ArgoCD
AWS
CI/CD
Datadog
Docker
Git
GitHub
GitHub Actions
GitOps
Grafana
Helm
Jenkins
Karpenter
KEDA
Kubernetes
OpenTelemetry
Prometheus
Terraform
Apply
$79k – $159k per year (Estimated) • In office • Full-Time • 6+ years exp • Lincoln
C#
C#
.NET
DevOps
AWS
Azure
CI/CD
Docker
Dynatrace
GCP
GitHub
GitHub Actions
Grafana
Jenkins
Kubernetes
Splunk
QA
Cypress
JMeter
k6
Pact
Playwright
Postman
Rest-Assured
Selenium
Supertest
WebDriverIO
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$87k – $130k per year • In office • Full-Time • 6+ years exp • Murray
Python
TypeScript
JavaScript
Python
Alembic
FastAPI
Pydantic
SQLAlchemy
Databases
PostgreSQL
Redis
AI/ML
Embeddings
LLM
LLM Guardrails
Ollama
RAG
Frontend
React Query
React Router
React.js
Vite
DevOps
AWS
CI/CD
Docker
Docker Compose
IAM
Terraform
Cybersecurity
FedRAMP
NIST 800-53
QA
Pytest
Apply
$152k per year • Equity • Remote • 8+ years exp
Lua
Python
Ruby
Rust
YARA
Databases
ElasticSearch
DevOps
Amazon ECS
Splunk
Cybersecurity
Cyber Kill Chain
Scapy
Snort
Suricata
Tcpdump
Wireshark
YARA
Zeek
Apply
$109k – $245k per year (Estimated) • Remote • 5+ years exp
SQL
Robotics
Apollo
Management
n8n
Zapier
Marketing
GA4
HubSpot
Apply
$121k – $228k per year (Estimated) • Remote • 5+ years exp
Python
SQL
DevOps
AWS
Azure
GCP
Kubernetes
Apply
$159k – $315k per year (Estimated) • Remote
AI/ML
AI Agents
LLM
Apply
Remote
PowerShell
Python
DevOps
AWS
CI/CD
GCP
Apply
Remote
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.