588,947open jobs
26,464companies
82,741added this week
Browse all
Location
In office
Seniority
Senior · 5+ years exp
Employment
Contractor
Overview
Company
Impact
Profile match
Save, secure, and scale your cloud with GCP, AWS, Azure, and/or Huawei. Partner with Deimos for expert guidance in hybrid multi-cloud solutions!

Deimos is a Cloud-native Developer and Security Operations technology services company. We help companies of all sizes adopt the Cloud for improved service delivery to their clients. We’re a fully remote African-based team of engineers who are passionate about implementing engineering best practices. We leverage the latest technologies while building globally competitive solutions for our clients. With Deimos being one of the two moons of Mars, we refer to ourselves as “Martians” who are on a mission to Mars, together.

Our teams value the ability to learn and adapt to technology changes while appreciating solid foundational design and the craft of software engineering. As such our engineers enjoy working with various clients who have different problems to solve. If this sounds like you then you would be an ideal fit for our environment. However, you must be based in one of the countries we currently hire in which are as follows: Kenya, Ghana, Nigeria, South Africa, and Senegal.

Role Overview

We are looking for an experienced Senior Site Reliability Engineer to join our Professional Services team and deliver Software and DevSecOps projects. You will report to a Site Reliability Engineering Manager.

SRE / DevOps is one of our core competencies. You will be part of a highly-skilled team that continuously innovates and delivers high value solutions to clients across various industries on all public clouds (AWS, Azure, GCP, etc). Technologies we work with daily include Kuberenetes, Helm, Terraform, GitOps, just to name a few.

What you will be doing

  • Enablement & RelOps Culture
    • Implement the Observability Ladder: Guide teams from basic monitoring to high-signal metric tracking. Work with product teams to define SLAs, SLIs, and SLOs, and build dashboards that track specific error budgets.
    • Empower Product Teams: Build frameworks and deployment tooling (e.g., CI/CD, internal tooling integrations) that allow teams to make data-driven decisions on deployment safety and automate rollbacks when error budgets are depleted.
    • Champion Reliability: Drive a blameless post-mortem culture focused on actionable takeaways, system improvements, and measurable metrics (MTBF, MTTR).

    Frameworks & Automation

    • Standardised Alerting & On-Call: Continuously improve company-wide alerting and on-call frameworks to reduce alert fatigue, ensuring alerts are highly actionable and symptom-based.
    • Disaster Recovery: Drive evolution of DR strategies from manual processes into fully automated runbooks-as-code, allowing teams to prove and improve service recoverability through autonomous, evidence-based testing.
    • Eliminate Toil: Develop systems, automations, and tooling for pre- and post-deployment verification, ensuring our hands-off reliability vision becomes a production reality, via Python (or similar).
    • Reliability-as-Code: Lead the drive to manage our entire reliability suite through IaC. Use Terraform to architect, deploy, and configure our observability stack including ELK, Grafana, Loki, Prometheus, and Tracing.

What you must have

  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • 5+ years of experience in Software Engineering, SRE, DevOps, or Platform Engineering, with demonstrable ownership of reliability standards at a team or company level.
  • Strong coding fluency: Proficiency in Python (or similar) with the ability to read, understand, reason about, and write production-grade automation code.
  • Cloud & IaC: Hands-on experience with AWS, and a solid understanding of Infrastructure as Code (Terraform or CloudFormation).
  • Deep Observability Knowledge: Demonstrable experience with monitoring tools (DataDog, Prometheus, ELK stack). Strong understanding of SRE concepts including Golden Signals, high-cardinality data handling, and error budget mathematics.
  • Systems Thinking: Strong grasp of designing for scale and resilience, including graceful failure, circuit breaking, connection pooling, and multi-AZ deployments.
  • Proven ability to define and drive reliability standards across multiple teams and drive a blameless post-mortem culture.

Qualities & Behaviours

  • Exceptional interpersonal and communication skills
  • A zest for automation.
  • Comfortable working as a remote team member.
  • Ability to keep up to date with DevOps/SRE best practices, trends and innovation.
  • Passionate about mentoring and growing technical skills within the team.

Expected Output for the role

  • Automate Azure infrastructure provisioning and configuration using PowerShell, YAML

    and Bicep.

  • Monitor and troubleshoot issues in the Azure environment, including network, storage, and compute resources.
  • Deploy and manage Azure Databricks infrastructure for data processing and analytics.
  • Attend to support tickets, which may arise due to product components not functioning

    as expected.

  • Develop and maintain technical support documentation of the product.
  • Promote innovations to support business requirements through activities that test, pilot

    and implement innovative concepts.

  • Responsible for support and troubleshooting DevOps tools and processes for

    stakeholders

About you

For us to achieve our ambitious vision together as a team, It is important for our Martians to lead at all levels, be self starters who take initiative and put their hands up for challenging tasks. A growth mindset is important to us and we encourage all our Martians to openly share knowledge, support and help each other, ask questions, get creative with new technologies and learn from setbacks.

Becoming a Martian means:

  • Comfortably working and learning from a fully remote, culturally diverse team based predominantly in South Africa, Kenya, Nigeria and Ghana.
  • Being an open, honest and respectful communicator.
  • You enjoy asking questions, identifying areas of improvement and proposing solutions, no matter your job title or whether you have been with us for a day, a month or years!
  • You are comfortable taking initiative and operating independently.
  • You thrive in a fast paced environment, where change is constant.
  • You find it exciting to work with various clients, from different industries, each with a different problem for you and your team to solve.
  • Intentionally sharing tech and industry trends that excite you with your peers.
  • Seeking continuous feedback and actively taking steps to continuously grow personally and professionally.

Want to know what you get by joining us?

  • Become a member of a team where we value each individual's contribution from day 1 and empower you to make suggestions, get involved and do what you love most!
  • Flexibility and the freedom to work remotely.
  • Work-life balance where you are not expected to work over weekends or after hours.
  • A forward thinking remote company that knows how important it is to stay connected as one team, by providing virtual social platforms for employee engagement.
  • A monthly work from home allowance which you can use to set yourself up to work comfortably from home. Whether that is pens, notebooks, new headphones or work snacks!
  • A MacBook or Windows laptop for you to do your best work on.
  • Become part of a team of exceptionally clever and talented people who like to share their knowledge and learnings.
  • We support your career growth and love to celebrate your successes and advancement!
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
588,947 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$18k – $41k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Hyderabad
Python
PowerShell
R
SAS
R
Shiny
AI/ML
Copilot
Claude Code
Streamlit
DevOps
Terraform
Ansible
SLURM
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Ubuntu
Platform Engineering
Configuration Management
Bitbucket
JFrog Artifactory
Amazon EKS
HPC
Management
Agile
Scrum
Apply
$35k – $58k per year (Estimated) • In office • Internship • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
Analytics
Tableau
Power BI
Microsoft Excel
Apply
$38k – $104k per year (Estimated) • In office • Full-Time • Master's Degree • Madrid
Python
Databases
ElasticSearch
AI/ML
Model Context Protocol
Embeddings
Scikit-learn
Prompt Engineering
AI Agents
NLP
TensorFlow
PyTorch
LLM
RAG
GraphRAG
DevOps
Azure
Git
AWS
Vector
Chips/EDA
PoC Library
Analytics
A/B Testing
Management
Agile
Apply
$17k – $43k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Pune
PowerShell
DevOps
Incident Management
Cybersecurity
Microsoft Entra ID
Chips/EDA
PoC Library
Apply
$18k – $40k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Hyderabad
SQL
Databases
Databricks
DevOps
AWS
Analytics
Tableau
Power BI
ETL/ELT
Apply
In office • Contractor • 10+ years exp
JavaScript
TypeScript
Node JS
Node JS
Nest.JS
Fastify
Databases
MySQL
PostgreSQL
Redis
DynamoDB
RabbitMQ
Apache Kafka
Frontend
GraphQL
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Twelve-Factor App
Incident Management
Management
Scrum
Apply
In office • Full-Time • 10+ years exp
C#
C#
.NET
Databases
MySQL
PostgreSQL
DynamoDB
RabbitMQ
Apache Kafka
DevOps
gRPC
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Twelve-Factor App
Incident Management
Management
Scrum
Apply
In office • Contractor • 4+ years exp • Bachelor's Degree
Java
C#
Databases
MySQL
PostgreSQL
SQLite
Apache Kafka
DevOps
Terraform
Ansible
GCP
Red Hat
Helm
GitHub Actions
Istio
Loki
AWS CDK
Prometheus
Azure
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Platform Engineering
GitHub
Management
Agile
Scrum
Apply
In office • Contractor • 5+ years exp • Bachelor's Degree
Python
Go
Java
AI/ML
LLM
LLM Guardrails
DevOps
Terraform
GCP
CloudFormation
Azure
CI/CD
AWS
Cybersecurity
Crowdstrike
SonarQube
ISO 27001
OWASP Top 10
SOC 2
Apply
In office • Contractor • 5+ years exp
JavaScript
TypeScript
Node JS
Node JS
Nest.JS
Fastify
Databases
MySQL
PostgreSQL
Redis
Frontend
GraphQL
DevOps
GCP
CI/CD
Git
AWS
Docker
Kubernetes
Management
Agile
Apply
See all jobs
This is one of many
588,947 more open roles from verified company boards, updated every day.