368,530open jobs
9,432companies
50,439added this week
Browse all
Location
Remote (Egypt)
Seniority
Senior · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
MOZN is an enterprise AI company that has helped 100+ organizations make critical and informed decisions through specialized AI, in two key areas: Financial Crime Prevention and Enterprise Knowledge Intelligence

About Mozn

MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence.

We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built on the relentless pursuit of excellence and meaningful impact.

If you’re passionate about working alongside exceptional talent on world-class AI, and you want the autonomy and runway to do the best work of your career, join us in shaping the future of intelligent enterprises.

About the role

We're hiring an AI Senior SRE: someone who carries a normal SRE workload - on-call rotation, incident response, root cause analysis, hands-on Kubernetes/cloud work - and builds the agentic layer that automates that workload over time. You are not exempt from operating production. You're the person best placed to know what should be automated, because you're the one doing it. This role exists because most reliability toil (triage, root-causing, remediation, SLO tracking, onboarding checks) is repetitive and well-defined enough to hand to an LLM-based agent with the right guardrails. Your job is to do the on-call/ops work like any SRE, then turn what you learn into agents - identify the workflow, build the agent, and earn the trust to let it act with increasing autonomy.

What you'll do

  • Carry a normal on-call rotation and act as a hands-on responder: investigate, fix, and document incidents yourself, exactly like any SRE on the team - especially for workflows that don' t have an agent yet.
  • Go deep on application-level reliability, not just infrastructure: read and debug service code, understand business logic well enough to find the real root cause, and ship fixes or PRs directly into application repos when the fix belongs in the app, not the platform.
  • Design, build, and ship LLM-based agents (using tools like Claude Code, OpenAI Codex, or Kimi K2/K3) that plug into our existing stack: Kubernetes, cloud APIs, Prometheus/Grafana/ELK/Datadog, PagerDuty/Slack.
  • Define the "tool" interface each agent needs - the specific APIs, scripts, and read/write actions - and build safe wrappers around them.
  • Set clear guardrails for every agent: what it may do autonomously vs. what it must propose for human approval, with a bias toward human-in-the-loop until an agent has earned trust.
  • Own agent evaluation - define what "correct" and "safe" look like per agent, and build test/backtest suites against real historical incidents.
  • Continuously tune prompts, context, and tool schemas as an agent' s scope grows.
  • Partner with the SRE/platform team to find good agent candidates: repetitive, well-scoped, auditable workflows.
  • Report on agent impact - MTTD/MTTR/MTTX movement, false positive/negative rates, and engineer-hours of toil removed.
  • Keep a security- and compliance-first posture: audit trails for every autonomous action, least-privilege access to production, and alignment with Saudi data residency/regulatory requirements.

Requirements

  • 3+ years building production software with LLMs - agentic workflows, tool/function calling, multi-step planning, RAG - not just personal use of a chat assistant.
  • Hands-on experience shipping real work with an agentic coding tool such as Claude Code, OpenAI Codex, or Kimi K2/K3.
  • Strong Python (or similar) for building agent tooling, API wrappers, and orchestration.
  • Real, hands-on SRE experience: comfortable being a primary on-call responder, running incident response, and doing root cause analysis under pressure - not just familiar with the concepts.
  • Application-level debugging skill, not just infra: able to read a service' s codebase, trace a failure back to the actual line/logic causing it, and ship a fix yourself - SRE work here isn' t limited to restarting pods or scaling nodes.
  • Solid hands-on Kubernetes and cloud provider experience (AWS/GCP/OCI/Azure) and fluency with observability tools (Prometheus, Grafana, Datadog, ELK) - both as an operator and as integration points for agents.
  • Understands guardrails for autonomous systems: permissioning, approval gates, rollback paths, auditability.
  • Can build trust with technical stakeholders - every agent starts with limited autonomy and has to earn more.

Nice to have

  • Experience in Saudi Arabia / MENA, ideally consulting across public or private sector clients.
  • Familiarity with Terraform/Ansible, Docker, VM/on-prem setups - useful context for the agents you' ll build, not the core job.
  • Experience building eval/backtest harnesses for LLM agents against historical incident data.
  • Background in ML engineering, LLMOps, or platform engineering.

Benefits

  • You will be at the forefront of an exciting time for the Middle East, joining a high-growth rocket-ship in an exciting space
  • You will be given a lot of responsibility and trust. We believe that the best results come when the people responsible for a function are given the freedom to do what they think is best
  • The fundamentals will be taken care of: competitive compensation, top-tier health insurance, and an enabling culture so that you can focus on what you do best
  • You will enjoy a fun and dynamic workplace working alongside some of the greatest minds in AI
  • We believe strength lies in difference, embracing all for who they are and empowered to be the best version of themselves
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Cairo
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$230k – $260k per year • Equity • Remote • Internship • Bachelor's Degree
Python
DevOps
Amazon EKS
AWS
Azure
CI/CD
GCP
Helm
Kubernetes
Terraform
Cybersecurity
FedRAMP
Orca Security
Apply
$100k – $200k per year • Equity 0.5–5% • In office • Full-Time • 1+ year exp • New York
Python
TypeScript
JavaScript
Python
FastAPI
Databases
DynamoDB
PostgreSQL
AI/ML
Claude
LLM
OpenAI
AI Agents
Frontend
Next.js
Tailwind CSS
React.js
DevOps
AWS
Docker
Vercel
GitHub
Management
Slack
Apply
$230k – $300k per year • Equity 0.1–0.2% • In office • Full-Time • 3+ years exp • PhD • New York
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
Devin
OpenAI Codex
Frontend
Next.js
tRPC
DevOps
Platform Engineering
Vercel
Apply
Engineering Manager 5 days ago
Remote • 10+ years exp
SQL
Databases
Cassandra
HBase
Apply
Engineering Manager 5 days ago
In office • 10+ years exp • Riyadh
SQL
Databases
Cassandra
HBase
Apply
CRM Support Analyst 29 days ago
In office • Full-Time • 1+ year exp • Riyadh
Marketing
Salesforce
Apply
Senior AI Engineer 1 month ago
In office • Full-Time • 5+ years exp • Bachelor's Degree • Riyadh
JavaScript
Node JS
Python
TypeScript
Python
Django
FastAPI
Flask
AI/ML
LLM
MLFlow
Ollama
ONNX
PyTorch
TensorFlow
Vertex AI
Amazon SageMaker
Frontend
Next.js
React.js
DevOps
Docker
Kubernetes
Rest API
Apply
Data Scientist III 1 month ago
In office • Full-Time • 3+ years exp • Bachelor's Degree • Riyadh
Python
SQL
AI/ML
PyTorch
Spark
DevOps
Docker
Git
GitHub
Apply
In office • Full-Time • Cairo
Apply
Apply
Remote • Full-Time • 3+ years exp • Cairo
C#
JavaScript
TypeScript
C#
.NET
AI/ML
Copilot
DevOps
Azure
Azure DevOps
CI/CD
Rest API
Analytics
ETL/ELT
Management
Power Apps
Power Automate
Apply
Remote • Full-Time • 3+ years exp • Cairo
C#
JavaScript
TypeScript
C#
.NET
AI/ML
Copilot
DevOps
Azure
Azure DevOps
CI/CD
Rest API
Analytics
ETL/ELT
Management
Power Apps
Power Automate
Apply
Remote • Full-Time • 1+ year exp • Cairo
C#
JavaScript
TypeScript
C#
.NET
AI/ML
Copilot
DevOps
Azure
Azure DevOps
CI/CD
Rest API
Analytics
ETL/ELT
Management
Power Apps
Power Automate
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.