368,910open jobs
9,449companies
47,822added this week
Browse all
Salary
$25k – $63k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Skan AI uses computer vision on employee desktops to build a picture of business processes without integrating into each system. Founded in 2015 in California, it produces process maps and time analytics that reveal where work stalls. Insurance and banking operations teams are its main users.

Be at the Forefront of the Agentic AI Revolution

At Skan AI, you'll be part of the team pioneering the context engine for human and agentic execution, bringing context from enterprise operators, systems, and processes to power how the world's largest organizations execute their most complex, mission-critical work.

Why Join Skan AI

We're in hyper-growth mode at exactly the right moment in history. As enterprises race to adopt agentic AI, we're uniquely positioned to deliver the clear signal they desperately need: a platform that trains and grounds AI Agents in trillions of real execution signals, enabling reliable, compliant automation of their most complex processes.

Backed by Dell Technologies Capital and other leading investors, we're the only company that can bridge the gap between AI's promise and enterprise reality, making us perfectly positioned to define the agentic era for modern enterprises.

Our diverse, collaborative team of 250+ innovators is solving category-defining challenges at the intersection of AI, process intelligence, and enterprise work. Diverse perspectives fuel breakthrough thinking, cross-functional collaboration is the norm, and our work directly transforms how Fortune 500 companies operate. We are shaping the future of work itself.

The Site Reliability Engineer (SRE) owns the reliability, availability, and operational health of Skan's cloud-hosted customer environments. Sitting within the Customer Cloud Ops team, this role bridges software and infrastructure - applying engineering discipline to automate toil, respond to incidents rapidly, and uphold the SLA commitments that protect Skan's enterprise relationships.

The SRE is accountable for keeping production customer environments running at the performance and availability standards enterprise clients expect. This means being on-call, being proactive about operational risk, and continuously eliminating the manual work that gets in the way of reliable operations.

WHY THIS ROLE EXISTS

As Skan's cloud-hosted customer base grows, the complexity and volume of environment management, incident response, and reliability engineering grows with it. Without dedicated SRE capability, operational toil accumulates, incidents take longer to resolve, and SLA commitments are at risk - damaging customer trust and creating costly escalations.

The SRE function applies software engineering practices to infrastructure and operations - replacing manual, reactive processes with automation, runbooks, and agentic workflows that scale.

KEY RESPONSIBILITIES

  • Monitor platform health and performance across all cloud-hosted customer environments using observability tooling (Prometheus, Grafana, Datadog, or equivalent)
  • Respond to and own P1/P2 incidents - lead triage, diagnosis, and resolution; drive MTTR reduction through structured post-incident review
  • Perform ongoing capacity and environment planning for customer cloud deployments - anticipating growth and preventing resource-related incidents
  • Design and implement SRE automation to eliminate repetitive operational toil - agentic alerting, auto-remediation scripts, and automated runbook execution
  • Manage change and release events that affect production customer environments - coordinating with DevOps and product teams to minimize risk
  • Maintain and improve runbooks for all known failure patterns and operational procedures
  • Contribute to HA/DR playbook validation
  • Participate in on-call rotation and respond to alerts within defined SLA windows
  • Track and report on SLO/SLA adherence - contributing to monthly operational reports and QBR data
  • Identify and escalate environment risks proactively - before they become customer-facing incidents
  • Collaborate with the Automation Engineering team to develop agentic workflows that automate triage, routing, and remediation

KEY DELIVERABLES

  • Incident response records and post-mortems for all P1/P2 events - with root cause, remediation steps, and prevention actions documented
  • SLA/SLO dashboards: availability and performance reports for all cloud customer environments - updated continuously
  • Capacity and environment sizing plans per customer - reviewed quarterly or on significant usage change
  • Automated toil-reduction scripts and agentic remediation workflows - measurable reduction in manual operational hours
  • Runbooks for all known failure patterns - maintained and validated against real incidents
  • Change management records for all production environment events
  • Target: ≥ 99.9% uptime across all cloud-hosted customer environments

KEY SKILLS & QUALIFICATIONS

  • 5-8 years of experience as a SRE engineer
  • Cloud platforms: AWS, Azure, or GCP - environment management, networking, IAM, and observability at scale
  • Observability and monitoring: Prometheus, Grafana, Datadog, or equivalent - building dashboards, alerts, and SLO tracking
  • Infrastructure as Code: Terraform, Ansible, or Pulumi - provisioning and configuration management
  • Incident management: on-call discipline, structured MTTR mindset, post-mortem culture, and blameless review practices
  • Scripting and automation: Python, Bash - automation of operational tasks and agentic workflow development
  • Linux systems administration: process management, log analysis, performance tuning
  • Container orchestration: Kubernetes and Docker - deployment management and debugging in production
  • SRE fundamentals: SLI/SLO/SLA definition, error budget management, toil measurement and reduction
  • Strong written documentation skills - clear, evidence-based runbooks and incident reports

Skan AI is an equal opportunity employer committed to building a diverse, inclusive, and respectful workplace around the world. We do not discriminate based on race, color, religion or belief, sex (including pregnancy, sexual orientation, gender identity, or gender expression), national origin, ancestry, age, disability, medical condition, genetic information, marital or family status, military or veteran status, or any other characteristic protected by applicable laws in the locations where we operate.

We welcome people from all backgrounds and provide reasonable accommodations throughout the hiring process.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,910 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
SOC Senior Analyst 2 days ago
In office • Full-Time • Bachelor's Degree • Riyadh
DevOps
AWS
Azure
GCP
Cybersecurity
MITRE ATT&CK
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $116k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$61k – $200k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
Solution Architect 3 days ago
$31k – $75k per year (Estimated) • In office • Full-Time
AI/ML
AI Agents
Hallucination
Robotics
Digital Twin
Apply
$31k – $76k per year (Estimated) • In office • Full-Time
DevOps
CI/CD
Incident Management
SLI/SLO/SLA
Splunk
Cybersecurity
HashiCorp Vault
Cryptography
Vault
Apply
Data Analyst III 25 days ago
$17k – $43k per year (Estimated) • In office • Full-Time • 10+ years exp
Python
SQL
Databases
Databricks
Snowflake
AI/ML
AI Agents
ChatGPT
Claude
Analytics
Power BI
Apply
$165k – $333k per year (Estimated) • In office • Full-Time • 10+ years exp • Menlo Park
AI/ML
AI Agents
Apply
$31k – $60k per year (Estimated) • In office • Full-Time • Bengaluru
Python
SQL
Databases
Databricks
Delta Lake
PostgreSQL
AI/ML
AI Agents
Spark
DevOps
Ansible
CI/CD
Docker
Kubernetes
Platform Engineering
Terraform
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
See all jobs
This is one of many
368,910 more open roles from verified company boards, updated every day.