823,562open jobs
53,068companies
134,459added this week
Browse all
Salary
$71k – $95k per year
Location
Hybrid (Bengaluru, India)
Seniority
Principal · 14+ years exp

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Aug 24, 2026.

Overview
Company
Impact
Profile match
Omnidian is a leading provider of AI-powered solutions for the financial services industry. We help financial institutions improve their operational efficiency, reduce costs, and enhance customer experience. Our products and services include: Fraud Detection: Identify and prevent fraudulent activity.

The Job

As the Principal SRE / DevOps Engineer, you will play a foundational role in architecting the future of our platform reliability and operational ecosystem, serving as a technical lead and strategist as we build a robust, scalable, and highly supportable infrastructure to support our clients and products.

This is a hands-on technical position with some project management and leadership responsibilities. In this role, you will work closely with our Software Product, Engineering, and Operations teams to document and evangelize a vision for platform reliability, observability, and automation. You will guide the team to break the high-level vision into well-defined milestones. You will assist in establishing and will champion and participate in best practices in workload and reliability management, including providing visibility to stakeholders and executives for progress toward our vision.

What You’ll Do

    At Omnidian we believe in trust and autonomy. How you create an impact is ultimately up to you. Here is an outline of the balance of responsibilities:

    Architecture Strategy and Vision (35%)

  • Formulate the multi-year technical roadmap for our platform reliability, infrastructure, and operational excellence, identifying where intelligent automation can significantly remove engineering and operational bottlenecks (e.g., routine toil, deployment friction, scaling overhead, and incident response).
  • Obtain executive and stakeholder buy-in for a documented vision for our SRE / DevOps deliverables. Communicate changes to and/or progress against the vision to executive and technical audiences.
  • Translate the technical vision to actionable and trackable work plans, including meaningful milestones. Work with Software Product, Go-To-Market, and other partner teams and stakeholders to align on appropriate milestones and timelines.
  • Translate business and reliability requirements to right-sized technical specifications; mentor team members to create technical documentation and clear success criteria (including SLIs/SLOs and error budgets).
  • Maintain a security-first mindset, ensuring all infrastructure, automation frameworks, and CI/CD pipelines are robustly defended against operational and security vulnerabilities.
  • Provide strong technical leadership and mentorship across teams, establishing clear, practical departmental standards for the safe and ethical use of automation and AI coding/ops assistants.
  • Develop an appropriate reliability and testing strategy (leveraging automation and AI where effective for chaos testing, integration validation, and coverage), ensuring LOE estimates include right-sized reliability and observability work.
  • SRE / DevOps Engineering (35%)

  • Design, implement, and maintain production-grade Kubernetes platforms, cluster management, and related orchestration.
  • Build and evolve Infrastructure as Code and configuration management using Ansible and Terraform for consistent, repeatable environments.
  • Own and continuously improve CI/CD pipelines, deployment strategies, and release automation to enable safe, frequent, and reliable software delivery.
  • Implement and refine observability stacks centered on Grafana (along with metrics, logging, and tracing systems) to provide actionable visibility into system health, performance, and reliability.
  • Create and maintain automation for operational toil reduction, self-healing systems, and infrastructure provisioning as needed.
  • Use AI coding and ops assistants daily for scripting, infrastructure code, refactoring, pipeline improvements, and reliability testing.
  • Project Management (25%)

  • Follow established agile ceremonies, including best practice metrics that are reviewed with the team to support continuous improvement (DORA metrics, SLO attainment, error budget burn, incident metrics, cycle time, technical debt, etc.).
  • Create and deliver on visible project plans, break down milestones into distinct work, plan and assign work, and ensure timely and accurate delivery.
  • Clearly and proactively communicate progress against the plan including status, blockers, dependencies, and risks.
  • Clear blockers to timely or accurate delivery; escalate to leadership as appropriate.
  • Identify and proactively communicate work needed from outside the team, obtain commitment, and follow through to ensure dependencies will be delivered to plan.
  • Set appropriate documentation expectations for the team (runbooks, architecture decision records, post-incident reviews); ensure this is included in LOE estimates.
  • Continuous Improvement (5%)

  • Identify, align stakeholders on, and implement improvements to our processes, codebases, and architecture.

Who You Are

  • You are an exceptional SRE / DevOps engineer who is passionate about reliability, observability, clean automation, robust infrastructure architecture, and production stability. You view generative AI as a pragmatic utility to accelerate engineering and operational velocity.
  • You understand that technology best practices and patterns are always shifting, and you love evaluating whether emerging frameworks, tools, or practices should evolve our roadmap.
  • You have experience delivering flexible, highly available platforms that provide meaningful reliability and operational insights.
  • You possess a security-oriented mindset and deeply understand the security boundaries required when exposing infrastructure and automation tooling.
  • You treat everyone with empathy and respect.
  • You are a strong communicator, effectively clarifying technical needs, reliability trade-offs, and operational impacts for cross-functional and non-technical partners.
  • You excel at balancing rapid responses to real business and operational needs with foundational, long-term architectural and reliability stability.

Experience You'll Need

  • 14+ years of experience building, operating, and optimizing large-scale distributed systems, cloud infrastructure, and production platforms.
  • Deep hands-on expertise with Kubernetes (cluster architecture, networking, scaling, security, and day-2 operations).
  • Strong experience with Ansible (or equivalent configuration management) and Terraform (or equivalent Infrastructure as Code tools) for infrastructure automation, consistency, provisioning, and managing cloud and infrastructure resources.
  • Proven ownership of modern CI/CD pipelines, deployment strategies, and release engineering practices.
  • Extensive experience designing and operating observability platforms, with strong proficiency in Grafana (dashboards, alerting, and integration with metrics/logging/tracing systems).
  • Solid scripting and automation skills (Python, Bash, or equivalent) plus Infrastructure as Code practices.
  • Demonstrated ability to establish and drive SRE practices including SLIs/SLOs, error budgets, incident management, and toil reduction.
  • Experience using advanced coding/ops assistants extensively for day-to-day automation, infrastructure code, testing, and the ability to establish team standards around their safe and effective use.
  • Experience designing systems with appropriate human oversight for automated or agentic operational workflows.

Experience Thats a Plus

  • 2+ years of hands-on experience integrating generative AI/LLM components or advanced automation frameworks into production infrastructure or operational systems.
  • Familiarity with advanced Kubernetes operators, service meshes, or multi-cluster management.
  • Experience implementing tracing frameworks and advanced observability for complex distributed systems.
  • Knowledge of event-driven architectures, time-series data, or high-cardinality metrics environments.
  • Exposure to IoT telemetry or similar high-volume data streams.
  • Solar industry experience

Work-Life & Culture

  • We offer a competitive total compensation package that includes monthly health insurance premiums, bonuses and long-term stock options for every employee
  • We love to lift each other up through company-wide slack channels such as #puppiesandpets, #omnidian-wellness, #praiseandbooms and #sustainablefuture
  • We are a passionate, mission driven team that believes in collaboration, mutual respect and trust. For examples, comeDiscover our Story!

Grow With Us

  • We mentor and invest in our employees and prioritize them for future opportunities. Check out ourInstagram reels to see a few career journey examples
  • Internal candidates: Check out our advice on Internal Transfer: Job Application Process
  • We’re a fast-growing startup, which means we’re constantly reinventing processes, adding new products, and asking people to use all of their skills and talents. That means there’s gonna be a lot of opportunities for you to grow, which also means you will likely be stretched in ways you’ve never experienced in a job before. If you are resilient, determined, and not afraid of a big challenge, come apply.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
823,562 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Bengaluru
≈ $29k – $69k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Hyderabad
JavaScript
Node JS
Node JS
Commander.js
DevOps
Terraform
CloudFormation
Datadog
PagerDuty
Azure
CI/CD
AWS
Kubernetes
Grafana
Opsgenie
Amazon EC2
Incident Management
SLI/SLO/SLA
IAM
Amazon CloudWatch
TCP/IP
DNS
DHCP
VPN
Cybersecurity
SIEM
Management
Confluence
Jira
ITIL
Apply
≈ $40k – $90k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Bengaluru
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Quantization
TensorFlow
PyTorch
Edge AI
LiteRT
DevOps
GCP
RTOS
Azure
AWS
Wi-Fi
IoT
FreeRTOS
Matter
Apply
Devops Lead 9 days ago
≈ $35k – $82k per year (Estimated) • In office • Full-Time • Bengaluru
DevOps
Terraform
GCP
Helm
CloudFormation
Prometheus
CI/CD
AWS
Kubernetes
Grafana
QA
Sentry
Apply
≈ $34k – $81k per year (Estimated) • In office • Bengaluru
Python
Bash
AI/ML
LLM
DevOps
Ansible
GCP
Red Hat
Podman
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Linux
Windows
Cybersecurity
Active Directory
LDAP
Apply
Dev Ops Engineer Lead 2 months ago
≈ $34k – $81k per year (Estimated) • Hybrid • Full-Time • Bengaluru
Python
Databases
PostgreSQL
AI/ML
AI Agents
Langfuse
LLM
DevOps
Terraform
Helm
Azure DevOps
GitHub Actions
Istio
OpenTelemetry
Prometheus
Azure
CI/CD
GitOps
ArgoCD
Git
Docker
Kubernetes
Grafana
Blue-Green Deployment
Service Mesh
Configuration Management
Azure AKS
Progressive Delivery
Cybersecurity
Snyk
Checkmarx
Apply
$64k – $85k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Toronto
PowerShell
Databases
Databricks
AI/ML
Machine Learning
DevOps
Terraform
Azure DevOps
Azure
CI/CD
Git
Platform Engineering
Management
Agile
Apply
≈ $33k – $77k per year (Estimated) • In office • Poland
Python
Bash
DevOps
RTOS
CI/CD
Git
Linux
Apply
≈ $15k – $29k per year (Estimated) • Remote (likely Bosnia and Herzegovina) • Internship • Sarajevo
Python
JavaScript
Java
TypeScript
Node JS
Bash
Java
Maven
Spring Boot
Node JS
Axios
Databases
PostgreSQL
AI/ML
NLP
Tokenization
Sentiment Analysis
Interpretability
Machine Learning
Frontend
Angular
React.js
Mobile
JUnit
DevOps
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Nginx
Linux
Design
Figma
Management
Trello
Jira
Agile
Scrum
QA
Playwright
Postman
Pytest
Apply
≈ $50k – $112k per year (Estimated) • In office • Contractor • Montreal
Python
SQL
AI/ML
Copilot
Analytics
Power BI
Microsoft Excel
Management
Power Automate
SharePoint
Apply
≈ $50k – $126k per year (Estimated) • In office • Texcoco
DevOps
CI/CD
Apply
$140k – $165k per year • Equity • Hybrid • Full-Time • Seattle
SQL
Analytics
Microsoft Excel
Management
Google Sheets
Marketing
YouTube
Instagram
Apply
$80k – $110k per year • Remote (United States) • Contractor • 2+ years exp
Apply
Hybrid • Full-Time • Costa Rica
SQL
Analytics
Microsoft Excel
Management
Google Sheets
Marketing
Instagram
Apply
$26k – $43k per year • Hybrid • 7+ years exp • Bengaluru
Python
Java
TypeScript
SQL
Java
Spring Boot
Databases
Databricks
Delta Lake
AI/ML
Claude
Time Series Forecasting
DevOps
Kubernetes
Management
Agile
Apply
$18k – $33k per year • Hybrid • 2+ years exp • Bengaluru
Python
Java
TypeScript
SQL
Java
Spring Boot
Databases
Databricks
Delta Lake
AI/ML
Copilot
Claude
Spark
DevOps
CI/CD
Management
Agile
Apply
≈ $35k – $76k per year (Estimated) • In office • Full-Time • Bengaluru
Python
AI/ML
Reinforcement Learning
AI Agents
Agentic Workflows
Machine Learning
DevOps
CI/CD
Apply
≈ $28k – $79k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
DevOps
Azure
AWS
Cloudflare
SLI/SLO/SLA
DNS
DHCP
VPN
BGP
OSPF
Cybersecurity
Zero Trust
Management
ITIL
Apply
In office • Full-Time • Bengaluru
Databases
SAP HANA
Management
ITIL
Apply
≈ $30k – $80k per year (Estimated) • In office • Bengaluru
Apply
≈ $26k – $58k per year (Estimated) • In office • 7+ years exp • Master's Degree • Bengaluru
Apply
See all jobs
This is one of many
823,562 more open roles from verified company boards, updated every day.