{"id":1264822,"url":"https://alion.io/job/omnidian-principal-sre-dev-ops-engineer","title":"Principal SRE / Dev Ops Engineer","company":{"id":3566,"name":"Omnidian","domain":"omnidian.com","url":"https://alion.io/company/omnidian","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Lever","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":null,"work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":71000,"max":95000,"currency":"USD","period":"year","gross":null,"usd_annual":95000},"salary_estimate":null,"experience_years_min":14,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"Agile","optional":false},{"name":"AI Agents","optional":false},{"name":"Ansible","optional":false},{"name":"CI/CD","optional":false},{"name":"Configuration Management","optional":false},{"name":"Error Budget","optional":false},{"name":"Grafana","optional":false},{"name":"Incident Management","optional":false},{"name":"Kubernetes","optional":false},{"name":"LLM","optional":false},{"name":"Python","optional":false},{"name":"Self-Healing","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"Terraform","optional":false},{"name":"Time Series Forecasting","optional":false}],"status":"live","first_seen_at":"2026-08-24T21:13:28Z","employer_posted_date":"2026-08-24","last_verified_at":"2026-09-29T02:54:38Z","board_verified":true,"closed_at":null,"days_open":35,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":35},"description":"The Job\nAs the Principal SRE / DevOps Engineer, you will play a foundational role in architecting the future of our platform reliability and operational ecosystem, serving as a technical lead and strategist as we build a robust, scalable, and highly supportable infrastructure to support our clients and products.\nThis is a hands-on technical position with some project management and leadership responsibilities. In this role, you will work closely with our Software Product, Engineering, and Operations teams to document and evangelize a vision for platform reliability, observability, and automation. You will guide the team to break the high-level vision into well-defined milestones. You will assist in establishing and will champion and participate in best practices in workload and reliability management, including providing visibility to stakeholders and executives for progress toward our vision.\nWhat You’ll Do\nAt Omnidian we believe in trust and autonomy. How you create an impact is ultimately up to you. Here is an outline of the balance of responsibilities:\nArchitecture Strategy and Vision (35%)\nFormulate the multi-year technical roadmap for our platform reliability, infrastructure, and operational excellence, identifying where intelligent automation can significantly remove engineering and operational bottlenecks (e.g., routine toil, deployment friction, scaling overhead, and incident response).\nObtain executive and stakeholder buy-in for a documented vision for our SRE / DevOps deliverables. Communicate changes to and/or progress against the vision to executive and technical audiences.\nTranslate the technical vision to actionable and trackable work plans, including meaningful milestones. Work with Software Product, Go-To-Market, and other partner teams and stakeholders to align on appropriate milestones and timelines.\nTranslate business and reliability requirements to right-sized technical specifications; mentor team members to create technical documentation and clear success criteria (including SLIs/SLOs and error budgets).\nMaintain a security-first mindset, ensuring all infrastructure, automation frameworks, and CI/CD pipelines are robustly defended against operational and security vulnerabilities.\nProvide strong technical leadership and mentorship across teams, establishing clear, practical departmental standards for the safe and ethical use of automation and AI coding/ops assistants.\nDevelop an appropriate reliability and testing strategy (leveraging automation and AI where effective for chaos testing, integration validation, and coverage), ensuring LOE estimates include right-sized reliability and observability work.\nSRE / DevOps Engineering (35%)\nDesign, implement, and maintain production-grade Kubernetes platforms, cluster management, and related orchestration.\nBuild and evolve Infrastructure as Code and configuration management using Ansible and Terraform for consistent, repeatable environments.\nOwn and continuously improve CI/CD pipelines, deployment strategies, and release automation to enable safe, frequent, and reliable software delivery.\nImplement and refine observability stacks centered on Grafana (along with metrics, logging, and tracing systems) to provide actionable visibility into system health, performance, and reliability.\nCreate and maintain automation for operational toil reduction, self-healing systems, and infrastructure provisioning as needed.\nUse AI coding and ops assistants daily for scripting, infrastructure code, refactoring, pipeline improvements, and reliability testing.\nProject Management (25%)\nFollow established agile ceremonies, including best practice metrics that are reviewed with the team to support continuous improvement (DORA metrics, SLO attainment, error budget burn, incident metrics, cycle time, technical debt, etc.).\nCreate and deliver on visible project plans, break down milestones into distinct work, plan and assign work, and ensure timely and accurate delivery.\nClearly and proactively communicate progress against the plan including status, blockers, dependencies, and risks.\nClear blockers to timely or accurate delivery; escalate to leadership as appropriate.\nIdentify and proactively communicate work needed from outside the team, obtain commitment, and follow through to ensure dependencies will be delivered to plan.\nSet appropriate documentation expectations for the team (runbooks, architecture decision records, post-incident reviews); ensure this is included in LOE estimates.\nContinuous Improvement (5%)\nIdentify, align stakeholders on, and implement improvements to our processes, codebases, and architecture.\n\nWho You Are\nYou are an exceptional SRE / DevOps engineer who is passionate about reliability, observability, clean automation, robust infrastructure architecture, and production stability. You view generative AI as a pragmatic utility to accelerate engineering and operational velocity.\nYou understand that technology best practices and patterns are always shifting, and you love evaluating whether emerging frameworks, tools, or practices should evolve our roadmap.\nYou have experience delivering flexible, highly available platforms that provide meaningful reliability and operational insights.\nYou possess a security-oriented mindset and deeply understand the security boundaries required when exposing infrastructure and automation tooling.\nYou treat everyone with empathy and respect.\nYou are a strong communicator, effectively clarifying technical needs, reliability trade-offs, and operational impacts for cross-functional and non-technical partners.\nYou excel at balancing rapid responses to real business and operational needs with foundational, long-term architectural and reliability stability.\nExperience You'll Need\n14+ years of experience building, operating, and optimizing large-scale distributed systems, cloud infrastructure, and production platforms.\nDeep hands-on expertise with Kubernetes (cluster architecture, networking, scaling, security, and day-2 operations).\nStrong experience with Ansible (or equivalent configuration management) and Terraform (or equivalent Infrastructure as Code tools) for infrastructure automation, consistency, provisioning, and managing cloud and infrastructure resources.\nProven ownership of modern CI/CD pipelines, deployment strategies, and release engineering practices.\nExtensive experience designing and operating observability platforms, with strong proficiency in Grafana (dashboards, alerting, and integration with metrics/logging/tracing systems).\nSolid scripting and automation skills (Python, Bash, or equivalent) plus Infrastructure as Code practices.\nDemonstrated ability to establish and drive SRE practices including SLIs/SLOs, error budgets, incident management, and toil reduction.\nExperience using advanced coding/ops assistants extensively for day-to-day automation, infrastructure code, testing, and the ability to establish team standards around their safe and effective use.\nExperience designing systems with appropriate human oversight for automated or agentic operational workflows.\nExperience Thats a Plus\n2+ years of hands-on experience integrating generative AI/LLM components or advanced automation frameworks into production infrastructure or operational systems.\nFamiliarity with advanced Kubernetes operators, service meshes, or multi-cluster management.\nExperience implementing tracing frameworks and advanced observability for complex distributed systems.\nKnowledge of event-driven architectures, time-series data, or high-cardinality metrics environments.\nExposure to IoT telemetry or similar high-volume data streams.\nSolar industry experience\nWork-Life & Culture\nWe offer a competitive total compensation package that includes monthly health insurance premiums, bonuses and long-term stock options for every employee\nWe love to lift each other up through company-wide slack channels such as #puppiesandpets, #omnidian-wellness, #praiseandbooms and #sustainablefuture\nWe are a passionate, mission driven team that believes in collaboration, mutual respect and trust. For examples, comeDiscover our Story!\nGrow With Us\nWe mentor and invest in our employees and prioritize them for future opportunities. Check out ourInstagram reels to see a few career journey examples\nInternal candidates: Check out our advice on Internal Transfer: Job Application Process\nWe’re a fast-growing startup, which means we’re constantly reinventing processes, adding new products, and asking people to use all of their skills and talents. That means there’s gonna be a lot of opportunities for you to grow, which also means you will likely be stretched in ways you’ve never experienced in a job before. If you are resilient, determined, and not afraid of a big challenge, come apply.","description_format":"text","description_chars":8791,"description_truncated":false,"requirements":{"experience_years_min":14,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Health insurance","Stock options"],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Solar Energy"],"lifecycle":[{"event":"open","at":"2026-09-25T22:04:58Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.842,"p_room":0.75,"age_days":34,"expected_fill_days":42,"reasons":["conf:3","win:late"],"computed_at":"2026-09-28T05:45:00Z"},"pay":{"stated_usd_annual":95000,"is_top_pay":false},"html_url":"https://alion.io/job/omnidian-principal-sre-dev-ops-engineer","json_url":"https://alion.io/job/omnidian-principal-sre-dev-ops-engineer.json","meta":{"generated_at":"2026-09-29T04:56:20Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4792,"day_limit":5000,"remaining_today":208,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}