{"id":1597688,"url":"https://alion.io/job/okta-principal-site-reliability-engineer","title":"Principal Site Reliability Engineer","company":{"id":4990,"name":"Okta","domain":"okta.com","url":"https://alion.io/company/okta","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"A","score":88,"open_postings":84,"ghost_share":0,"stale_share":0.488,"repost_share":0,"time_to_fill_p50_days":56,"computed_at":"2026-10-06T05:45:30Z"}},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":36000,"max_usd":93000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":12},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"AWS","optional":false},{"name":"CI/CD","optional":false},{"name":"DNS","optional":false},{"name":"GCP","optional":false},{"name":"GitOps","optional":false},{"name":"Go","optional":false},{"name":"Helm","optional":false},{"name":"IAM","optional":false},{"name":"Incident Management","optional":false},{"name":"Kubernetes","optional":false},{"name":"LLM Guardrails","optional":false},{"name":"MySQL","optional":false},{"name":"Okta","optional":false},{"name":"OpenSearch","optional":false},{"name":"Platform Engineering","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Python","optional":false},{"name":"Redis","optional":false},{"name":"Terraform","optional":false},{"name":"Agentic Workflows","optional":true},{"name":"Amazon EKS","optional":true},{"name":"ArgoCD","optional":true},{"name":"Datadog","optional":true},{"name":"Git","optional":true},{"name":"Google GKE","optional":true},{"name":"Machine Learning","optional":true},{"name":"Splunk","optional":true}],"status":"live","first_seen_at":"2026-06-15T05:08:24Z","employer_posted_date":"2026-09-21","last_verified_at":"2026-10-07T00:25:19Z","board_verified":true,"closed_at":null,"days_open":113,"trust":{"level":"ok","repost_count":0,"flags":[],"days_open":112},"description":"Secure Every Identity, from AI to Human\n Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.\nThis is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.\nGet to know Okta\nOkta is The World’s Identity Company. We free everyone to safely use any technology-anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth. \nAt Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box, we’re looking for lifelong learners and people who can make us better with their unique experiences. \nJoin our team! We’re building a world where Identity belongs to you. \nThe Engineering Opportunity\nWe are seeking a Principal Site Reliability Engineer to serve as a technical leader for reliability engineering within Okta's Emerging Products Group (EPG).\nThis role extends beyond operating production systems. You will define technical strategy, influence platform architecture, establish reliability standards, and lead transformational initiatives that improve scalability, resilience, security, and operational excellence for one of Okta's fastest-growing product areas.\nInitially, you will partner closely with the Spera / Identity Security Posture Management (ISPM) engineering organization to establish reliability strategy, operational excellence, and platform maturity. Over time, you will help drive broader reliability initiatives across EPG and contribute to the evolution of reliability engineering practices across multiple products including Workflows, IGA, PAM, and ISPM.\nYou will work closely with engineering leadership, product leadership, architects, and Staff engineers to shape the future of Okta's cloud infrastructure and reliability practices.\nThe ideal candidate combines deep technical expertise with strong organizational influence and has a proven track record of leading large-scale engineering initiatives that drive measurable business outcomes.\nWhat You'll Be Doing\nReliability Strategy & Architecture\nDefine and drive the reliability strategy for critical product and platform services.\nEstablish standards for availability, resilience, observability, incident management, and operational readiness.\nLead architecture reviews for critical services and platform initiatives.\nPartner with engineering leaders to ensure reliability objectives align with business priorities and customer expectations.\nCreate frameworks, standards, and operational guardrails that enable engineering teams to operate safely at scale.\nGuide service architecture toward simplicity, scalability, resilience, and operational excellence.\nDrive major initiatives that improve platform maturity and long-term sustainability.\nProduct & Platform Leadership\nOwn reliability architecture and operational excellence for the Spera / ISPM product area.\nCollaborate closely with engineering leadership to establish reliability objectives and technical roadmaps.\nLead large-scale scalability, resiliency, and performance initiatives.\nPartner with platform and product engineering teams to build self-service operational capabilities that improve developer productivity while strengthening reliability and security.\nInfluence technical direction through data-driven recommendations, engineering expertise, and collaborative leadership.\nSupport highly available, large-scale cloud environments as part of an on-call rotation.\nEngineering & Automation\nDesign, build, and operate large-scale cloud infrastructure and production services.\nDevelop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.\nEliminate operational toil through automation, tooling, and platform engineering.\nImprove deployment safety, operational workflows, and platform consistency through GitOps and Infrastructure-as-Code practices.\nCollaborate on modernizing existing workloads and aligning them with evolving platform capabilities.\nLead complex engineering initiatives from conception through production rollout and long-term operational ownership.\n\nTechnical Leadership\nMentor Staff and Senior engineers across multiple teams and organizations.\nLead technical reviews, design reviews, and operational readiness assessments.\nBuild engineering consensus across teams with differing priorities and objectives.\nHelp develop the next generation of technical leaders within Okta.\nDrive adoption of reliability engineering best practices across EPG.\nShare patterns, tooling, and operational practices across Workflows, Inbox, PAM, and ISPM teams.\nInfluence technical direction through expertise, collaboration, and execution rather than organizational authority.\nAI & Agentic Operations\nLead the exploration and adoption of AI-assisted reliability engineering practices across EPG.\nDesign and champion agentic systems that accelerate troubleshooting, incident response, root-cause analysis, and operational decision-making.\nEvaluate emerging AI technologies and identify practical opportunities to improve reliability engineering workflows.\nEstablish best practices for safe, effective, and measurable use of AI within production operations.\nDrive initiatives that reduce operational toil and improve engineering productivity through intelligent automation.\nOur Tech Stack\nInfrastructure/Orchestration: Kubernetes (EKS/GKE), Terraform, Helm, Git, ArgoCD, Gitops\nProgramming: Golang, Python\nObservability: Datadog, Splunk\nData Stores: PostgreSQL, Redis, OpenSearch\nWhat We Are Looking For\nTechnical Excellence\nExtensive experience designing and operating large-scale production systems in AWS and/or GCP.\nDeep expertise with Kubernetes in production environments.\nExperience designing reliability strategies for Kubernetes-based platforms.\nStrong expertise troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle challenges.\nExtensive experience with Infrastructure as Code technologies such as Terraform and Helm.\nStrong software engineering skills in Golang and/or Python.\nExperience building internal platforms, developer tooling, and operational automation.\nDeep understanding of distributed systems architecture and cloud-native application design.\nStrong understanding of cloud networking fundamentals including DNS, service discovery, ingress, load balancing, TLS, traffic management, and multi-region architectures.\nExperience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.\nExperience establishing observability standards, monitoring strategies, and operational best practices across engineering organizations.\nExperience with or strong interest in AI-assisted engineering and operational automation.\nReliability & Operational Excellence\nStrong expertise operating customer-facing production systems at scale.\nDeep understanding of reliability engineering principles including SLIs, SLOs, error budgets, capacity planning, and resilience engineering.\nExperience leading major incident response efforts and driving long-term operational improvements.\nStrong understanding of CI/CD, GitOps, deployment strategies, and automation-first operational practices.\nProven success driving large-scale reliability transformations and architectural modernization efforts.\nAbility to balance reliability, scalability, security, customer experience, and engineering velocity.\nSecurity & Compliance\nStrong understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design.\nExperience operating systems within security-sensitive or regulated environments is a plus.\nFamiliarity with operational controls, compliance requirements, and security best practices in cloud-native environments.\nLeadership & Influence\nDemonstrated success leading complex technical initiatives across multiple teams and organizations.\nDemonstrated ability to drive technical strategy and influence engineering outcomes across multiple organizations and leadership teams.\nProven ability to influence technical direction without direct organizational authority.\nExperience working effectively within globally distributed engineering organizations spanning multiple timezones and cultures.\nStrong collaboration, communication, and stakeholder management skills.\nExperience mentoring Staff engineers and helping develop future technical leaders.\nAbility to translate business objectives into technical strategy and measurable reliability outcomes.\nExperience working closely with engineering leadership, architects, and product leaders to drive organizational change.\nPreferred Qualifications\nExperience operating SaaS platforms serving millions of users.\nExperience supporting globally distributed production environments.\nExperience leading platform engineering or reliability transformation initiatives.\nExperience implementing AI-assisted operational tooling, agentic workflows, or intelligent automation platforms.\nPrior experience as a Staff, Principal, or equivalent senior technical leader in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or Cloud Operations.\n#P25309_3469886\nThe Okta Experience\nSupporting Your Well-Being \nDriving Social Impact \nDeveloping Talent and Fostering Connection + Community\nWe are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.\nOkta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.\nIf reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation.\nNotice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.","description_format":"text","description_chars":11103,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Cybersecurity","Identity Management","E-learning"],"lifecycle":[{"event":"open","at":"2026-10-01T18:15:48Z"}],"visa":[],"liveness":{"score":18,"band":"cold","label":"Long shot","p_open":1,"p_active":0.647,"p_room":0.28,"age_days":113,"expected_fill_days":56,"reasons":["conf:0","velocity","win:tail","crowd:brand"],"computed_at":"2026-10-06T05:45:30Z"},"pay":null,"html_url":"https://alion.io/job/okta-principal-site-reliability-engineer","json_url":"https://alion.io/job/okta-principal-site-reliability-engineer.json","meta":{"generated_at":"2026-10-07T01:14:41Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2051,"day_limit":5000,"remaining_today":2949,"minute_limit":60,"resets_at":"2026-10-08T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":4990},"rest":"https://alion.io/mcp/rest/get_company?id=4990"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fokta-principal-site-reliability-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fokta-principal-site-reliability-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fokta-principal-site-reliability-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/okta-principal-site-reliability-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fokta-principal-site-reliability-engineer"}]}