{"id":1830767,"url":"https://alion.io/job/cover-genius-senior-site-reliability-engineer-platform-engineeringdevops","title":"Senior Site Reliability Engineer (Platform Engineering/DevOps)","company":{"id":174848,"name":"Cover Genius","domain":"covergenius.com","url":"https://alion.io/company/cover-genius","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Kula","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Sydney, Australia"],"countries":["AU"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":72000,"max_usd":184000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":2444},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"AWS","optional":false},{"name":"Bash","optional":false},{"name":"CI/CD","optional":false},{"name":"Claude Code","optional":false},{"name":"Cursor","optional":false},{"name":"Datadog","optional":false},{"name":"Docker","optional":false},{"name":"ElasticSearch","optional":false},{"name":"FinOps","optional":false},{"name":"GCP","optional":false},{"name":"Gemini","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"Terraform","optional":false},{"name":"LLM","optional":true},{"name":"Stripe","optional":true}],"status":"live","first_seen_at":"2026-10-03T12:09:34Z","employer_posted_date":"2026-10-03","last_verified_at":"2026-10-09T00:27:26Z","board_verified":true,"closed_at":null,"days_open":5,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":5},"description":"About the Company\nCover Genius is the global infrastructure for embedded protection. Active in over 60 countries and all 50 US States, we protect the customers of the world’s largest digital companies, including Klarna, Revolut, Stripe, Priceline, Agoda, Booking.com, Turkish Airlines, Tongcheng Travel, eBay, and Uber, with seamless, end-to-end experiences. Cover Genius has protected more than 73M customers globally across 240M policies with USD $3.2BN in gross written sales.\nComing off a stellar year with 40% YoY revenue growth and a recent $100M capital raise, putting our valuation at $1.9BN, we are accelerating into our next phase of growth. As part of our team, you’ll help drive our AI-first roadmap, developing hyper-personalization engines, agentic distribution, and automated claims infrastructure, while building the scalable technology powering the fast-growing $70B embedded protection market.\nOur people are: Accountable, customer-obsessed, collaborative, driven\nOur people are not: Passive, defensive, siloed, hesitant\nAbout the Role\nAs a Senior Site Reliability Engineer, you'll own reliability and infrastructure strategy across the organisation - not just for a single team. Decisions you make on system design, tooling, and process will directly affect every engineering team's ability to operate and ship at scale.\nTo drive success in this role, you will have a strong background in cloud infrastructure and platform engineering, with experience across infrastructure-as-code, CI/CD and release automation, observability, security, and disaster recovery. You should possess strong technical judgement, the ability to set standards other engineers adopt, and a proactive approach to eliminating operational risk before it becomes a problem.\nRegular collaboration with software engineering teams, security teams, and other relevant stakeholders will be key in ensuring the reliability and efficiency of our production systems are achieved.\nKey Responsibilities\nDrive infrastructure and reliability strategy that connects directly to business outcomes - revenue protection, customer trust, and developer velocity\n\nAnalyze, test, and evolve systems to improve reliability and performance at an architectural/infrastructure level, setting technical direction other engineers build against\n\nArchitect multi-region, multi-AZ infrastructure with clear failover and disaster recovery strategies, applying deep AWS and GCP expertise to govern cloud infrastructure across multiple teams and projects\n\nDefine observability strategy and standards, and develop the tooling and dashboards other teams build on\n\nDefine and own SLOs and error budgets for services in your area, and use them to prioritise reliability work against feature velocity\n\nTake a leading role in major incidents and lead troubleshooting on the most complex production issues, driving deployment safety and process improvements while building a strong postmortem and continuous-improvement culture across the organisation\n\nReduce operational toil by building automation and self-service platforms, rather than absorbing repetitive work yourself\n\nDevelop and maintain design, troubleshooting, and runbook standards that other engineers can follow without tribal knowledge\n\nMentor other engineers and raise the bar on production ownership, testing, and code review across teams\n\nApply AI-assisted development to infrastructure problems, and help build the tooling and practices that make the wider team more effective with it\n\nDrive cloud cost optimisation at the organisational level - reserved capacity, right-sizing, FinOps practices\n\nSkills & Experience\nWhat you will bring:\n5+ years of experience in SRE, Platform Engineering, DevOps or other related roles\n\nDeep understanding of SRE and platform engineering principles, with a track record of setting them as team or organisational standards\n\nExtensive experience using, configuring, and setting standards for modern observability tools such as Datadog, Elasticsearch, Prometheus, Grafana\n\nExpert-level experience with cloud native and container technology such as Docker, and hands-on experience designing and managing Kubernetes clusters at scale\n\nDeep experience defining infrastructure-as-code standards and module libraries using tools such as Terraform\n\nComfortable scripting and developing internal tooling with Bash and at least one programming language (e.g. Python, Go)\n\nFluent with AI-driven development environments like Cursor, Claude Code, or Gemini, with a proven ability to leverage these tools within production engineering workflows\n\nExperience working with Linux\n\nStrong understanding of networking, distributed systems, and system architecture at scale\n\nProven experience deploying, scaling, and monitoring web applications and databases across multi-region or high-availability environments\n\nExpert-level knowledge of AWS and/or GCP platforms, with experience driving cloud cost optimisation and platform decisions at an organisational level\n\nBachelor's degree in Computer Science/Engineering, a postgraduate degree and/or record of academic achievement is also desirable\n\nWhat you will have:\nOwnership & Delivery\nTakes ambiguous problems and drives them to shipped outcomes - not just code, but results, and takes accountability even without a clear owner\n\nBalances speed with quality - knows when to iterate fast and when to invest in durability\n\nManages risk proactively - identifies failure modes and mitigates before they bite\n\nCommunication & Influence\nCreates clarity from ambiguity; documents decisions so others can build on your work\n\nInfluences through evidence and collaboration, not authority - mentors and unblocks teammates\n\nCommunicates technical concepts clearly to engineers, product, and business stakeholders\n\nAI-First Mindset\nTreats AI tools as essential infrastructure, not optional add-ons - continuously experiments with new capabilities\n\nUnderstands LLM strengths and limitations - knows when to prompt and when to build differently\n\nThinks in leverage: automates the repetitive, focuses human attention on judgement calls\n\nHelps others adopt AI workflows, shares what works, and raises the floor for the whole team\n\nWhy Cover Genius?\nAt Cover Genius, we create magic by turning the archaic into the extraordinary. We take one of the world’s oldest and most complex industries and reinvent it with world-class technology. Cover Genius doesn’t just disrupt legacy insurance; we make the impossible feel effortless.Our operating principles guide our mission:\nMake today matter: We deliver with urgency and excellence. We act decisively, move with intention, and hold a high bar.\n\nAct with accountability: We own our commitments and take pride in delivering results that move us forward.\n\nGrow together: We are a collective of curious minds. We learn from wins and setbacks, share knowledge generously, and elevate each other every day.\n\nInspire Each Other: We push each other to think bigger and pull together to go further.\n\nChampion Our Customers: We lead with empathy and center the customer, turning complex disruptions into seamless moments of trust.\n\nFlexible Work Environment - Our teams are hybrid. We work from home on a Wednesday and Thursday and attend the office on Monday, Tuesday and Friday with flexibility around start/finish times.\n\nDon't just take our word for it - hear about our culture straight from our people here: story time.\n\nReady to make an impact? If you’re looking for a place where you’ll be challenged, trusted, and empowered, we’d love to meet you.\nThe Legal & Privacy Stuff\nCover Genius promotes diversity and inclusivity. We don't tolerate discrimination, demeaning treatment of anyone, or harassment due to race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or any other legally protected status.\nBy submitting your application, you acknowledge that we may collect, store, and process your personal data for recruitment purposes. To ensure a fair evaluation, we may use AI to assist in sorting applications, but all final decisions are made by our hiring team and no candidate dispositions are automated. We will keep your information on file for three years from the date of your application. For detailed information about how we handle your data and our use of AI, please review our full Privacy Policy.","description_format":"text","description_chars":8381,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Flexible schedule"],"hiring_locations":[{"name":"Australia","iso":"AU","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["InsurTech"],"lifecycle":[{"event":"open","at":"2026-10-04T12:09:34Z"}],"visa":[],"liveness":{"score":17,"band":"cold","label":"Long shot","p_open":1,"p_active":0.462,"p_room":0.368,"age_days":4,"expected_fill_days":2,"reasons":["conf:2","velocity","win:tail"],"computed_at":"2026-10-08T05:49:30Z"},"pay":null,"html_url":"https://alion.io/job/cover-genius-senior-site-reliability-engineer-platform-engineeringdevops","json_url":"https://alion.io/job/cover-genius-senior-site-reliability-engineer-platform-engineeringdevops.json","meta":{"generated_at":"2026-10-09T00:37:20Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":850,"day_limit":5000,"remaining_today":4150,"minute_limit":60,"resets_at":"2026-10-10T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":174848},"rest":"https://alion.io/mcp/rest/get_company?id=174848"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fcover-genius-senior-site-reliability-engineer-platform-engineeringdevops"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fcover-genius-senior-site-reliability-engineer-platform-engineeringdevops"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fcover-genius-senior-site-reliability-engineer-platform-engineeringdevops"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/cover-genius-senior-site-reliability-engineer-platform-engineeringdevops\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fcover-genius-senior-site-reliability-engineer-platform-engineeringdevops"}]}