{"id":1494494,"url":"https://alion.io/job/altera-aiops-observabilitysre-lead","title":"AIOPs Observability/SRE Lead","company":{"id":5328,"name":"Altera","domain":"altera.com","url":"https://alion.io/company/altera","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":91,"open_postings":16,"ghost_share":0,"stale_share":0.125,"repost_share":0.063,"time_to_fill_p50_days":80,"computed_at":"2026-10-01T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Jose, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":187000,"max":270700,"currency":"USD","period":"year","gross":null,"usd_annual":270700},"salary_estimate":null,"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AIOps","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"Incident Management","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"SRE","optional":false},{"name":"VPN","optional":false},{"name":"Chaos Engineering","optional":true},{"name":"HPC","optional":true}],"status":"live","first_seen_at":"2026-09-29T00:00:00Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-01T15:41:46Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"Job Details:\nJob Description:\nOnsite Requirement: This position requires regular in-office work and is an onsite role based in San Jose, CA. Candidates must be able to work onsite in San Jose, CA.\nAbout Altera\nAt Altera™, our independence as the world’s largest pure-play FPGA solutions provider gives us the focus, speed, and agility to innovate without compromise. With more than four decades of industry-leading FPGA expertise, our singular mission is to deliver high-performance, flexible FPGA solutions that enable customers to solve their most complex computing challenges.\nAs Altera continues to evolve and scale as an independent company, our IT organization is transforming its global infrastructure to support the needs of the business, engineering organizations, laboratories, and employees around the world.\nAbout the Role\nWe are seeking a Site Reliability Engineering Manager to lead the SRE function. This role builds and manages a team of reliability engineers who drive platform stability, observability, automation, and operational excellence across the semiconductor company's critical IT and engineering systems.\nIn this role, you will define and implement scalable, secure, and high-performance monitoring architectures that support both current and future business requirements. You will work closely with IT, Engineering, Integration teams, security teams, and external service providers to modernize network infrastructure, enable hybrid cloud connectivity, and ensure a smooth transition from existing environments to the future-state network.\nThe ideal candidate brings deep expertise in enterprise SRE architecture and transformation, strong hands-on knowledge of routing and switching technologies, and experience integrating cloud environments such as AWS and Azure with large-scale on-premises infrastructure.\nKey Responsibilities\nLead and grow the SRE team including hiring, mentoring, and developing reliability engineering capabilities.\nDefine SRE practice including SLIs, SLOs, error budgets, and reliability targets across critical platforms.\nDrive automation initiatives to eliminate toil and improve platform reliability and scalability.\nOversee incident management, blameless post-mortems, and systematic reliability improvement programs.\nCollaborate with cloud, infrastructure, and application teams to embed reliability into platform design.\nEstablish observability standards including unified monitoring, alerting, logging, and tracing strategies.\nManage on-call processes, escalation procedures, and team wellbeing for 24x7 operations.\nReport on platform reliability, SLO compliance, and operational maturity to IT leadership.\nCollaborate with Integration teams, IT infrastructure teams, cybersecurity teams, and external service providers to design and implement reliable connectivity solutions.\nEvaluate network technologies and solutions and provide technical recommendations based on business requirements, scalability, performance, security, and cost.\nEnsure network transformation initiatives align with applicable security, regulatory, and compliance requirements.\nPartner with cybersecurity teams to incorporate appropriate security controls, segmentation, firewall policies, VPN connectivity, and access controls into network architecture.\nDevelop and maintain comprehensive documentation for AIOps architectures, topologies, configurations, standards, policies, migration plans, and operational procedures.\nEstablish architecture standards, design principles, and best practices that promote consistency, scalability, reliability, and operational efficiency.\nProvide technical leadership and guidance to Observability engineering and operations teams throughout architecture, implementation, migration, and optimization activities.\nIdentify opportunities for continuous improvement, automation, standardization, and modernization of network operations.\nTroubleshoot and provide architectural guidance for complex Observability services\nStay current with emerging enterprise SRE, cloud networking, automation, and security technologies and assess their applicability to Altera’s environment.\nSalary Range\nThe pay range below is for Bay Area California only. Actual salary may vary based on a number of factors including job location, job-related knowledge, skills, experiences, trainings, etc. We also offer incentive opportunities that reward employees based on individual and company performance.\n$187,000 - $270,700 USD\nWe use artificial intelligence to screen, assess, or select applicants for the position. Applicants must be eligible for any required U.S. export authorizations.\nQualifications:\nMinimum Qualifications\nBachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field, with 10+ years of professional experience in enterprise networking.\n10+ years of experience in site reliability engineering or DevOps with a focus on platform reliability.\nProven experience leading technical engineering teams in a senior or management capacity.\nDeep understanding of SRE principles including SLIs, SLOs, error budgets, and reliability frameworks.\nStrong background in observability, incident management, and blameless post-mortem culture.\nExperience driving automation and toil reduction programs across infrastructure and platform teams.\nKnowledge of cloud platforms (Azure, AWS) and containerized environments.\nExcellent communication skills for stakeholder reporting and cross-functional collaboration.\nPreferred Qualifications\nExperience establishing SRE practices in semiconductor, HPC, or EDA-dependent environments.\nFamiliarity with AIOps platforms and AI-assisted incident management tooling.\nKnowledge of chaos engineering practices and resiliency testing frameworks.\nJob Type:\nRegularShift:\nShift 1 (United States of America)Primary Location:\nSan Jose, California, United StatesAdditional Locations:\nPosting Statement:\nAll qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.","description_format":"text","description_chars":6356,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Processors, MCUs & AI Chips"],"lifecycle":[{"event":"open","at":"2026-09-30T01:44:04Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":2,"expected_fill_days":80,"reasons":["conf:18","velocity","win:early","comp:brand"],"computed_at":"2026-10-01T05:45:00Z"},"pay":{"stated_usd_annual":270700,"is_top_pay":true},"html_url":"https://alion.io/job/altera-aiops-observabilitysre-lead","json_url":"https://alion.io/job/altera-aiops-observabilitysre-lead.json","meta":{"generated_at":"2026-10-02T01:57:35Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2532,"day_limit":5000,"remaining_today":2468,"minute_limit":60,"resets_at":"2026-10-03T00:00:00Z"}}}