{"id":1638298,"url":"https://alion.io/job/isone-senior-site-reliability-engineer","title":"Senior Site Reliability Engineer","company":{"id":3854059,"name":"isone","domain":"isone.com","url":"https://alion.io/company/isone","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Dayforce","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Holyoke, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":134000,"max":170000,"currency":"USD","period":"year","gross":null,"usd_annual":170000},"salary_estimate":null,"experience_years_min":5,"visa_sponsorship":true,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Dynatrace","optional":false},{"name":"Opsgenie","optional":false},{"name":"Platform Engineering","optional":false},{"name":"PowerShell","optional":false},{"name":"Python","optional":false},{"name":"Splunk","optional":false},{"name":"Terraform","optional":false},{"name":"ITIL","optional":true}],"status":"live","first_seen_at":"2026-08-17T04:00:00Z","employer_posted_date":"2026-10-01","last_verified_at":"2026-10-08T01:24:35Z","board_verified":true,"closed_at":null,"days_open":51,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":51},"description":"The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and application teams to reduce operational toil, improve service resilience, and implement scalable automation solutions. \n\nThis role has a strong emphasis on observability engineering, automation, Splunk administration, and Infrastructure as Code (IaC). The ideal candidate will possess hands-on experience with Splunk or demonstrate a strong willingness to develop expertise in the platform. Experience with Terraform, automation technologies such as Python and PowerShell, and the ability to leverage AI-assisted development tools to accelerate engineering solutions are key components of the role.\n\nWhat we offer you:\nA stable, mission-driven workplace where your impact truly matters\nA highly engaged work environment that values inclusion, collaboration, and employee safety and wellbeing\nCompetitive compensation with a base salary + performance bonus\nRobust benefits package, including:\nEnhanced 401(k) and financial planning support\nTuition reimbursement and professional development \nWellness programs, including an onsite gym\nFlexible work hours\nEmployee Business Networks \nFree coffee at our onsite café\nHybrid work environment (3 days/week onsite)\nDistance-based relocation assistance available\n\nHow you will make an Impact\nBuild and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage) \nAdminister, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflows \nDevelop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overhead \nDevelop meaningful KPIs and dashboards for business and IT service health \nEngineer and implement resilience patterns including HA, DR, and automated failover \nPartner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverage \nConduct performance testing, capacity modeling, forecasting, and right-sizing \nParticipate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvements \nIdentify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutions \nDesign, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, PowerShell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiency \nDesign, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainability \nIdentify gaps in observability coverage and drive engineering solutions to close them \nCollaborate with architecture and application teams to ensure production readiness \nLeverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectiveness \nReduce repeat incidents by engineering permanent fixes and driving continuous improvement \n\nWhat we are looking for\n5+ years of experience in SRE, DevOps, systems engineering, platform engineering, or IT operations \nExperience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness and aptitude to develop expertise in Splunk administration, engineering, and automation. \nExperience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practices \nStrong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environments \nDemonstrated experience designing, developing, and supporting automation solutions that measurably reduced manual operational effort in an enterprise environment \nAbility to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platforms \nKnowledge of distributed systems, networking, enterprise infrastructure, and cloud platforms \nFamiliarity with SRE principles including SLOs, error budgets, observability, and toil reduction \nAbility to analyze and troubleshoot complex technical systems \nPreferred Qualifications \nExperience in mission-critical, highly available, or regulated environments \nExperience utilizing AI-assisted development tools to accelerate automation, operational engineering, or platform management activities \nKnowledge of ITIL processes and/or SRE best practices \nExperience with performance testing, capacity planning, resilience testing, or disaster recovery validation \n\nThis employer will not sponsor applicants for work visas for this position (ex: H-1B, F-1/CPT/OPT, O-1, E-3, TN, J, etc.).\n\nThe expected salary range for this position is $134,000 - $170,000 per year, for a Senior to Lead level candidate. This role is also eligible for an annual performance bonus, comprehensive health insurance (medical, dental and vision), flexible spending and health savings accounts, a 401(k) plan with generous employer contributions and a student debt benefit, life and AD&D insurance, disability insurance, critical illness and hospital indemnity benefits, paid time off, paid leave, a wellness program, an employee assistance program and other great company perks.","description_format":"text","description_chars":6279,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Flexible schedule","Health insurance","Hybrid work","Professional development","Relocation assistance","Wellness"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":true,"industries":[],"lifecycle":[{"event":"open","at":"2026-10-01T21:41:30Z"}],"visa":[],"liveness":{"score":41,"band":"fade","label":"Fading","p_open":1,"p_active":0.739,"p_room":0.55,"age_days":51,"expected_fill_days":39,"reasons":["conf:2","velocity","win:tail"],"computed_at":"2026-10-07T05:47:15Z"},"pay":{"stated_usd_annual":170000,"is_top_pay":true},"html_url":"https://alion.io/job/isone-senior-site-reliability-engineer","json_url":"https://alion.io/job/isone-senior-site-reliability-engineer.json","meta":{"generated_at":"2026-10-08T01:47:05Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3121,"day_limit":5000,"remaining_today":1879,"minute_limit":60,"resets_at":"2026-10-09T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3854059},"rest":"https://alion.io/mcp/rest/get_company?id=3854059"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fisone-senior-site-reliability-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fisone-senior-site-reliability-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fisone-senior-site-reliability-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/isone-senior-site-reliability-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fisone-senior-site-reliability-engineer"}]}