{"id":1495375,"url":"https://alion.io/job/peraton-lead-site-reliability-engineer","title":"Lead Site Reliability Engineer","company":{"id":722,"name":"Peraton","domain":"peraton.com","url":"https://alion.io/company/peraton","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"iCIMS","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":112000,"max":179000,"currency":"USD","period":"year","gross":null,"usd_annual":179000},"salary_estimate":null,"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon CloudWatch","optional":false},{"name":"AWS","optional":false},{"name":"C#","optional":false},{"name":"CI/CD","optional":false},{"name":"CloudFormation","optional":false},{"name":"Datadog","optional":false},{"name":"Docker","optional":false},{"name":"GitHub Actions","optional":false},{"name":"Go","optional":false},{"name":"IAM","optional":false},{"name":"Java","optional":false},{"name":"Kubernetes","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Progressive Delivery","optional":false},{"name":"Python","optional":false},{"name":"Terraform","optional":false},{"name":"GitLab","optional":true},{"name":"Jenkins","optional":true},{"name":"Zero Trust","optional":true}],"status":"live","first_seen_at":"2026-09-30T01:56:13Z","employer_posted_date":"2026-09-30","last_verified_at":"2026-10-03T23:36:29Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"Responsibilities\nPeraton is seeking a Lead Site Reliability Engineer to join our team of qualified, diverse individuals. The ideal candidate will play a critical role in ensuring the reliability, resilience, and recoverability of mission essential platforms by leading infrastructure level disaster recovery drills, platform rebuild validation, and automated deployment processes. This engineer will partner closely with cross functional teams to maintain and enhance complex cloud-based environments, integrate modern automation solutions, and support largescale modernization and continuity efforts across high visibility programs.\nResponsibilities\nThe Lead Site Reliability Engineer ’s responsibilities shall include, but are not limited to:\nSupporting full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks.\nExecuting infrastructure level drill activities to ensure the platform can be fully rebuilt within the 48hour recovery target.\nVerifying end to end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations.\nIdentifying exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and driving corrective actions with engineering teams.\nDesigning, implementing, and supporting automated IaaC workflows utilizing Terraform, AWS CloudFormation, and standardized CI/CD pipelines.\nManaging and optimizing Kubernetes clusters and containerized workloads (Docker), including cluster provisioning, scaling, and workload reliability improvements.\nBuilding and maintaining observability solutions using CloudWatch, Datadog, and other monitoring/alerting tools to ensure service reliability and proactive incident response.\nDeveloping automation, tooling, and scripts using Python or Java to reduce manual processes and enhance operational repeatability.\nCollaborating with platform engineering, security, applications, and data teams to ensure consistent, secure, and compliant platform operations.\nParticipating in on-call rotations, root cause analyses, and incident response activities to improve system resilience and operational excellence.\nQualifications\nRequired Qualifications\nBachelor’s degree and 8-10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; or 12 years of experience with a high school diploma. \nExpert level hands on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.\nStrong experience with Infrastructure as Code (Terraform, CloudFormation) and infrastructure automation principles.\nExperience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub actions or similar tools\n Deep understanding of Kubernetes administration, container orchestration, and Docker based deployments.\nProven experience validating DR processes, performing system rebuilds, and conducting data integrity checks.\nExperience building monitoring tools like dashboards, metrics, logs, and alerting systems using CloudWatch, Datadog, or similar observability tools.\nProficiency with programming/scripting languages such as Python, Java, or C# or Go.\nExperience debugging complex failure modes, including cascading failures, network partitions, backpressure, and eventual consistency issues.\nStrong analytical and documentation skills with the ability to clearly communicate technical findings to cross functional teams.\nAbility to work in a fast-paced environment supporting high visibility, mission critical systems.\nAbility to obtain a Public Trust clearance. \nUS Citizen or Green Card Holder. \nPreferred Qualifications\nAWS DevSecOps Engineer certification (preferred).\nAdditional AWS certifications (Solutions Architect, SysOps, Developer) and/or Kubernetes certifications (CKA, CKAD).\nFamiliarity with Zero Trust security models and cloud security best practices.\nExperience with GitLab, Jenkins, or similar CI/CD platforms.\nExperience with highly regulated environments (healthcare, finance, DHS, DoD, CMS, etc.).\nExperience supporting federal, defense, or largescale enterprise programs involving legacy-to-cloud modernization.\nPrior involvement in large-scale DR drills, continuity of operations (COOP), or portability/executable readiness assessments.\nPeraton Overview\nPeraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.\nTarget Salary Range\n$112,000 - $179,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.EEO\nEEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.","description_format":"text","description_chars":5844,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"high_school","optional":false},"security_clearance":true,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["National Security","Homeland Security"],"lifecycle":[{"event":"open","at":"2026-09-30T01:56:13Z"}],"visa":[],"liveness":{"score":39,"band":"fade","label":"Fading","p_open":0.9,"p_active":0.781,"p_room":0.55,"age_days":3,"expected_fill_days":3,"reasons":["conf:66","velocity","win:tail","comp:brand"],"computed_at":"2026-10-03T05:45:00Z"},"pay":{"stated_usd_annual":179000,"is_top_pay":false},"html_url":"https://alion.io/job/peraton-lead-site-reliability-engineer","json_url":"https://alion.io/job/peraton-lead-site-reliability-engineer.json","meta":{"generated_at":"2026-10-04T01:44:22Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2303,"day_limit":5000,"remaining_today":2697,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}