{"id":1198858,"url":"https://alion.io/job/leidos-site-reliability-engineer-4","title":"Site Reliability Engineer","company":{"id":7791,"name":"Leidos","domain":"leidos.com","url":"https://alion.io/company/leidos","size_band":"5000+","is_staffing_agency":false,"employer_type":"services","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":86,"open_postings":93,"ghost_share":0,"stale_share":0.559,"repost_share":0,"time_to_fill_p50_days":22,"computed_at":"2026-09-26T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Baltimore, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":107900,"max":195050,"currency":"USD","period":"year","gross":null,"usd_annual":195050},"salary_estimate":null,"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"Amazon Aurora","optional":false},{"name":"Amazon S3","optional":false},{"name":"Ansible","optional":false},{"name":"API Gateway","optional":false},{"name":"AWS","optional":false},{"name":"AWS Glue","optional":false},{"name":"AWS Lambda","optional":false},{"name":"CI/CD","optional":false},{"name":"CloudFormation","optional":false},{"name":"Configuration Management","optional":false},{"name":"FedRAMP","optional":false},{"name":"PostgreSQL","optional":false},{"name":"SRE","optional":false},{"name":"Terraform","optional":false},{"name":"Zero Trust","optional":false},{"name":"Amazon CloudWatch","optional":true},{"name":"Amazon ECS","optional":true},{"name":"Amazon EKS","optional":true},{"name":"Docker","optional":true},{"name":"Kubernetes","optional":true},{"name":"New Relic","optional":true},{"name":"Splunk","optional":true}],"status":"live","first_seen_at":"2026-09-24T19:43:10Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-27T03:09:46Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services (CMS) modernize its legacy Contact Center CRM platform within the CMS AWS Enclave and Pega Cloud for Government.\nPrimary Responsibilities\nThe SRE/Senior Cloud Engineer shall design, build, and operate the highly available cloud infrastructure supporting the CRM modernization effort.\nResponsibilities shall include, but are not limited to:\nDesign, build, and operate highly available AWS infrastructure within the CMS AWS Enclave (FedRAMP Moderate), applying AWS Well-Architected Framework best practices.\nArchitect secure, scalable multi-account VPC interconnectivity between the CMS AWS Enclave and Pega Cloud for Government (PCFG) in AWS GovCloud US-West, including PrivateLink, API Gateway, and Direct Connect.\nSupport container and serverless architectures (e.g., AWS Lambda, Glue) for data integration, batch processing, and API layers supporting the modernized CRM.\nApply Site Reliability Engineering (SRE) practices, including defining and tracking Service Level Objectives (SLOs), error budgets, and reliability metrics aligned to contract SLAs (e.g., ≥99.9% availability).\nBuild and maintain observability across the AWS and Pega environments, including centralized logging, monitoring, and alerting, to enable proactive detection of performance and availability issues.\nLead incident response and root-cause analysis for production issues, and long-term reliability improvements.\nAutomate infrastructure provisioning, configuration, and environment build-out using Infrastructure as Code (e.g., Terraform, CloudFormation, Ansible).\nDesign and test Disaster Recovery capability for cloud-based workloads, including backup, failover, and Multi-AZ/Multi-Region resilience.\nSupport performance testing and capacity planning to validate the platform's ability to scale to 20,000 concurrent CSR sessions and peak Open Enrollment Period (OEP) volumes.\nSupport continuous security monitoring, vulnerability remediation, and Zero Trust alignment across the AWS and Pega environments.\nPartner with the DevOps Lead/Configuration Manager to build and maintain CI/CD pipelines, ensuring automated testing, security scanning, and deployment across all SDLC environments.\nSupport cloud connectivity and data movement for the AWS-based data migration pipeline (e.g., AWS Glue, S3, RDS/Aurora PostgreSQL) between legacy Siebel and the modernized CRM.\nCoordinate with the CMS Hybrid Cloud Team on cloud environment provisioning, patching, and lifecycle management activities.\nManage release coordination and change windows in support of OEP blackout periods and other critical operational periods, minimizing risk of service disruption.\nCollaborate with the Release Train Engineer, Solution Architect, and Agile delivery teams to align infrastructure readiness with sprint and PI planning.\nSupport integration of Genesys Cloud CX infrastructure and telephony/chat channels with the modernized CRM environment.\nContinuously identify opportunities to reduce operational toil through automation of repetitive tasks, log analysis, and routine operational activities.\nDocument cloud architecture, operational runbooks, and disaster recovery procedures to support the Transition-Out Plan and audit readiness.\nProvide technical mentoring and knowledge-sharing to other engineers on cloud architecture, automation, and reliability engineering best practices.\nCommunicate technical status, risks, and dependencies to CMS leadership, and Leidos management.\nSupport requirements traceability and technical documentation related to infrastructure and integration architecture.\nActively participate in planning sessions, requirements gathering activities, design sessions, Agile sessions, and other events supporting the CRM modernization effort.\nRequired Qualifications:\nBachelor’s degree and a minimum of 6-8 years of relevant experience in cloud engineering, site reliability engineering, or infrastructure, or an equivalent combination of education and experience\nExperience building highly available AWS infrastructure based on industry best practices and the AWS Well-Architected Framework\nExperience with Infrastructure as Code, automation, and configuration management of cloud-based resources\nExperience designing Disaster Recovery for cloud-based workloads, including Multi-AZ/Multi-Region resilience\nAbility to obtain Public Trust\nPreferred Qualifications:\nAWS Solutions Architect or SysOps Administrator Certification\nExperience with AWS GovCloud and FedRAMP-authorized cloud environments\nExperience supporting Pega Cloud for Government (PCFG) or similar SaaS platform connectivity (e.g., AWS PrivateLink, VPC peering)\nFamiliarity with container (Docker, ECS, EKS) and serverless (Lambda) architectures\nExperience with observability/monitoring tooling (e.g., Splunk, New Relic, CloudWatch) and incident response practices\nExperience supporting federal contact center or other 24x7 mission-critical government systems\nAgile delivery experience\nIf you're looking for comfort, keep scrolling. At Leidos, we outthink, outbuild, and outpace the status quo - because the mission demands it. We're not hiring followers. We're recruiting the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We're already at step 30 - and moving faster than anyone else dares.\nOriginal Posting:\nSeptember 24, 2026For U.S. Positions: While subject to change based on business needs, Leidos reasonably anticipates that this job requisition will remain open for at least 3 days with an anticipated close date of no earlier than 3 days after the original posting date as listed above.\nPay Range:\nPay Range $107,900.00 - $195,050.00The Leidos pay range for this job level is a general guideline only and not a guarantee of compensation or salary. Additional factors considered in extending an offer include (but are not limited to) responsibilities of the job, education, experience, knowledge, skills, and abilities, as well as internal equity, alignment with market data, applicable bargaining agreement (if any), or other law.","description_format":"text","description_chars":6285,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Equity"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["National Security","C4ISR","Defense Manufacturing","Engineering Services"],"lifecycle":[{"event":"open","at":"2026-09-24T19:43:10Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":1,"expected_fill_days":22,"reasons":["conf:2","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-09-26T05:45:00Z"},"pay":{"stated_usd_annual":195050,"is_top_pay":true},"html_url":"https://alion.io/job/leidos-site-reliability-engineer-4","json_url":"https://alion.io/job/leidos-site-reliability-engineer-4.json","meta":{"generated_at":"2026-09-27T04:08:30Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3885,"day_limit":5000,"remaining_today":1115,"minute_limit":60,"resets_at":"2026-09-28T00:00:00Z"}}}