{"id":1570200,"url":"https://alion.io/job/jumpmind-senior-cloud-operations-engineer","title":"Senior Cloud Operations Engineer","company":{"id":1773336,"name":"Jumpmind","domain":"jumpmind.com","url":"https://alion.io/company/jumpmind","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"ADP","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Columbus, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":105000,"max_usd":204000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":1020},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Amazon CloudWatch","optional":false},{"name":"AWS","optional":false},{"name":"Datadog","optional":false},{"name":"FinOps","optional":false},{"name":"Grafana","optional":false},{"name":"Kubernetes","optional":false},{"name":"LLM","optional":false},{"name":"PCI DSS","optional":false},{"name":"Prometheus","optional":false},{"name":"SOC 2","optional":false},{"name":"Terraform","optional":false},{"name":"Model Context Protocol","optional":true},{"name":"Python","optional":true},{"name":"SIEM","optional":true},{"name":"Slack","optional":true}],"status":"live","first_seen_at":"2026-09-30T13:32:00Z","employer_posted_date":"2026-09-30","last_verified_at":"2026-10-04T23:49:51Z","board_verified":true,"closed_at":null,"days_open":4,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":4},"description":"Senior Cloud Operations Engineer Role: Responsibilities, Qualifications, and AI Integration\nSenior Cloud Operations Engineer\nRole summary\nThe Senior Cloud Operations Engineer leads enterprise cloud platform operations, governance, security, reliability, and optimization across Amazon Web Services (AWS) and related cloud technologies. This person serves as the technical anchor of the CloudOps function, mentors the junior team members, and ensures cloud services remain stable, secure, scalable, cost-effective, and aligned with business priorities.\nThis role establishes platform standards, operational practices, service objectives, and accountability across the cloud environment. This person partners closely with Cloud Engineering, Security, and Enterprise IT to enable modernization while protecting operational stability and business continuity. AI fluency is core to this role: Jumpmind expects its CloudOps team to actively use AI tools and build lightweight agents to reduce toil, speed up diagnosis, and automate operational workflows, not just run infrastructure manually.\nDay-to-day responsibilities\nMonitor production infrastructure health, availability, and performance across cloud environments; own alerting and dashboards (uptime, latency, error rates, capacity)\n\nLead incident response for production issues - triage, coordinate, drive root-cause analysis, and own postmortems/corrective actions\n\nOperate and troubleshoot Kubernetes clusters in production - node health, pod scheduling issues, resource limits/requests, cluster upgrades, and workload scaling\n\nManage patching, OS/dependency updates, configs, and vulnerability remediation timelines for infrastructure in coordination with Security\n\nOwn capacity planning and scaling decisions ahead of peak retail traffic events\n\nDrive cloud cost optimization (FinOps) - rightsizing, reserved capacity, waste elimination, monthly spend reviews\n\nManage backups, disaster recovery testing, and recovery runbooks\n\nExecute infrastructure changes defined by Cloud Engineering's IaC (Terraform) - apply, validate, and operate what Engineering builds\n\nBuild and refine operational runbooks, on-call procedures, and escalation paths\n\nServe as the technical mentor for the Junior Cloud Operations Engineers - pairing, code/config review, on-call shadowing\n\nPartner with Cloud Engineering on handoffs from build to run; flag operational gaps (missing alerting, fragile deploy patterns) back to Engineering\n\nSupport audit and compliance evidence requests (SOC 2, PCI DSS) related to operational controls, access, and change management\n\nParticipate in an on-call rotation\n\nBuild and maintain AI agents and AI-assisted automations that handle operational toil - e.g., alert triage/summarization, log analysis, runbook execution, auto-remediation of known failure patterns\n\nEvaluate and integrate AI/LLM-powered tools into the operations workflow (incident summarization, on-call copilots, ChatOps assistants) and drive adoption across the team\n\nUse AI coding/agent tools day-to-day to write and maintain operational scripts, IaC change validation, and internal tooling faster\n\nRequired qualifications\n5+ years in cloud operations, SRE, DevOps, or infrastructure engineering roles\n\nDeep hands-on experience with a major cloud provider (AWS preferred)\n\nStrong troubleshooting skills across networking, compute, storage, and containerized workloads\n\nExperience owning incident response and writing postmortems\n\nComfort reading and operating Terraform managed infrastructure (not necessarily authoring modules from scratch)\n\nExperience with monitoring/observability tooling (e.g., CloudWatch, Datadog, Grafana, Prometheus)\n\nFamiliarity with compliance-driven environments (SOC 2, PCI DSS) and working alongside a security team\n\nExcellent written communication - this role will produce runbooks, postmortems, and audit evidence\n\nHands-on experience using AI tools in daily technical work, and demonstrated experience building or configuring AI agents/automations\n\nExperience leveraging AI to automate operational workflows - examples might include auto-generated incident summaries, AI-assisted root cause analysis, or agentic remediation scripts\n\nPreferred qualifications\nExperience operating infrastructure for a SaaS or retail/commerce platform with high-availability requirements\n\nScripting ability (Python, Bash, or Go) for operational tooling and automation\n\nPrior experience mentoring or leading a junior engineer\n\nExposure to SIEM or security monitoring tooling\n\nExperience with agent frameworks or protocols (e.g., MCP) and integrating AI agents with internal systems (Slack, ticketing, monitoring)\n\nFamiliarity with AI governance/security concepts (model access controls, data exposure risks in AI tooling)","description_format":"text","description_chars":4769,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Commerce","Data & Analytics"],"lifecycle":[{"event":"open","at":"2026-10-01T07:39:58Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":3,"expected_fill_days":24,"reasons":["conf:0","win:early"],"computed_at":"2026-10-04T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/jumpmind-senior-cloud-operations-engineer","json_url":"https://alion.io/job/jumpmind-senior-cloud-operations-engineer.json","meta":{"generated_at":"2026-10-05T01:17:21Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1690,"day_limit":5000,"remaining_today":3310,"minute_limit":60,"resets_at":"2026-10-06T00:00:00Z"}}}