Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 30, 2026. Jumpmind scores B on the Alion truth index.
Senior Cloud Operations Engineer Role: Responsibilities, Qualifications, and AI Integration
Senior Cloud Operations Engineer
Role summary
The Senior Cloud Operations Engineer leads enterprise cloud platform operations, governance, security, reliability, and optimization across Amazon Web Services (AWS) and related cloud technologies. This person serves as the technical anchor of the CloudOps function, mentors the junior team members, and ensures cloud services remain stable, secure, scalable, cost-effective, and aligned with business priorities.
This role establishes platform standards, operational practices, service objectives, and accountability across the cloud environment. This person partners closely with Cloud Engineering, Security, and Enterprise IT to enable modernization while protecting operational stability and business continuity. AI fluency is core to this role: Jumpmind expects its CloudOps team to actively use AI tools and build lightweight agents to reduce toil, speed up diagnosis, and automate operational workflows, not just run infrastructure manually.
Day-to-day responsibilities
Monitor production infrastructure health, availability, and performance across cloud environments; own alerting and dashboards (uptime, latency, error rates, capacity)
Lead incident response for production issues - triage, coordinate, drive root-cause analysis, and own postmortems/corrective actions
Operate and troubleshoot Kubernetes clusters in production - node health, pod scheduling issues, resource limits/requests, cluster upgrades, and workload scaling
Manage patching, OS/dependency updates, configs, and vulnerability remediation timelines for infrastructure in coordination with Security
Own capacity planning and scaling decisions ahead of peak retail traffic events
Drive cloud cost optimization (FinOps) - rightsizing, reserved capacity, waste elimination, monthly spend reviews
Manage backups, disaster recovery testing, and recovery runbooks
Execute infrastructure changes defined by Cloud Engineering's IaC (Terraform) - apply, validate, and operate what Engineering builds
Build and refine operational runbooks, on-call procedures, and escalation paths
Serve as the technical mentor for the Junior Cloud Operations Engineers - pairing, code/config review, on-call shadowing
Partner with Cloud Engineering on handoffs from build to run; flag operational gaps (missing alerting, fragile deploy patterns) back to Engineering
Support audit and compliance evidence requests (SOC 2, PCI DSS) related to operational controls, access, and change management
Participate in an on-call rotation
Build and maintain AI agents and AI-assisted automations that handle operational toil - e.g., alert triage/summarization, log analysis, runbook execution, auto-remediation of known failure patterns
Evaluate and integrate AI/LLM-powered tools into the operations workflow (incident summarization, on-call copilots, ChatOps assistants) and drive adoption across the team
Use AI coding/agent tools day-to-day to write and maintain operational scripts, IaC change validation, and internal tooling faster
Required qualifications
5+ years in cloud operations, SRE, DevOps, or infrastructure engineering roles
Deep hands-on experience with a major cloud provider (AWS preferred)
Strong troubleshooting skills across networking, compute, storage, and containerized workloads
Experience owning incident response and writing postmortems
Comfort reading and operating Terraform managed infrastructure (not necessarily authoring modules from scratch)
Experience with monitoring/observability tooling (e.g., CloudWatch, Datadog, Grafana, Prometheus)
Familiarity with compliance-driven environments (SOC 2, PCI DSS) and working alongside a security team
Excellent written communication - this role will produce runbooks, postmortems, and audit evidence
Hands-on experience using AI tools in daily technical work, and demonstrated experience building or configuring AI agents/automations
Experience leveraging AI to automate operational workflows - examples might include auto-generated incident summaries, AI-assisted root cause analysis, or agentic remediation scripts
Preferred qualifications
Experience operating infrastructure for a SaaS or retail/commerce platform with high-availability requirements
Scripting ability (Python, Bash, or Go) for operational tooling and automation
Prior experience mentoring or leading a junior engineer
Exposure to SIEM or security monitoring tooling
Experience with agent frameworks or protocols (e.g., MCP) and integrating AI agents with internal systems (Slack, ticketing, monitoring)
Familiarity with AI governance/security concepts (model access controls, data exposure risks in AI tooling)

