{"id":1524823,"url":"https://alion.io/job/nvidia-senior-devops-engineer-aiops","title":"Senior DevOps Engineer, AIOps","company":{"id":6,"name":"NVIDIA","domain":"nvidia.com","url":"https://alion.io/company/nvidia","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":89,"open_postings":302,"ghost_share":0.017,"stale_share":0.387,"repost_share":0.033,"time_to_fill_p50_days":29,"computed_at":"2026-10-04T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Tel Aviv, Israel","Israel"],"countries":["IL"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":101000,"max_usd":263000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":2013},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"AIOps","optional":false},{"name":"Amazon S3","optional":false},{"name":"Ansible","optional":false},{"name":"Bash","optional":false},{"name":"CI/CD","optional":false},{"name":"ClickHouse","optional":false},{"name":"Datadog","optional":false},{"name":"DNS","optional":false},{"name":"Docker","optional":false},{"name":"FastAPI","optional":false},{"name":"GitLab CI","optional":false},{"name":"Grafana","optional":false},{"name":"Helm","optional":false},{"name":"JFrog Artifactory","optional":false},{"name":"Kubernetes","optional":false},{"name":"Langfuse","optional":false},{"name":"LangGraph","optional":false},{"name":"Linux","optional":false},{"name":"LLM","optional":false},{"name":"Model Context Protocol","optional":false},{"name":"Node JS","optional":false},{"name":"OpenShift","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Platform Engineering","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"React.js","optional":false},{"name":"Redis","optional":false},{"name":"SQL","optional":false},{"name":"SRE","optional":false},{"name":"TCP/IP","optional":false},{"name":"Temporal","optional":false},{"name":"Terraform","optional":false},{"name":"Go","optional":true},{"name":"JavaScript","optional":true},{"name":"LangChain","optional":true}],"status":"live","first_seen_at":"2026-09-29T00:00:00Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-05T00:20:27Z","board_verified":true,"closed_at":null,"days_open":6,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":6},"description":"NVIDIA is powering the world’s most advanced AI factories, where resilient infrastructure is essential to keep accelerated computing environments running at scale. The Agentic AIOps team is building a mission-critical observability and prediction platform - delivered as both a high-scale SaaS solution and a robust on-premises deployment for NVIDIA’s largest enterprise customers.\nAs a Senior DevOps Engineer, you’ll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIA’s AI infrastructure.\nWhat You'll Be Doing:\nOwn the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations. \nBuild and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints. \nEngineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory. \nAutomate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity. \nOperate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations. \nStrengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows. \nDeliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening. \nPartner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes. \nWhat We Need to See:\nBachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent experience. \n5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices. \nStrong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting. \nStrong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible. \nExperience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices. \nPractical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting. \nStrong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting. \nSound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment. \nWays To Stand Out From the Crowd:\nExperience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost. \nFamiliarity with Temporal, LangGraph, Model Context Protocol (MCP), Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage. \nDeep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana. \nExperience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers. \nExperience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain. \nWith competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you are passionate about building mission-critical systems at the frontier of AI infrastructure, we want to hear from you.","description_format":"text","description_chars":4572,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Processors, MCUs & AI Chips","Servers & Data Center Hardware","Computer Components","AI Chips & Accelerators"],"lifecycle":[{"event":"open","at":"2026-09-30T13:36:32Z"}],"visa":[],"liveness":{"score":88,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.881,"p_room":1,"age_days":5,"expected_fill_days":29,"reasons":["conf:0","velocity","win:early","comp:brand"],"computed_at":"2026-10-04T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/nvidia-senior-devops-engineer-aiops","json_url":"https://alion.io/job/nvidia-senior-devops-engineer-aiops.json","meta":{"generated_at":"2026-10-05T00:52:19Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1083,"day_limit":5000,"remaining_today":3917,"minute_limit":60,"resets_at":"2026-10-06T00:00:00Z"}}}