{"id":43829,"url":"https://alion.io/job/nvidia-principal-engineer-cloud-site-reliability-engineering","title":"Principal Engineer, Cloud Site Reliability Engineering","company":{"id":6,"name":"NVIDIA","domain":"nvidia.com","url":"https://alion.io/company/nvidia","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":87,"open_postings":369,"ghost_share":0.014,"stale_share":0.458,"repost_share":0.03,"time_to_fill_p50_days":29,"computed_at":"2026-10-08T05:49:30Z"}},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Santa Clara, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":272000,"max":431250,"currency":"USD","period":"year","gross":null,"usd_annual":431250},"salary_estimate":null,"experience_years_min":15,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Apache Kafka","optional":false},{"name":"Cassandra","optional":false},{"name":"Chef","optional":false},{"name":"CI/CD","optional":false},{"name":"Docker","optional":false},{"name":"ElasticSearch","optional":false},{"name":"Git","optional":false},{"name":"Hadoop","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Machine Learning","optional":false},{"name":"MySQL","optional":false},{"name":"OpenStack","optional":false},{"name":"Puppet","optional":false},{"name":"Python","optional":false},{"name":"Rest API","optional":false},{"name":"SQL","optional":false},{"name":"Windows","optional":false}],"status":"live","first_seen_at":"2026-08-05T00:00:00Z","employer_posted_date":"2026-08-05","last_verified_at":"2026-10-09T02:40:27Z","board_verified":true,"closed_at":null,"days_open":65,"trust":{"level":"ok","repost_count":1,"flags":[],"days_open":64},"description":"NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. These cloud services provide almost half a million automated jobs per day on thousands of servers helping with the efficiency of thousands of NVIDIA's software engineers worldwide. The cloud hosts various machines and devices with operating systems like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers unified CI/CD solutions and cloud-based software development. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?\nWhat you'll be doing:\nServe as an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing.\n\nEvaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.\n\nArchitecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software.\n\nCustomer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud.\n\nIdentify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.\n\nLeading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions.\n\nLooking for problems within software systems and resolving the issues\n\nCraft and implement critical metrics using various analytics methods and dashboards.\n\nWhat we need to see:\nBS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).\n\n15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.\n\nExperience of maintaining cloud infrastructure and highly available production environment.\n\nStrong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.\n\nExperience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch.\n\nExcellent knowledge and working experience with Docker containers and Virtual Machines.\n\nGood background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.\n\nAbility to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment.\n\nWays to stand out from the crowd:\nDepth in AI, Machine Learning and Deep Learning algorithms and techniques.\n\nStrong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.\n\nExperience developing large-scale software systems using modular architecture under real-time performance requirements.\n\nBackground in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.\n\nYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.\nApplications for this job will be accepted at least until August 9, 2026.This posting is for an existing vacancy.\nNVIDIA uses AI tools in its recruiting processes.\nNVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.","description_format":"text","description_chars":4435,"description_truncated":false,"requirements":{"experience_years_min":15,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Equity"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Processors, MCUs & AI Chips","Servers & Data Center Hardware","Computer Components","AI Chips & Accelerators"],"lifecycle":[{"event":"open","at":"2026-08-05T00:00:00Z"},{"event":"close","at":"2026-08-07T22:42:34Z"},{"event":"reopen","at":"2026-09-03T08:08:33Z"}],"visa":[],"liveness":{"score":17,"band":"cold","label":"Long shot","p_open":1,"p_active":0.624,"p_room":0.28,"age_days":64,"expected_fill_days":29,"reasons":["conf:1","velocity","win:tail","crowd:brand"],"computed_at":"2026-10-08T05:49:30Z"},"pay":{"stated_usd_annual":431250,"is_top_pay":true},"html_url":"https://alion.io/job/nvidia-principal-engineer-cloud-site-reliability-engineering","json_url":"https://alion.io/job/nvidia-principal-engineer-cloud-site-reliability-engineering.json","meta":{"generated_at":"2026-10-09T04:07:26Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3601,"day_limit":5000,"remaining_today":1399,"minute_limit":60,"resets_at":"2026-10-10T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":6},"rest":"https://alion.io/mcp/rest/get_company?id=6"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fnvidia-principal-engineer-cloud-site-reliability-engineering"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fnvidia-principal-engineer-cloud-site-reliability-engineering"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fnvidia-principal-engineer-cloud-site-reliability-engineering"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/nvidia-principal-engineer-cloud-site-reliability-engineering\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fnvidia-principal-engineer-cloud-site-reliability-engineering"}]}