{"id":1427408,"url":"https://alion.io/job/nvidia-senior-software-engineer-nemo-core-platform-2","title":"Senior Software Engineer, NeMo Core Platform","company":{"id":6,"name":"NVIDIA","domain":"nvidia.com","url":"https://alion.io/company/nvidia","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":89,"open_postings":299,"ghost_share":0.017,"stale_share":0.381,"repost_share":0.033,"time_to_fill_p50_days":29,"computed_at":"2026-10-03T05:45:00Z"}},"role":"Backend","role_family":"Backend","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Toronto, Canada","Canada"],"countries":["CA"],"hiring_countries":["CA"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":141000,"max_usd":275000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":104},"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Docker","optional":false},{"name":"Kubernetes","optional":false},{"name":"NVIDIA NeMo","optional":false},{"name":"Python","optional":false},{"name":"SLURM","optional":false}],"status":"live","first_seen_at":"2026-09-28T00:00:00Z","employer_posted_date":"2026-09-28","last_verified_at":"2026-10-04T00:54:03Z","board_verified":true,"closed_at":null,"days_open":6,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":6},"description":"We are looking for a Senior Software Engineer to help build NeMo Platform, NVIDIA’s product for developing, evaluating, deploying, and operating AI systems at scale. This role is for a senior engineer/architect for our Core team which owns and ships an open source plugin-based AI platform for running and optimizing Agents targeting multiple compute backends (local/docker, Kubernetes, Slurm, etc.).\nAs AI systems become more autonomous and more deeply integrated into real workflows, teams need robust APIs and orchestration systems for running, monitoring, and optimizing Agents at scale. Increasingly, the users of these systems are themselves autonomous or semi-autonomous Agents. The NeMo Platform group is building a sophisticated agent execution framework to enable agents to automatically run hundreds of experiments in parallel to find the most efficient agent architectures for our customers. This is important product engineering research for making agents more sustainable. AI systems are not yet nearly as efficient as they can be, and systems like NeMo Platform will allow large scale AI consumers to automatically tune their agents to use fewer tokens and rely on more efficient models with better throughput.\nWhat you'll be doing:\nWorking in a product research environment where we place big bets on where the future is heading, adapting in real time as we build alongside an industry that is constantly evolving with us. This means fast iteration, high ownership, pragmatic decisions, and performance-minded implementation under production constraints\n\nDesigning an Agentic Execution system that are flexible enough to work in many environments (local, k8s, Slurm, on prem / air-gapped)\n\nProvide senior technical leadership through design reviews, code reviews, mentoring, and ownership of ambiguous cross-component problems\n\nBuilding and maintain our Core Platform APIs for running jobs, storing data, entities, secrets, and RBAC and Auth\n\nExtending our flexible Plugin Architecture that makes it easy for many teams and external customers to install new capabilities into our system\n\nBuilding in the open in our OSS repo, keeping up the high standards that the open source community demands\n\nShipping code at the speed of light with an unlimited token budget using best in class agentic coding tools\n\nImproving reliability, observability, debuggability, and performance across NeMo Platform, SDKs, plugins, jobs, and developer workflows\n\nBuilding strong test coverage across unit, integration, E2E, Docker, and Kubernetes workflows\n\nWhat we need to see:\nBS, MS, or equivalent experience in Computer Science, Computer Engineering, or a related technical field\n\n10+ years of professional software engineering experience building production systems\n\nComfort working in a very fast and ambiguous environment\n\nExceptional communication, both verbal and written. This includes the ability to produce and review high quality architectural RFCs, and to discuss them clearly with the right level of technical detail for the right people (Engineer, Product, Marketing, etc.)\n\nStrong system design skills, with a pragmatic flexibility and phenomenal instincts to invent robust systems quickly without over-complicating. Strong understanding of reliability, scalability, security, and performance tradeoffs in production infrastructure\n\nExperience with distributed systems, cloud-native services, containers, Kubernetes, and job orchestration\n\nExcellent Python engineering skills, including API design, typing, testing, debugging, performance analysis, and maintainable software design\n\nExperience designing SDKs, libraries, plugins, CLIs, or other developer-facing interfaces\n\nAbility to work independently, define technical scope, break down ambiguous problems, and drive work across team boundaries\n\nWays to stand out from the crowd:\nExperience building, deploying, and iterating on production agentic AI systems at scale in Kubernetes\n\nExperience with sophisticated plugin architectures\n\nStrong ability to connect technical evaluation work to business outcomes, product quality, user experience, reliability, or operational efficiency\n\nExperience with enterprise AI systems where measurement, regression testing, observability, governance, and continuous improvement are required for production deployment\n\nYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 170,000 CAD - 220,000 CAD for Level 4, and 225,000 CAD - 275,000 CAD for Level 5.You will also be eligible for equity and benefits.\nApplications for this job will be accepted at least until October 2, 2026.This posting is for an existing vacancy.\nNVIDIA uses AI tools in its recruiting processes.","description_format":"text","description_chars":4759,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Equity"],"hiring_locations":[{"name":"Canada","iso":"CA","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Processors, MCUs & AI Chips","Servers & Data Center Hardware","Computer Components","AI Chips & Accelerators"],"lifecycle":[{"event":"open","at":"2026-09-29T00:12:53Z"}],"visa":[],"liveness":{"score":89,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.894,"p_room":1,"age_days":5,"expected_fill_days":29,"reasons":["conf:6","velocity","win:early","comp:brand"],"computed_at":"2026-10-03T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/nvidia-senior-software-engineer-nemo-core-platform-2","json_url":"https://alion.io/job/nvidia-senior-software-engineer-nemo-core-platform-2.json","meta":{"generated_at":"2026-10-04T02:47:58Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4152,"day_limit":5000,"remaining_today":848,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}