Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Aug 20, 2026. Upscale AI scores B on the Alion truth index.
About the role
AI training and inference clusters live or die on the network fabric. ScaleUp builds the software that programs our high-performance Ethernet switch silicon - SDK, SAI, simulation, and the CI that keeps that stack shippable.
We’re looking for a DevOps engineer at the intersection of AI infrastructure, data-center networking, and silicon-aware software: reliable pipelines for an ASIC SDK, fast feedback for developers, and release paths worthy of cloud- and AI-scale deployment.
You’ll partner with SDK, SAI, and QA engineers, and work in GitHub + Jira day to day. Success means CI is trusted, simulation and software gates are clear, and infrastructure never blocks the next generation of AI networking features.
Responsibility
- Own CI/CD for the ScaleUp stack that powers AI/HPC Ethernet switching (build, unit/integration, gating, artifacts, release promotion).
- Scale pipelines across multi-repo dependencies: switch SDK, SAI adapters, shared test frameworks, and switch simulation / model targets.
- Improve build caching, shared runners, and nightlies so large C/C++ and Python SDK builds stay fast and predictable.
- Own release / promote workflows (branch policies, artifact publish, pre-gate vs post-gate steps) so SDK drops are repeatable.
- Shorten commit → green for teams building programmable data-plane, QoS, ACL, RoCE-era fabrics, and AI-workload traffic patterns.
- Add and maintain quality gates: lint, unit tests, model smoke, coverage where useful, overnight soak - high signal, low flake.
- Manage artifacts and caches (CI artifacts, object storage as needed) with clear retention and failure handling.
- Operate Linux build/test farms for C/C++ SDKs and Python harnesses (toolchains, containers, lab/sim hosts).
- Improve pipeline observability (dashboards, alerts, runbooks, postmortems).
- Secure secrets and access for GitHub Actions, artifact stores, and internal services.
- Partner with QA on PR / nightly / soak jobs that validate end-to-end switch behavior before customer and cloud deployments.
- Publish reusable pipeline templates so new ScaleUp / AI-networking repos onboard quickly.
- Jira: keep engineering work traceable - epics/stories/bugs for CI and infra, link PRs and releases to tickets, support sprint and release planning with SDK/SAI/QA, tighten ticket hygiene (status, components, labels) so blockers and CI debt are visible.
Qualification
- Strong CI/CD experience (GitHub Actions and/or Jenkins; GitLab CI also fine) in multi-repo environments.
- Deep Linux fluency: shells, packaging, toolchains, debugging native builds and shared-library / SDK load paths.
- Automation in Python and bash; enough CMake/Make to unblock switch-SDK CI.
- Containers (Docker) and self-hosted or cloud runners at meaningful scale.
- Proven work reducing flaky tests and improving CI signal-to-noise.
- Hands-on Jira (or equivalent): workflows, boards, linking commits/PRs to issues, release/version fields - comfortable driving process with developers, not only “keeping the lights on.”
- Clear communication with hardware-adjacent software teams.
- Excitement about AI data centers, Ethernet fabrics, and silicon-software co-design - not only generic cloud DevOps.
Nice to have
- Background in networking ASICs, switch SDKs, SAI, DPDK, or NIC/SmartNIC software.
- Exposure to AI/ML cluster networking (GPU fabrics, RoCE/RDMA, congestion control, telemetry).
- Simulation / hardware-in-the-loop CI, pytest at scale, junit/Allure-style reporting.
- Build-cache and large-monorepo / multi-repo release patterns; backport or branch-gating workflows.
- Static analysis / lint gates in CI (e.g. cppcheck, language linters).
- IaC (Terraform/Ansible) and cloud runners (AWS/GCP).
- Atlassian suite beyond Jira (Confluence runbooks, Jira + GitHub automation).
- Release hygiene: versioning, artifacts, SBOM/signing, branch protection / merge gates.
- Your pipelines sit under real AI networking product software - not a side CRUD service.
- You influence how fast we ship switch SDK + SAI features that AI clusters depend on.
- Small team, high ownership: changes land and developers feel them the next day.

