{"id":1459621,"url":"https://alion.io/job/firmus-technologies-network-engineer-ai-cluster-commissioning","title":"Network Engineer, AI Cluster Commissioning","company":{"id":63306,"name":"Firmus Technologies","domain":"firmus.ai","url":"https://alion.io/company/firmus-technologies","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"A","score":95,"open_postings":3,"ghost_share":0,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":90,"computed_at":"2026-10-01T05:45:00Z"}},"role":"Networking","role_family":"Networking","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["Melbourne, Australia"],"countries":["AU"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":87000,"max_usd":214000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":359},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Ansible","optional":false},{"name":"BGP","optional":false},{"name":"CI/CD","optional":false},{"name":"HPC","optional":false},{"name":"ISO 27001","optional":false},{"name":"Least Privilege","optional":false},{"name":"Linux","optional":false},{"name":"NCCL","optional":false},{"name":"Python","optional":false},{"name":"SIEM","optional":false},{"name":"SOC 2","optional":false},{"name":"Zero Trust","optional":false},{"name":"InfiniBand","optional":true}],"status":"live","first_seen_at":"2026-09-29T08:24:09Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-01T07:56:20Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"Infrastructure Automation Engineer, AI Cluster Commissioning\nFirmus Technologies\nFirmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.\nFounded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. \nAt Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.\nFirmus AI Cloud\nOur large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.\nIt empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.\nWhy Firmus?\nAs an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.\nWe are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap. \nWork alongside founders and experts in AI infrastructure, energy systems and next-generation compute.\nWhat we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.\nConsidering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.\nRole Summary\nFirmus Technologies is seeking a skilled Network Engineer to join our Commissioning team. This position will play a crucial role in the deployment, commissioning, and configuration of our network designs for AI infrastructure projects. This role offers an exciting opportunity to work at the forefront of AI networking technology and contribute to the growth of AI infrastructure.\nKey Responsibilities\nDeploy and Commission High-Performance Networks\nDeploy, test, and commission low-latency, high-throughput interconnects (e.g.: Ethernet 100/200/400/800 GbE) for AI workloads.\nConfigure, test, diagnose, remediate, and benchmark performance across multi-node clusters, high-speed storage fabrics and parallel computing AI environments.\nWork with partners and vendors to respond to and resolve network issues identified during the bring-up and commissioning of large-scale AI platforms.\nAnalyse logs, run diagnostics and coordinate with internal teams, partners, and vendors as required.\nDevelop and maintain monitoring tools to proactively identify bottlenecks, errors and abnormal behaviours.\nTest configurations, performance, redundant paths and other aspects as set out in the commissioning test plan and acceptance tests.\nEthernet and RDMA Networking\nDesign, monitor and troubleshoot Ethernet fabrics and RDMA-enabled transport layers.\nMaintain network configurations and ensure fabric health and topology visibility using vendor and open source tools.\nPerform verification and acceptance tests for the new network fabric.\nUnderstand various RoCE optimisation protocols and mechanisms (e.g.: SHARP, CollNet) and use performance tools (e.g.: ibstat, perfquery, ib_write_bw, nccl, etc) to monitor fabric health, congestion and link errors.\nEnable and test optimised networking for AI frameworks and work closely with partners and vendors to ensureefficient multi-node communication.\nPerform deep dive diagnostics to resolve layer 1-4 issues across HPC and AI workloads.\nNetwork as Code and Automation\nDevelop and maintain automated network configurations using Infrastructure as Code (IaC) tools (e.g.: Ansible, Netbox, bash and Python scripts).\nImplement CI/CD pipelines for network changes to improve speed, consistency, and auditability.\nAutomate routine tasks such as provisioning, backups and compliance checks.\nProject Management and Stakeholder Management\nSupport the deployment team in their project management and resource allocation for the network portion of AI cluster installations.\nCollaborate and work closely with the Global Operations Centre, Software Defined Infrastructure team, Data Centre Infrastructure team and Solution Architects to integrate new deployments.\nWork closely with both the Firmus Engineering and Operations teams to align network infrastructure with customers’ requirements.\nFacilitate knowledge sharing and communication between teams and create and maintain comprehensive technical documentation.\nMaintain and build strong relationships with key technology partners and vendors and proactively manage and coordinate partner engagement on site.\nNetwork Security\nImplement authentication and access control for fabric and out-of-band management networks.\nImplement secure configurations for RDMA/RoCEv2 fabrics, including multi-tenant workload isolation.\nSupport vulnerability management: coordinate scanning, patching, and remediation tracking for network infrastructure.\nCollaborate with Security and Risk team to enforce policies and respond to security incidents.\nDesign and implement zero-trust network architecture across corporate and AI compute fabric, including segmentation and least-privilege access.\nParticipate in security incident response - detection, containment, root cause analysis, and post-incident reporting - with the Security and Risk team.\nContribute to compliance and audit activity (e.g. ISO 27001, SOC 2) relating to network controls.\nIntegrate network telemetry and logs with SIEM/observability tooling for security monitoring.\nTechnology Expertise\nPhysical network hardware and advanced networking technologies, including NVIDIA Spectrum Ethernet Platform, RDMA over Converged Ethernet (RoCE), DPU/SmartNICs.\nFamiliarity with open-source network operating systems such as Cumulus Linux and Sonic as well as network simulation environments like NVIDIA Air.\nTesting and benchmark tools and processes to validate network topologies, network performance, and acceptance tests.\nProvide technical support and troubleshooting for advanced networking technologies, escalating to vendors as needed.\nSkills & Experience\nBachelor’s degree in network engineering, computer science, or a related technical field.\n5+ years of experience in network engineering.\nExperience in Linux systems especially host network configuration.\nExperience with high performance Ethernet networks, IPv4, IPv6, BGP, RoCE.\nStrong project management skills and experienced in complex technical projects.\nExcellent problem-solving and analytical skills.\nAbility to work independently and as part of a team.\nStrong communication skills, both written and verbal.\nWillingness to undertake international and/or domestic travel for on-site deployments and commissioning as required.\nSolid understanding of advanced networking technologies, particularly those related to AI would be highly advantageous.\nHands-on experience with NVIDIA Spectrum Ethernet Platform and RDMA over Converged Ethernet (RoCE) preferred.\nWillingness to undertake international and/or domestic travel for on-site deployments and commissioning as required.\nHighly Desirable Experiences\nNVIDIA Infiniband networking technology, subnet manager configuration, UFM, multi-tenancy configurations.\nNetwork security, firewall configuration and management, network segmentation, secure network design, zerotrust architecture\nAuthentication and acess controls\nLocation & Reporting\nThis role is based in Australia or Singapore with regular visits to current and future project sites in Australia and SE Asia.\nReport to: Head of AI Cluster Commissioning\nEmployment Basis: Full-time","description_format":"text","description_chars":8472,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Australia","iso":"AU","kind":"country"},{"name":"Singapore","iso":"SG","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Data Centers & Colocation","Cloud Platforms (IaaS & PaaS)","Solar Energy","AI Compute & Inference"],"lifecycle":[{"event":"open","at":"2026-09-29T11:38:24Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":1,"expected_fill_days":90,"reasons":["conf:11","velocity","win:early"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/firmus-technologies-network-engineer-ai-cluster-commissioning","json_url":"https://alion.io/job/firmus-technologies-network-engineer-ai-cluster-commissioning.json","meta":{"generated_at":"2026-10-01T19:40:07Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2330,"day_limit":5000,"remaining_today":2670,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}