{"id":2084919,"url":"https://alion.io/job/flexion-robotics-ml-engineer-infrastructure","title":"ML Engineer - Infrastructure","company":{"id":673914,"name":"Flexion Robotics","domain":"flexion.ai","url":"https://alion.io/company/flexion","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workable","truth_index":{"grade":"A","score":85,"open_postings":5,"ghost_share":0,"stale_share":0.4,"repost_share":0,"time_to_fill_p50_days":78,"computed_at":"2026-10-09T06:01:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Zurich, Switzerland"],"countries":["CH"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":103000,"max_usd":233000,"period":"year","method":"role_country_seniority_unknown","sample_n":19},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"FSDP","optional":false},{"name":"GCP","optional":false},{"name":"Kubernetes","optional":false},{"name":"NCCL","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"Reinforcement Learning","optional":false},{"name":"Reinforcement Learning","optional":false},{"name":"SLURM","optional":false},{"name":"Ansible","optional":true},{"name":"Configuration Management","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-10-08T11:44:49Z","employer_posted_date":"2026-10-08","last_verified_at":"2026-10-09T22:03:03Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"About Flexion\nAt Flexion, we are building the autonomy stack for humanoid robots. Our mission is to drive the transition from fragile prototypes to real-world deployments of humanoids. We were founded by leading scientists in robot reinforcement learning (ex-Nvidia, ex-ETH Zürich) and backed by leading international VC firms. In just months, we went from our first line of code to deploying real humanoid capabilities with our customers, leveraging simulation and reinforcement learning. Today, we are rapidly expanding the capabilities of our autonomy stack, our customer base, and our team.\nThe role\nWe are looking for an experienced ML engineer to join Flexion’s experienced infrastructure team and take ownership of Flexion’s GPU compute platforms. This is a senior, on-site role with significant scope.\nAt Flexion, we are building the brain for humanoid robots, which involves training foundation models with vast amounts of data on large GPU clusters. You will own the design, bring-up, operation and optimization of performant clusters. You will work with AI engineers to help them optimize their training speed and hardware utilization. You will also influence strategic compute planning and contribute to new tools and platforms for iterating on our AI models efficiently. This will put you at the heart of Flexion’s AI development and allow you to directly impact the execution of our ambitious roadmap. You will closely collaborate with the company’s leadership, engineers of the infrastructure team and AI engineers across the company.\nKey responsibilities\nArchitect, run and continuously improve existing and future cloud-based GPU clusters. Select the best frameworks and tooling to run our clusters efficiently. Work on cluster provisioning, job schedulers and monitoring systems.\nHelp AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML libraries.\nContribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them. \nOptimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offs.\nRaise the bar on engineering practices, including testing, code quality, documentation, and system reliability.\nRequirements\nDegree in Computer Science, Electrical Engineering or Software Engineering (or equivalent practical experience) plus significant industry experience.\nHands-on experience with the training or inference of large models (billions of parameters) on distributed multi-node GPU hardware. This can include bringing up and running the cluster, writing and optimizing training/inference code, building ML pipelines, etc.\nProficiency in Python and working knowledge of PyTorch.\nDeep understanding of distributed training concepts (DDP, FSDP, NCCL).\nExperience with at least one cloud platform (AWS, GCP, Azure or neoclouds) or large-scale on-premises GPU infrastructure.\nExperience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay.\nNice-to-haves\nFamiliarity with profilers (e.g., PyTorch Profiler, Dynolog, HTA, Nsight).\nExperience with high-performance or parallel file systems (e.g., Lustre).\nExperience provisioning compute on multiple cloud providers.\nExperience with infrastructure-as-code and configuration management (Terraform, Ansible).\nBenefits\nCompetitive Compensation\nJoining a leading robotics team & exposure to never-done-before research\nEnergetic, collaborative culture with a bias for action and regular community events\nZurich\nEnhanced pension plan\nRelocation & permit sponsorship\nEnhanced holiday & paid leave perks\nCentral Zürich office with top-tier robotics testing facilities and infrastructure\nSan Franciso\n401(k) with company contributions\nHealth, dental & vision coverage with the flexibility to choose your own plan\nOpen PTO policy & paid company holidays","description_format":"text","description_chars":3887,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"master","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":true,"industries":["Robotic Software & Control Systems","Humanoid Robots"],"lifecycle":[{"event":"open","at":"2026-10-08T11:44:49Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":78,"reasons":["conf:0","win:early"],"computed_at":"2026-10-09T06:01:00Z"},"pay":null,"html_url":"https://alion.io/job/flexion-robotics-ml-engineer-infrastructure","json_url":"https://alion.io/job/flexion-robotics-ml-engineer-infrastructure.json","meta":{"generated_at":"2026-10-10T00:38:03Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1025,"day_limit":5000,"remaining_today":3975,"minute_limit":60,"resets_at":"2026-10-11T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":673914},"rest":"https://alion.io/mcp/rest/get_company?id=673914"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fflexion-robotics-ml-engineer-infrastructure"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fflexion-robotics-ml-engineer-infrastructure"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fflexion-robotics-ml-engineer-infrastructure"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/flexion-robotics-ml-engineer-infrastructure\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fflexion-robotics-ml-engineer-infrastructure"}]}