{"id":1158669,"url":"https://alion.io/job/proxima-technologies-principal-ml-performance-engineer-gpu-optimization","title":"ML Performance Engineer (GPU Optimization)","company":{"id":77316,"name":"Proxima Technologies","domain":"proxima.ae","url":"https://alion.io/company/proxima-ae","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"C","score":62,"open_postings":11,"ghost_share":0.636,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-07T05:47:15Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":["New York, United States","Boston, United States","Zurich, Switzerland"],"countries":["US","CH"],"hiring_countries":["US"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":162000,"max_usd":293000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":132},"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"C++","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"DeepSpeed","optional":false},{"name":"FSDP","optional":false},{"name":"GCP","optional":false},{"name":"HPC","optional":false},{"name":"Machine Learning","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"PyTorch C++","optional":false},{"name":"TensorRT","optional":false},{"name":"Triton","optional":false},{"name":"XLA","optional":false},{"name":"Kubernetes","optional":true}],"status":"live","first_seen_at":"2026-09-23T18:56:23Z","employer_posted_date":"2026-10-07","last_verified_at":"2026-10-08T00:32:13Z","board_verified":true,"closed_at":null,"days_open":14,"trust":{"level":"ok","repost_count":0,"flags":["company_stale"],"days_open":13},"description":"Principal ML Performance Engineer (GPU Optimization)\nAbout Proxima\nProxima is a frontier AI and data generation company discovering the next generation of proximity therapeutics by making protein interactions programmable. Our platform brings together foundation-model machine learning, a scalable data generation engine, and a partnership track record exceeding $5B in collaborations across the world’s leading biopharma and tech organizations. We’ve recently closed an oversubscribed seed round with an elite group of VCs including DCVC, NVIDIA’s NVentures, AIX, Yosemite among others.\nNeo-1 is our all-atom foundation model that combines state-of-the-art structure prediction and molecular generation in a single system. Neo-1 enables rapid exploration of chemical and structural space for high value, previously intractable targets, and in particular unlocks small molecule proximity therapeutics like molecular glues with AI for the first time.\nIn parallel, we are developing an advanced structural interactomics platform built on proprietary XLMS technology and a lab equipped with next-generation mass spectrometry instrumentation. This platform produces proteome-scale maps of protein interactions and helps identify small molecules that modulate proximity. Together with Neo-1, it creates an integrated system capable of co-folding protein complexes while generating candidate small molecules to influence those interactions.\nProximity-based therapeutics represent one of the most promising frontiers in modern drug discovery with the potential to treat previously intractable diseases and target ‘undruggable’ proteins. We’re building the tech and the team to make that happen. Come join us!\nWhat you'll do\nProfile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning\n\nWrite and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA) when beneficial\n\nScale distributed training across 32-64 nodes, employing FSDP, DeepSpeed, tensor and pipeline parallelism, and mixed precision\n\nReduce inference cost by optimizing memory scaling for large complexes, improving diffusion sampling efficiency, batching ragged inputs, and maximizing throughput across up to 1000 GPUs\n\nManage GPU cluster efficiency on GCP, focusing on scheduling, utilization, spot strategy, and cost reporting\n\nDevelop benchmarks and profiling tools for the research team\n\nWhat we need\nMinimum of 6+ years experience in ML systems, HPC, or performance engineering, with a BS/MS/PhD in CS, EE, or related field\n\nDemonstrated ability to set technical direction beyond coding: selecting infrastructure, influencing research teams, and mentoring engineers\n\nDeep knowledge of PyTorch internals with hands-on experience profiling and fixing real bottlenecks\n\nExperience with CUDA and Triton, skilled at reading Nsight output, and strong understanding of memory bandwidth and occupancy\n\nExperience with distributed training at multi-node scale\n\nStrong proficiency in Python and C++\n\nAble to name a model they made materially faster and quantify the improvement\n\nNice to haves\nExperience in geometric deep learning, equivariant networks, or protein structure models such as AlphaFold, ESM, or RFdiffusion\n\nExperience writing kernels for structure-model primitives, including triangle attention, triangle multiplicative updates, cuEquivariance, or FlashAttention for pair bias\n\nExperience orchestrating large batch inference and managing Kubernetes GPU scheduling","description_format":"text","description_chars":3542,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"},{"name":"New York","iso":null,"kind":"city"},{"name":"Zurich","iso":null,"kind":"city"},{"name":"Boston","iso":null,"kind":"city"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Managed IT Services (MSP)","Systems Integrators","Video Surveillance"],"lifecycle":[{"event":"open","at":"2026-09-23T22:41:31Z"}],"visa":[],"liveness":{"score":46,"band":"ok","label":"Likely open","p_open":1,"p_active":0.511,"p_room":0.9,"age_days":13,"expected_fill_days":26,"reasons":["conf:0","stale_co","win:mid"],"computed_at":"2026-10-07T05:47:15Z"},"pay":null,"html_url":"https://alion.io/job/proxima-technologies-principal-ml-performance-engineer-gpu-optimization","json_url":"https://alion.io/job/proxima-technologies-principal-ml-performance-engineer-gpu-optimization.json","meta":{"generated_at":"2026-10-08T02:32:49Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4390,"day_limit":5000,"remaining_today":610,"minute_limit":60,"resets_at":"2026-10-09T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":77316},"rest":"https://alion.io/mcp/rest/get_company?id=77316"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fproxima-technologies-principal-ml-performance-engineer-gpu-optimization"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fproxima-technologies-principal-ml-performance-engineer-gpu-optimization"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fproxima-technologies-principal-ml-performance-engineer-gpu-optimization"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/proxima-technologies-principal-ml-performance-engineer-gpu-optimization\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fproxima-technologies-principal-ml-performance-engineer-gpu-optimization"}]}