{"id":1202309,"url":"https://alion.io/job/tower-research-capital-machine-learning-performance-engineer-training","title":"Machine Learning Performance Engineer, Training","company":{"id":58857,"name":"Tower Research Capital","domain":"tower-research.com","url":"https://alion.io/company/tower-research-capital","size_band":"501-1000","is_staffing_agency":false,"is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"B","score":74,"open_postings":34,"ghost_share":0,"stale_share":0.824,"repost_share":0,"time_to_fill_p50_days":69,"computed_at":"2026-09-24T05:45:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"middle","employment_type":null,"work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["New York, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":200000,"max":null,"currency":"USD","period":"year","gross":null,"usd_annual":200000},"salary_estimate":null,"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"C++","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"cuDNN","optional":false},{"name":"CUTLASS","optional":false},{"name":"DeepSpeed","optional":false},{"name":"FSDP","optional":false},{"name":"HPC","optional":false},{"name":"InfiniBand","optional":false},{"name":"JAX","optional":false},{"name":"Machine Learning","optional":false},{"name":"Megatron-LM","optional":false},{"name":"NCCL","optional":false},{"name":"NVLink","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"PyTorch C++","optional":false},{"name":"Triton","optional":false},{"name":"XLA","optional":false},{"name":"Kubernetes","optional":true},{"name":"Ray","optional":true},{"name":"Reinforcement Learning","optional":true},{"name":"SLURM","optional":true},{"name":"Time Series Forecasting","optional":true}],"status":"live","first_seen_at":"2026-09-24T19:50:12Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-25T01:06:20Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.\nTower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization. \nEngineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.\nAt Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do - combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.\nAt Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.\nSummary\nYou will bridge the gap between quantitative research and high-performance computing, building and optimizing the systems used to train machine learning models at scale. You will focus on accelerating the end-to-end training lifecycle-from data ingestion and distributed execution to kernel performance and hardware utilization-enabling researchers to iterate more quickly across increasingly complex models and datasets.\nResponsibilities\nTraining Performance and Benchmarking\nBenchmark model-training workloads across CPUs, GPUs, and other accelerator platforms to identify bottlenecks and guide Tower’s compute infrastructure decisions.\nDevelop performance models and standardized benchmarks for measuring throughput, utilization, scalability, and time to convergence.\nDistributed Training Optimization\nDesign and optimize distributed training strategies, including data, tensor, pipeline, and model parallelism.\nImprove communication efficiency across multi-GPU and multi-node environments by optimizing collective operations, topology awareness, and computation-communication overlap.\nEnd-to-End Training Efficiency\nAnalyze and improve the full training pipeline, including data loading, preprocessing, memory management, forward and backward passes, optimizer execution, checkpointing, and experiment recovery.\nIdentify bottlenecks across compute, memory, storage, networking, and interconnects to increase accelerator utilization and researcher productivity.\nGPU Kernel and Framework Development\nDevelop and optimize GPU kernels and performance-critical framework components for quantitative machine learning workloads.\nIntegrate specialized libraries, compilers, and execution techniques to improve throughput, memory efficiency, and numerical performance.\nModel and Numerical Optimization\nApply techniques such as mixed-precision training, gradient accumulation, activation checkpointing, operator fusion, and memory-efficient optimizers.\nEvaluate tradeoffs among training speed, numerical stability, reproducibility, model quality, and infrastructure cost.\nTraining Infrastructure\nPartner with HPC and infrastructure teams to optimize workload scheduling, resource allocation, observability, fault tolerance, and reproducibility across shared compute environments.\nHelp define the architecture and tooling required to support large-scale experimentation across on-premises and cloud-based infrastructure.\nCross-Functional Collaboration\nWork closely with ML Researchers, Quantitative Researchers, HPC Engineers, Systems Engineers, and hardware specialists to translate research requirements into highly efficient training systems.\nQualifications\n3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments.\nDeep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models, compilation paths, autograd systems, and distributed-training capabilities.\nStrong programming skills in Python and C++, with experience developing or optimizing performance-critical systems.\nProven experience with GPU kernel development and optimization using technologies such as CUDA, Triton, CUTLASS, cuBLAS, cuDNN, or related libraries.\nStrong understanding of GPU architecture, including streaming multiprocessor execution, warp scheduling, tensor cores, and the memory hierarchy from registers through HBM.\nExperience with distributed-training technologies and communication libraries such as NCCL, FSDP, DeepSpeed, Megatron-LM, XLA, or equivalent systems.\nProficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, or comparable tracing and profiling platforms.\nUnderstanding of high-performance networking, storage, and accelerator interconnects, including technologies such as InfiniBand, RDMA, NVLink, or NVSwitch.\nDemonstrated ability to benchmark heterogeneous compute platforms and make rigorous, data-driven recommendations about performance, scalability, and cost.\nPreferred Qualifications\nExperience optimizing training workloads for transformer-based, time-series, reinforcement-learning, or other computationally intensive models.\nExperience with cluster orchestration and scheduling technologies such as Kubernetes, Slurm, Ray, or similar platforms.\nFamiliarity with fault-tolerant distributed training, large-scale checkpointing, experiment reproducibility, and GPU-cluster observability.\nPractical experience with specialized accelerators, custom hardware, or compiler technologies for machine learning.\nPrior experience in financial trading is not required.\nAnticipated New York annual base salary of $200,000, plus eligible for discretionary bonus.\nBenefits\nTower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world. \nAt Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive - without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.\nOur benefits include:\nGenerous paid time off policies\nSavings plans and other financial wellness tools available in each region\nHybrid working opportunities\nFree breakfast, lunch and snacks daily \nIn-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more) \nVolunteer opportunities and charitable giving \nSocial events, happy hours, treats and celebrations throughout the year\nWorkshops and continuous learning opportunities\nAt Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work - together.\nTower Research Capital is an equal opportunity employer.","description_format":"text","description_chars":7526,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Continuous learning","Hybrid work"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Asset Management","FinTech","Capital Markets"],"lifecycle":[{"event":"open","at":"2026-09-24T21:34:16Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":69,"reasons":["conf:0","win:early"],"computed_at":"2026-09-25T01:34:55Z"},"pay":{"stated_usd_annual":200000,"is_top_pay":true},"html_url":"https://alion.io/job/tower-research-capital-machine-learning-performance-engineer-training","json_url":"https://alion.io/job/tower-research-capital-machine-learning-performance-engineer-training.json","meta":{"generated_at":"2026-09-25T01:34:55Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1604,"day_limit":5000,"remaining_today":3396,"minute_limit":60,"resets_at":"2026-09-26T00:00:00Z"}}}