This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a AI Researcher - Training Optimization based in Brazil.
This research-focused role offers the opportunity to improve the efficiency, stability, and scalability of large-scale AI model training.
You will work at the intersection of machine learning research and systems engineering to develop practical training optimization techniques.
Your work will target areas such as convergence, training cost, memory efficiency, numerical stability, and model quality.
You will design rigorous experiments, analyze large-scale results, and turn research insights into measurable improvements.
The role also provides opportunities to publish research and contribute to the broader machine learning community.
You will collaborate closely with infrastructure and inference teams to ensure research decisions translate effectively into real-world performance.
This is an ideal environment for an independent researcher who enjoys solving challenging problems and turning novel ideas into practical results.
Accountabilities:
- Design, implement, and evaluate novel training optimization techniques for large-scale neural networks, including optimization algorithms, schedulers, normalization methods, and curriculum strategies.
- Investigate approaches for improving training efficiency, stability, convergence speed, and overall model quality across long training runs and large datasets.
- Research and implement techniques involving optimizer and scheduler innovations, mixed- and low-precision training, memory-efficient training, gradient noise reduction, scaling laws, and convergence analysis.
- Explore training-time regularization and robustness techniques that can improve model performance and reliability.
- Design and execute large-scale experiments, analyze results rigorously, and translate findings into actionable improvements to training systems and methodologies.
- Author or co-author research papers, technical reports, blog posts, and other forms of high-quality technical communication.
- Collaborate with infrastructure and inference engineering teams to connect training decisions with real-world production performance and constraints.
- Independently investigate emerging research directions and contribute to core model training decisions.
- Strong background in machine learning research, with particular expertise in training dynamics, optimization, and large-scale model training.
- Demonstrated experience training large neural networks, including large language models, multimodal models, or other large sequence models.
- Publication experience at machine learning or natural language processing venues such as NeurIPS, ICML, ICLR, ACL, EMNLP, COLM, arXiv, or through equivalent high-quality open research.
- Strong understanding of optimization theory and practice, backpropagation, gradient flow, training stability, distributed training, and large-batch training.
- Strong proficiency in Python and experience with modern machine learning frameworks, particularly PyTorch.
- Ability to independently formulate research questions, design experiments, interpret complex datasets, and reason from empirical evidence.
- Experience with non-standard architectures, such as RNN variants, long-context models, or hybrid systems, is a plus.
- Experience optimizing large-scale GPU training using technologies such as FSDP, ZeRO, or custom kernels is desirable.
- Contributions to open-source machine learning projects or research codebases are advantageous.
- Ability to work effectively in a fast-moving, ambiguous environment while maintaining strong scientific rigor and technical quality.
- Strong communication and collaboration skills, with the ability to explain complex research findings and work effectively across research and engineering teams.
- Full-time opportunity within a research-focused AI environment.
- Remote working arrangement with the flexibility to contribute from locations worldwide.
- Direct influence over core model training strategies and technical decisions.
- Freedom to investigate novel research ideas and publish meaningful research.
- Direct access to large-scale experiments and real-world production constraints.
- Opportunity to work on advanced problems involving training optimization, efficiency, convergence, and model quality.
- Close collaboration with infrastructure and inference engineering teams.
- Opportunity to contribute to research papers, technical reports, and open technical work.
- Small, senior-level team that values deep technical thinking and thoughtful execution.
Requirements:
Benefits:

