This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior ML Solutions Architect - Token Factory based in France.
Join a fast-growing AI infrastructure team building a serverless platform for running and customizing open-source LLMs in production.
Help customers move from AI prototypes to scalable, reliable production applications without building their own complex inference stacks.
Design optimized inference workflows and customized LLM solutions across multiple models and modalities.
Work hands-on with inference, fine-tuning, evaluation, prompt engineering, and retrieval-augmented generation.
Partner directly with customers to understand technical challenges and translate them into effective AI architectures.
Collaborate closely with product and engineering teams to turn customer feedback into platform improvements.
Work remotely across Europe in an international environment focused on high-impact AI projects, technical ownership, and continuous innovation.
Accountabilities
- Optimize LLM inference workflows across different modalities to deliver measurable business value and meet customer requirements.
- Support customers with supervised and reinforcement-learning-based fine-tuning approaches to improve model quality and performance.
- Design and implement LLM-powered solutions using serverless inference services and served open-source models.
- Build production-ready applications using LLM APIs, including multimodal models covering text, vision, audio, and domain-specific use cases.
- Provide technical guidance on prompt engineering, RAG architectures, model selection, inference optimization, and deployment strategies.
- Guide customers through the transition from proof of concept to production, with a focus on performance, reliability, scalability, and cost efficiency.
- Work closely with product and engineering teams to communicate customer needs, identify platform gaps, and contribute to roadmap development.
- Help customers select appropriate models, inference configurations, and fine-tuning strategies based on their use cases and technical constraints.
- Contribute to improving the platform and its capabilities by sharing practical insights from customer implementations and production workloads.
- 5+ years of professional experience working with ML/AI systems, including at least 2 years focused specifically on LLMs and generative AI.
- Deep understanding of the modern LLM ecosystem, including model architectures, inference approaches, and fine-tuning techniques.
- Hands-on experience running LLMs in production, including deploying and operating inference workloads at scale.
- Strong practical experience with LLM fine-tuning, including supervised fine-tuning, SFT, LoRA, and data preparation or curation; experience with reinforcement-learning-based fine-tuning is a strong advantage.
- Experience building LLM evaluation frameworks, including task-specific benchmarks, offline and online evaluation pipelines, and LLM-as-a-judge approaches.
- Practical experience with modern inference frameworks and ML libraries such as vLLM, SGLang, TensorRT-LLM, or Transformers.
- Experience deploying LLM-powered applications through APIs from providers such as OpenAI or Anthropic, as well as open-source models.
- Strong Python programming skills and the ability to develop practical, production-oriented AI solutions.
- Excellent communication skills, with the ability to explain complex technical concepts clearly to customers, engineers, product teams, and other audiences.
- Experience working with multimodal AI models, such as vision-language or speech models, is a plus.
- Familiarity with DevOps technologies including Docker, Kubernetes, and Git is beneficial.
- Contributions to open-source ML or AI projects are an additional advantage.
- Familiarity with cloud AI platforms such as AWS SageMaker or Bedrock, Google Vertex AI, or Azure ML is welcome.
- Competitive compensation.
- Career growth and ongoing learning opportunities.
- Flexible working arrangements with a high degree of autonomy and ownership.
- Fully remote work opportunity from Europe.
- Collaborative and innovative international working environment.
- Opportunity to work on impactful AI infrastructure and production-grade LLM projects.
- Exposure to advanced technologies across LLM inference, fine-tuning, evaluation, retrieval, and multimodal AI.
- Opportunity to work closely with talented AI, engineering, product, and customer-facing teams.
- Meaningful technical ownership and the opportunity to influence the evolution of an emerging AI platform.
- Inclusive workplace committed to equal employment opportunities and a diverse working environment.
- Reasonable accommodations are available during the application process where required.
- Candidates must be authorized to work in the country where they apply and may need to provide proof of employment eligibility.

