This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Software Engineer - Backend based in United States.
This role is a senior technical leadership opportunity focused on building the infrastructure and backend platforms that power production machine learning and generative AI workloads. You will design and evolve scalable, reliable services that enable advanced models to operate efficiently at global scale. Working closely with ML Scientists, Data Engineers, Product teams, and engineers, you will turn sophisticated models into robust production systems. The role combines hands-on software engineering with architectural leadership, technical strategy, and mentorship. You will also contribute to observability, GPU infrastructure, workload optimization, and cloud cost efficiency. This is an opportunity to shape modern ML infrastructure while influencing engineering practices across a globally distributed organization.
Accountabilities:
- Lead the design, architecture, and evolution of infrastructure, platforms, and backend services supporting machine learning and generative AI workloads.
- Build and operate highly scalable, reliable, and observable distributed systems and microservices for production environments.
- Deploy and manage ML and generative AI workloads using serving technologies such as vLLM, Triton, TorchServe, SageMaker Endpoints, or comparable frameworks.
- Develop cloud-native infrastructure and services on AWS, supporting high-availability workloads across services such as ECS, EKS, Lambda, DynamoDB, S3, IAM, and CloudWatch.
- Establish and improve observability practices using technologies such as OpenTelemetry, Prometheus, Grafana, and CloudWatch.
- Work with vector databases, feature stores, caching systems such as Valkey/Redis, and infrastructure-as-code technologies including CDK, CloudFormation, or Terraform.
- Contribute to GPU infrastructure management, workload scheduling, performance optimization, and cloud cost efficiency.
- Drive complex technical initiatives, influence architectural decisions, and establish engineering standards across distributed teams.
- Partner closely with ML Scientists, Data Engineers, Product teams, and other stakeholders to translate machine learning capabilities into reliable production solutions.
- Mentor engineers and serve as a technical leader, helping elevate engineering quality and accelerate software delivery through AI-assisted development tools.
- 7+ years of professional software engineering experience building and operating large-scale, production-grade distributed systems and microservices.
- Strong proficiency in Python, Go, and/or TypeScript, with deep knowledge of API design, system architecture, scalability, resiliency, and performance optimization.
- Hands-on experience developing CI/CD pipelines, cloud-native applications, and production infrastructure on AWS.
- Strong experience with containerization and orchestration technologies such as Docker, Kubernetes, ECS, or comparable platforms.
- Experience designing and operating highly available production workloads in distributed environments.
- Experience with machine learning or generative AI infrastructure and model-serving technologies is highly valuable.
- Familiarity with observability frameworks and tools such as OpenTelemetry, Prometheus, Grafana, and CloudWatch.
- Experience with vector databases, feature stores, caching technologies, and infrastructure-as-code solutions is preferred.
- Understanding of GPU infrastructure, workload scheduling, performance tuning, and cloud cost optimization is an advantage.
- Proven ability to lead complex technical initiatives, influence architecture, and collaborate effectively with diverse technical and business stakeholders.
- Demonstrated experience mentoring engineers and providing technical leadership within distributed or globally collaborative teams.
- Experience using AI-assisted development tools to improve engineering productivity and accelerate software delivery is preferred.
- Strong communication skills and the ability to operate effectively in a fast-paced, evolving environment.
- Remote position within the United States, with occasional visits to an office for team events or meetings.
- Flexible remote and hybrid work options depending on team and location.
- Estimated base salary ranging from $140,000-$210,000 USD for most U.S. locations.
- Estimated base salary of $157,000-$235,000 USD for Austin, D.C. Metro, non-Bay Area California, Hawaii, Illinois, Massachusetts, New Hampshire, Oregon, Virginia, and Washington.
- Estimated base salary of $166,800-$250,200 USD for the New York City Metro and Kirkland/Seattle areas.
- Estimated base salary of $182,000-$273,000 USD for the Bay Area and Los Angeles.
- Potential eligibility for additional compensation, including corporate bonuses and/or equity awards.
- Medical, dental, and vision insurance.
- 401(k) retirement savings plan.
- Paid sick time, flexible paid time off, and paid holidays.
- Paid parental leave and wellness days.
- Life insurance, short- and long-term disability, and AD&D insurance.
- Mental health and Employee Assistance Program resources.
- Tuition assistance and financial education and advice.
- Adoption, surrogacy, and fertility benefits.
- Dependent daycare and backup care benefits.
- Employee stock purchase plan.
- Opportunities to work remotely with a globally distributed engineering organization.
- Benefits and compensation may vary based on geographic location and applicable eligibility requirements.
Requirements:
Benefits:

