This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff (Software Engineer, Infrastructure) based in the United States.
This role focuses on building and operating the foundational infrastructure behind real-time search, retrieval, model serving, and agent workloads.
You will tackle complex infrastructure challenges that span multiple technical domains rather than being limited to a single platform area.
The position offers opportunities to own production systems end to end, from architecture and implementation through reliability and ongoing operations.
You will improve performance, scalability, availability, and cost efficiency across both online traffic and background workloads.
A key part of the role involves identifying gaps between existing platforms and creating shared abstractions, automation, and tooling that make infrastructure safer and easier to use.
You will also help shape technical direction, lead high-impact programs, and collaborate with infrastructure, product, AI, and security teams.
This is an excellent opportunity for an experienced infrastructure engineer who thrives on ambiguity, cross-cutting problems, and high-impact technical ownership
Accountabilities:
- Own complex, cross-cutting infrastructure problems spanning compute, storage, networking, data, deployment, and reliability, including removing bottlenecks across retrieval and serving paths and resolving failures that cross platform boundaries.
- Design, build, deploy, and operate distributed infrastructure supporting consumer, AI, and enterprise workloads, maintaining ownership throughout the full system lifecycle.
- Identify limitations and gaps across existing infrastructure platforms and develop shared abstractions, automation, and internal tooling that improve usability, safety, and engineering efficiency.
- Improve system performance, availability, scalability, reliability, and cost efficiency across real-time request traffic and background workloads.
- Investigate and resolve complex production incidents that span application, service, and infrastructure layers, translating incident findings into durable architectural and engineering improvements.
- Develop technical solutions for ambiguous and cross-functional infrastructure challenges, balancing reliability, performance, scalability, operational complexity, and development velocity.
- Set technical direction for complex infrastructure systems through architecture decisions, design reviews, technical standards, and long-term engineering strategy.
- Lead high-impact infrastructure programs that require coordination across multiple engineering teams and technical domains.
- Collaborate closely with infrastructure, product, AI, security, and other engineering partners to establish scalable and sustainable technical solutions.
- Mentor engineers, contribute to design and architecture reviews, and help raise engineering standards across the broader organization.
- Develop sufficient technical depth in unfamiliar systems to diagnose problems effectively, make sound architectural decisions, and drive issues through to lasting resolution.
- 4+ years of professional software engineering experience building and operating production backend, platform, infrastructure, or distributed systems.
- Demonstrated experience owning complex production systems end to end, from architecture and implementation through deployment, operations, and continuous improvement.
- Proven ability to deliver sustained technical impact across teams and address problems that extend beyond the boundaries of a single engineering function.
- Strong software engineering skills in Python or another systems/backend language such as Go, Rust, C++, or Java.
- Meaningful experience across at least two infrastructure domains, such as cloud platforms, distributed systems, Kubernetes, storage, databases, networking, data systems, developer infrastructure, or production reliability.
- Strong understanding of software and infrastructure interactions, with the ability to reason across multiple layers of a production technology stack.
- Demonstrated ability to establish technical direction through architecture, design reviews, technical standards, and engineering best practices.
- Strong leadership through influence, including the ability to coordinate complex initiatives, build alignment across teams, and mentor other engineers.
- Ability to investigate complex production incidents, identify underlying causes, and turn short-term fixes into durable technical and architectural improvements.
- Strong problem-solving skills and comfort working with ambiguous requirements and unfamiliar systems.
- Ability to develop technical depth quickly when entering new infrastructure domains or working with systems outside your existing specialization.
- Excellent collaboration and communication skills, with the ability to work effectively across infrastructure, product, AI, security, and engineering teams.
- A strong sense of ownership, curiosity, and willingness to take responsibility for high-impact infrastructure challenges.
- Candidates are encouraged to apply even if they do not meet every listed qualification, particularly when they bring strong infrastructure experience across multiple domains.
- Opportunity to work on foundational infrastructure supporting high-scale search, retrieval, model serving, and AI agent workloads.
- Broad technical scope spanning multiple infrastructure domains, with the flexibility to focus on cross-cutting systems challenges.
- High level of ownership over production systems, from architecture and design through deployment, operations, and continuous improvement.
- Opportunity to influence technical strategy and establish engineering standards across multiple teams.
- Exposure to distributed systems, cloud infrastructure, Kubernetes, storage, databases, networking, data platforms, reliability, and developer infrastructure.
- Collaboration with highly technical teams across infrastructure, AI, product, security, and engineering.
- Opportunity to solve technically challenging problems where performance, latency, scalability, reliability, and cost directly affect the end-user experience.
- Environment that supports technical growth, mentorship, architecture leadership, and learning across unfamiliar systems.
- Opportunity to lead high-impact, cross-functional infrastructure programs and make durable improvements to engineering systems and practices.
- United States-based role with location and work-arrangement details determined by the hiring team.
- Compensation and additional benefits are determined by the hiring company and will be discussed during the recruitment process.

