About the Role
This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context layer that makes AI agents reliable, accurate, and secure for mission-critical business operations. You'll own the systems that keep those agents running fast and reliably in production - from design through deployment - working closely with ML and infrastructure teams to scale inference at increasing concurrency.
What You'll Do
Own inference and model-serving infrastructure end to end, from architecture design through production deployment.
Build and scale systems that enable AI agents to run reliably and efficiently under high concurrency in production environments.
Collaborate with ML and infrastructure teams to ensure seamless integration and drive performance optimization.
Identify infrastructure bottlenecks and lead the engineering effort to resolve them.
What We're Looking For
5+ years of experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
Hands-on experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
Demonstrated ability to optimize production ML systems for latency, throughput, and reliability at scale.
Strong proficiency with containerization and orchestration technologies - Docker and Kubernetes - for deploying ML workloads.
Experience building or maintaining distributed systems that handle concurrent requests and manage resource allocation under load.
Solid command of monitoring, observability, and debugging tooling for production systems (e.g., Prometheus, Grafana, ELK, distributed tracing).
Experience deploying and managing ML systems on cloud platforms such as AWS, GCP, or Azure.
Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.
Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune) is a plus.
Familiarity with real-time or low-latency inference systems, agentic AI pipelines, or enterprise data infrastructure is a plus.
Location
On-site in San Mateo, California, United States. Visa sponsorship is not available for this role.

