{"id":1937689,"url":"https://alion.io/job/sierra-engineer-inference","title":"Software Engineer, Inference","company":{"id":37541,"name":"Sierra","domain":"sierra.ai","url":"https://alion.io/company/sierra-ai","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"A","score":90,"open_postings":74,"ghost_share":0,"stale_share":0.203,"repost_share":0,"time_to_fill_p50_days":104,"computed_at":"2026-10-08T05:49:30Z"}},"role":"Backend","role_family":"Backend","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":230000,"max":390000,"currency":"USD","period":"year","gross":null,"usd_annual":390000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"AI Agents","optional":false},{"name":"Google Maps","optional":false},{"name":"Google Workspace","optional":false},{"name":"OpenAI","optional":false},{"name":"Post-training","optional":false},{"name":"SGLang","optional":false},{"name":"Speculative Decoding","optional":false},{"name":"vLLM","optional":false}],"status":"live","first_seen_at":"2026-10-05T23:30:38Z","employer_posted_date":"2026-10-05","last_verified_at":"2026-10-08T22:07:19Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"About us\nSierra is the leading platform for customer-facing AI agents, working with many of the world's biggest brands - including The GAP, Rocket Mortgage, SoFi, Sutter Health, and SoftBank - to transform how they serve customers and grow their businesses. We are primarily an in-person company based in San Francisco, with growing offices across North America, Europe, and Asia.\nWe are guided by a set of values that are at the core of our actions and define our culture: Trust, Customer Obsession, Craftsmanship, Intensity, and a commitment to balancing Family along the way. These values are the foundation of our work, and we are committed to upholding them in everything we do.\nOur co-founders areBret Taylor and Clay Bavor. Bret currently serves as Board Chair of OpenAI. Previously, he was co-CEO of Salesforce (which had acquired the company he founded, Quip) and CTO of Facebook. Bret was also one of Google's earliest product managers and co-creator of Google Maps. Before founding Sierra, Clay spent 18 years at Google, where he most recently led Google Labs. Earlier, he started and led Google’s AR/VR effort, Project Starline, and Google Lens. Before that, Clay led the product and design teams for Google Workspace.\nAbout the role\nSierra’s AI agents depend on foundation models to reason and act in real time. The Inference team builds the systems that make those models fast, reliable, and efficient at scale.\nAs a Software Engineer on Inference, you’ll help define Sierra’s inference architecture across both self-hosted models and third-party inference providers. You’ll work on the systems responsible for serving and routing inference, managing capacity and quota, and optimizing for latency, reliability, and cost.\nThis is a systems-first role at the intersection of distributed infrastructure and AI. You don’t need to be an ML researcher-we’re looking for engineers who love complex systems problems and are excited to apply that expertise to one of the fastest-moving areas of AI infrastructure.\nWhat you'll do\nPartner with frontier labs and providers. At our scale, we rely on frontier labs, and inference providers to supply capacity, training and inference infrastructure.\n\nShape Sierra’s inference architecture. Design how inference traffic flows across models, infrastructure, and providers, including new serving and proxy layers as Sierra scales.\n\nBuild for low latency and high reliability. Develop systems for routing, failover, capacity management, and quota that keep inference performant and available across large-scale production workloads.\n\nBuild and operate self-hosted inference. Run models on GPU infrastructure, from building containers and operating inference engines to managing the underlying compute capacity.\n\nOptimize inference performance. Work with the Applied Research team on techniques such as speculative decoding and serving-engine optimizations that improve latency, throughput, and cost.\n\nBuild across a hybrid inference stack. Work with both Sierra-managed infrastructure and leading inference platforms, making architectural decisions about where and how workloads should run.\n\nPush the serving stack forward. Work closely with inference providers to tune engines and infrastructure for Sierra’s workloads.\n\nSupport the broader model lifecycle. Contribute to infrastructure that enables post-training while partnering closely with our Models and Agent Runtime teams.\n\nWhat you'll bring\nDeep systems thinking and strong distributed systems fundamentals.\n\nExperience designing, building, and operating large-scale production systems.\n\nStrong judgment around tradeoffs involving latency, reliability, capacity, and cost.\n\nExperience taking ownership of complex infrastructure from architecture through production operation.\n\nExcitement about applying systems expertise to AI infrastructure and learning quickly as the underlying technology evolves.\n\nEven better\nExperience with ML infrastructure, MLOps, or production inference systems.\n\nExperience serving LLMs or other large models at scale.\n\nExperience operating self-hosted inference and GPU infrastructure.\n\nFamiliarity with inference frameworks such as vLLM or SGLang.\n\nExperience with post-training infrastructure or inference-performance optimization.\n\nOur values\nTrust: We build trust with our customers with our accountability, empathy, quality, and responsiveness. We build trust in AI by making it more accessible, safe, and useful. We build trust with each other by showing up for each other professionally and personally, creating an environment that enables all of us to do our best work.\n\nCustomer Obsession: We deeply understand our customers’ business goals and relentlessly focus on driving outcomes, not just technical milestones. Everyone at the company knows and spends time with our customers. When our customer is having an issue, we drop everything and fix it.\n\nCraftsmanship: We get the details right, from the words on the page to the system architecture. We have good taste. When we notice something isn’t right, we take the time to fix it. We are proud of the products we produce. We continuously self-reflect to continuously self-improve.\n\nIntensity: We know we don’t have the luxury of patience. We play to win. We care about our product being the best, and when it isn’t, we fix it. When we fail, we talk about it openly and without blame so we succeed the next time.\n\nFamily: We know that balance and intensity are compatible, and we model it in our actions and processes. We are the best technology company for parents. We support and respect each other and celebrate each other’s personal and professional achievements.\n\nWhat we offer\nWe want our benefits to reflect our values and offer the following to full-time employees:\nFlexible (unlimited) paid time off\n\nMedical, dental, and vision benefits for you and your family\n\nLife insurance and disability benefits\n\nRetirement plan dependent on country of employment\n\nParental leave\n\nFertility and family building benefits through Carrot\n\nLunch, as well as delicious snacks and coffee to keep you energized\n\nDiscretionary benefit stipend giving people the ability to spend where it matters most\n\nFree alphorn lessons\n\nThese benefits are further detailed in Sierra's policies, may vary by region, and are subject to change at any time, consistent with the terms of any applicable compensation or benefits plans. Eligible full-time employees can participate in Sierra's equity plans subject to the terms of the applicable plans and policies.\nBe you, with us\nWe're working to bring the transformative power of AI to every organization in the world. To do so, it is important to us that the diversity of our employees represents the diversity of our customers. We believe that our work and culture are better when we encourage, support, and respect different skills and experiences represented within our team. We encourage you to apply even if your experience doesn't precisely match the job description. We strive to evaluate all applicants consistently without regard to race, color, religion, gender, national origin, age, disability, veteran status, pregnancy, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.","description_format":"text","description_chars":7253,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity","Life insurance","Parental leave","Retirement plans"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Commerce","LLM & Generative AI","AI Agents"],"lifecycle":[{"event":"open","at":"2026-10-06T00:18:00Z"}],"visa":[],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":2,"expected_fill_days":104,"reasons":["conf:0","velocity","win:early","comp:brand"],"computed_at":"2026-10-08T05:49:30Z"},"pay":{"stated_usd_annual":390000,"is_top_pay":true},"html_url":"https://alion.io/job/sierra-engineer-inference","json_url":"https://alion.io/job/sierra-engineer-inference.json","meta":{"generated_at":"2026-10-09T00:58:30Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1494,"day_limit":5000,"remaining_today":3506,"minute_limit":60,"resets_at":"2026-10-10T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":37541},"rest":"https://alion.io/mcp/rest/get_company?id=37541"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fsierra-engineer-inference"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fsierra-engineer-inference"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fsierra-engineer-inference"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/sierra-engineer-inference\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fsierra-engineer-inference"}]}