{"id":1482961,"url":"https://alion.io/job/bespokelabs-mts-post-training-enterprise","title":"MTS, Post-Training (Enterprise)","company":{"id":669257,"name":"BespokeLabs","domain":"bespokelabs.ai","url":"https://alion.io/company/bespokelabs-2","size_band":"1-10","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Mountain View, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":300000,"max":350000,"currency":"USD","period":"year","gross":null,"usd_annual":350000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":true,"relocation_package":false,"has_equity":false,"technologies":[{"name":"DeepSeek","optional":false},{"name":"Fine-tuning","optional":false},{"name":"Llama","optional":false},{"name":"LLM","optional":false},{"name":"Mistral","optional":false},{"name":"Post-training","optional":false},{"name":"Qwen","optional":false},{"name":"Reinforcement Learning","optional":false},{"name":"Reward Modeling","optional":false},{"name":"Tool Use","optional":false}],"status":"live","first_seen_at":"2026-09-29T20:24:47Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-04T00:44:54Z","board_verified":true,"closed_at":null,"days_open":4,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":4},"description":"About Bespoke Labs\nBespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.\nRecently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.\nBespoke is uniquely positioned to capture a large market share of data and RL environment curation.\nAbout the Role\nThis is a delivery-heavy role focused on standing up our enterprise post-training capability. Demand is inbound, and our goal is to ship custom, high-performing models to 2-3 paying enterprise customers by year-end. Over the first 6-12 months, you will do the execution work required to solve real-world enterprise problems while helping build the scalable product underneath.\nYou will not be running abstract experiments or training models solely for benchmarks. You will sit directly at the intersection of enterprise demand and applied post-training-curating data, building rigorous eval suites, fine-tuning models, and proving their value to enterprise stakeholders. We will measure you on the production impact, robustness, and delivery of models shipped to real users, not on published papers.\nThe thing we care about most is whether you have done this before. If you have post-trained an LLM, shipped it to production users, managed regression risks, and owned the evals from end-to-end, we want to talk.\nWhat You'll Do\nShip enterprise-grade models. Post-train, fine-tune, and align open-weight and proprietary base models for complex business domains, ensuring they perform reliably in production.\n\nBuild bespoke evaluation suites. Define what \"quality\" means for subjective domain-specific tasks, create custom benchmarks, and calibrate LLM judges against human domain experts.\n\nCurate and filter high-impact datasets. Build and run production data flywheels combining real production traces, human labeling, and synthetic augmentation with strict filtering standards.\n\nManage and mitigate regression risk. Rigorously track downstream performance to ensure fine-tuning for new behaviors doesn't silently degrade core capabilities or reasoning.\n\nEngage directly with stakeholders. Sit in front of product and enterprise customers to understand their requirements, translate vague domain preferences into technical eval metrics, and explain model behavior clearly.\n\nDeploy for cost and privacy. Fine-tune open-weight architectures (e.g., Llama, Qwen, Mistral, DeepSeek) to hit strict enterprise latency, cost, and privacy targets.\n\nDirect frontier tools and workflows. Leverage state-of-the-art post-training techniques, preference tuning, and data curation tooling to maximize output quality and delivery speed.\n\nWhat We're Looking For\nA record of shipped models. You have post-trained at least one LLM that was deployed to real users in production, and you can show how you measured its success.\n\nEnd-to-end eval ownership. Demonstrated experience building benchmarks, creating eval datasets, and getting stakeholders to agree on clear metrics for complex or subjective tasks.\n\nDeep understanding of regression risks. You can instinctively explain how training a model on new behaviors impacts existing capabilities and how to prevent it.\n\nCustomer-facing or product empathy. Experience collaborating directly with non-ML stakeholders, enterprise customers, or product managers to turn requirements into model behavior.\n\nStrong software and ML fundamentals. Fluency in modern post-training frameworks, data processing pipelines, and code bases designed for production deployment.\n\nOwnership mindset. You take complete responsibility for the full post-training lifecycle-from raw data to model deployment and failure analysis-without requiring close supervision.\n\nYou May Be a Good Fit If You Also\nHave post-trained conversational or task-oriented assistants (e.g., support agents, multi-turn chat, tool-using agents)\n\nHave built LLM judges or reward models and calibrated them against human raters\n\nHave operated a production data flywheel: traces → labeling → synthetic augmentation → retrain\n\nHave extensive hands-on experience with open-weight models (Llama, Qwen, Mistral, DeepSeek) for cost, latency, or privacy optimization\n\nCome from forward-deployed engineer (FDE), founder, or early-stage startup backgrounds\n\nWhat We Offer\nLocation: Mountain View, CA (Preferred) or San Francisco, CA (Onsite); Remote considered\n\nBase Salary: $300,000 - $350,000 USD / year\n\nAdditional Comp: 25% performance-based bonus + equity\n\nBenefits & Perks:\nHealth, dental, and vision coverage\n\n401(k)\n\nDaily onsite lunch provided\n\nVisa sponsorship and relocation support available\n\nDirect impact on how the industry trains and evaluates agents\n\nWe value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.","description_format":"text","description_chars":4991,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":true,"industries":["Artificial Intelligence","AI Agents"],"lifecycle":[{"event":"open","at":"2026-09-29T21:48:54Z"}],"visa":[{"country":"US","licensed_sponsor":true,"evidence":"H-1B filings in 12 months: 8","filings_12m":8,"filings_prev_12m":1,"green_card_filings_12m":0,"median_offered_wage_usd":160000,"route":null,"cap_exempt":false,"checked_at":"2026-10-03T21:08:04+00:00","sources":["US Department of Labor: LCA disclosure data (H-1B, H-1B1, E-3)"],"filings_for_role_12m":5}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":3,"expected_fill_days":236,"reasons":["conf:6","win:early"],"computed_at":"2026-10-03T05:45:00Z"},"pay":{"stated_usd_annual":350000,"is_top_pay":true},"html_url":"https://alion.io/job/bespokelabs-mts-post-training-enterprise","json_url":"https://alion.io/job/bespokelabs-mts-post-training-enterprise.json","meta":{"generated_at":"2026-10-04T01:35:59Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2071,"day_limit":5000,"remaining_today":2929,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}