{"id":743441,"url":"https://alion.io/job/simile-evaluations-engineering-member-of-technical-staff","title":"Evaluations Engineering - Member of Technical Staff","company":{"id":720753,"name":"Simile","domain":"simile.com","url":"https://alion.io/company/simile-2","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"B","score":75,"open_postings":3,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-25T05:45:01Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"staff","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":200000,"max":400000,"currency":"USD","period":"year","gross":null,"usd_annual":400000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"LLM","optional":false},{"name":"Human-in-the-Loop","optional":true}],"status":"live","first_seen_at":"2026-07-17T00:00:46Z","employer_posted_date":"2026-07-17","last_verified_at":"2026-09-25T21:22:43Z","board_verified":true,"closed_at":null,"days_open":71,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":71},"description":"About the Company\nSimile is The Simulation Company. We simulate human behavior to keep people at the center of the decisions that shape the world. With AI, anyone can create a product, a campaign, a policy, or a script - the bottleneck has moved upstream. The hard question is no longer whether you can create something, but what to create, for whom, and how to bring it to life. Those are fundamentally human decisions, and they shouldn't be left to chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans in an increasingly agentic world. Our mission is to simulate all eight billion people on earth.\nWe launched five months ago. Since then we've grown revenue 5x, built a new foundation model for human behavior that has run tens of millions of simulations for F100 enterprises, trained a first-of-its-kind confidence model that predicts the accuracy of every simulation, and released the first product that lets organizations verifiably predict the future. The world's leading companies use Simile to make business-critical decisions - from consumer leaders like CVS Health and Wealthfront to professional services organizations like Deloitte and Gallup - strategizing product launches, entering new markets, and forecasting earnings calls.\nWe've raised over $200M at a $2B post-money valuation led by Greenoaks, with Index Ventures, Hanabi, A*, Bain Capital Ventures, and CVS Health Ventures. We've grown from a small home in Palo Alto to a global team of 50+, and we're building a team of the best researchers, engineers, designers, and operators in the world. The future is too important to be left to chance.\nAbout the Role\nAs a Member of Technical Staff in Evaluations Engineering, you will build the systems that enable Simile to evaluate whether our simulations of human behavior are accurate, trustworthy, and improving over time.\nYou will work across data and evaluation infrastructure, evaluation execution workflows, backend services, automation, and internal tooling. Your initial focus will include streamlining how evaluations are run across models; strengthening evaluation versioning, data models, and access controls; and automating customer validations, survey operations, and human data workflows.\nEvaluation at Simile presents unusual engineering challenges. Our models predict distributions of human behavior, and the ground truth used to evaluate them can be noisy and heterogeneous. You will partner closely with Evals, Modeling, Product Engineering, and Data Operations to turn complex methods and inputs into systems that are reproducible, scalable, and useful for model development and business decisions.\nIn this role, you will:\nBuild evaluation execution infrastructure: Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.\n\nStrengthen evaluation data systems: Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.\n\nAutomate validation and data collection: Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.\n\nBuild human data workflows: Create labeling and review tools that enable external experts and operators to contribute high-quality judgments to evaluation campaigns.\n\nDevelop evaluation tooling: Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.\n\nRequirements\nMust Haves\nStrong Engineering Fundamentals: Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.\n\nData and Systems Experience: Experience building backend services, data pipelines, automation workflows, and relational data models.\n\nEnd-to-End Execution: Ability to work across data, backend, and interface layers and take ambiguous projects from technical design through deployment and adoption.\n\nEvaluation Judgment: Strong intuition for what makes evaluation infrastructure reliable, including versioning, provenance, reproducibility, holdout integrity, noisy ground truth, and meaningful model comparisons.\n\nML and LLM Fluency: Familiarity with modern model-development and evaluation workflows sufficient to partner effectively with modeling and evaluation researchers.\n\nProduct and User Judgment: Ability to build clear, efficient tools for researchers, engineers, data operators, and other expert users.\n\nOwnership and Communication: A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.\n\nNice to Haves\nWe do not expect one person to have all of these. We are hiring a team with complementary strengths.\nModel-Evaluation Infrastructure: Experience building LLM or ML evaluation systems, benchmark platforms, regression suites, experiment-tracking tools, or model-quality dashboards.\n\nResearch and Internal Tools: Experience developing technical surfaces for ML engineers, researchers, data scientists, or operations teams.\n\nHuman Data Systems: Experience with labeling platforms, expert-review workflows, LLM-as-judge systems, grader calibration, or other human-in-the-loop evaluation methods.\n\nData-Collection Automation: Experience automating surveys, experiments, customer-data ingestion, or other human data collection workflows.\n\nStatistical Fluency: Comfort reasoning about sampling error, uncertainty, calibration, confidence intervals, and distributional metrics.\n\nSensitive Data and Access Controls: Experience designing permissions, auditability, and data-governance systems for human or customer data.\n\nAgentic Engineering: Experience using modern AI coding tools to accelerate development while independently testing and validating their output.\n\nYou might be a great fit if you have worked on ML or evaluation infrastructure, data platforms, backend systems, experiment tracking, research tooling, workflow orchestration, internal tools, or human-data systems. You do not need to have held an “Evals Engineer” title, but you should have several years of experience building reliable production software and be excited to apply that experience to model quality.\nYou do not need to match every bullet. If you do not perfectly see yourself in this JD but believe you would be exceptional at building the measurement layer for behavioral simulation, we would love to hear from you.\nCompensation & Benefits\nAt Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.\nSalary Range: $200,000 - $400,000 USD\nNote: Final offers are based on experience, specialized skills, interview performance, and relevant training.\n\nEquity: Grants are available for eligible roles, subject to board approval.\n\nHealth & Wellness: Comprehensive medical, dental, and vision coverage.\n\nTime Off: Flexible time off policies to support work-life balance.\n\nOur Process\nWe prioritize thoughtful conversations and clear examples of past work. Our hiring journey is designed to help both sides align on fit, working style, and expectations.\nReapplication Policy: To ensure a fair and thorough evaluation for all applicants, Simile observes a 90-day waiting period before reconsidering candidates for the same role.\nCommitment to Diversity & Inclusion\nEqual Opportunity: Simile is an equal opportunity workplace. We welcome applicants of all backgrounds and identities, valuing an environment where everyone can contribute authentically.\nAccommodations: If you require support or reasonable accommodations during the application process due to a disability, please let us know. We are happy to assist.","description_format":"text","description_chars":7869,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity","Flexible time off"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","AI Agents"],"lifecycle":[{"event":"open","at":"2026-09-11T13:26:28Z"}],"liveness":{"score":14,"band":"cold","label":"Long shot","p_open":1,"p_active":0.484,"p_room":0.28,"age_days":70,"expected_fill_days":30,"reasons":["conf:2","win:tail","crowd:"],"computed_at":"2026-09-25T05:45:01Z"},"pay":{"stated_usd_annual":400000,"is_top_pay":true},"html_url":"https://alion.io/job/simile-evaluations-engineering-member-of-technical-staff","json_url":"https://alion.io/job/simile-evaluations-engineering-member-of-technical-staff.json","meta":{"generated_at":"2026-09-26T00:18:05Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":225,"day_limit":5000,"remaining_today":4775,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}