{"id":2190263,"url":"https://alion.io/job/expedock-benchmark-evaluation-engineer","title":"Benchmark & Evaluation Engineer","company":{"id":3785151,"name":"Expedock","domain":"expedock.com","url":"https://alion.io/company/expedock","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"BambooHR","truth_index":null},"role":"Industrial Engineering","role_family":"Industrial Engineering","seniority":null,"employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon Redshift","optional":false},{"name":"BigQuery","optional":false},{"name":"Google BigQuery","optional":false},{"name":"LLM","optional":false},{"name":"LLM Evaluation","optional":false},{"name":"NumPy","optional":false},{"name":"Polars","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Python","optional":false},{"name":"Snowflake","optional":false},{"name":"SQL","optional":false},{"name":"Synthetic Data","optional":false},{"name":"Amazon S3","optional":true},{"name":"Cursor","optional":true},{"name":"dbt","optional":true},{"name":"DuckDB","optional":true},{"name":"Spark","optional":true}],"status":"live","first_seen_at":"2026-10-09T17:42:09Z","employer_posted_date":"2026-10-09","last_verified_at":"2026-10-11T20:20:13Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"Type:Full-time\nLocation:Remote (Philippines)\nSchedule:US Shift\nAbout Expedock\nWe are a tech-enabled workforce augmentation platform leveraging top 1% offshore talent & cutting edge technology to enable businesses to unlock their full potential.\nAbout the Client\nOur client is an innovative AI platform company developing industry-standard benchmarks to evaluate how AI analysts handle complex, production-scale enterprise data. They are remote-first, fast-growing, and building products that bridge natural language analytics, data modeling, and end-to-end data systems.\nWho We Need\nWe are looking for a Benchmark & Evaluation Engineer with a strong mix of data engineering, advanced SQL, and AI evaluation experience. You should have a \"production data\" sensibility-knowing how complex, messy, and multi-system enterprise environments operate-and a bias for rigor and honest scoring.\nWhat You'll Do\nBuild Realistic Data Environments: Design and stand up multi-system enterprise environments (warehouses, lakes, APIs, DBs) at scale across key domains like healthcare, finance, and product analytics. \nEngineer Benchmark Datasets: Generate calibrated, large-scale datasets with realistic noise, seasonality, drift, and synthetic PII to test query scalability and reasoning. \nCurate Task Libraries: Write and verify benchmark tasks, golden SQL, and rubric-scored reasoning keys while eliminating data leakage and ambiguity. \nOwn the Evaluation Harness: Extend grading systems (deterministic checks, LLM judges), maintain regression-tracked leaderboards, and ensure scoring integrity. \nWhat You Need\nNon-Negotiable Qualifications:\nAdvanced SQL & Data Modeling:Expert proficiency in production-grade SQL (CTEs, window functions, query plans, cardinality) and data modeling (star/snowflake schemas, SCDs, referential integrity).\nData Engineering & Systems: Strong Python skills (pandas, NumPy, Polars) and hands-on experience standing up and loading data into at least one major warehouse (Snowflake, BigQuery, Redshift, or Postgres).\nEvals & Benchmark Experience:Proven experience designing evaluation sets or benchmark tasks, including golden-answer/golden-SQL verification, rubric/LLM-judge calibration, avoiding data leakage, and tracking regressions. \nMulti-System Data Environments: Ability to construct realistic multi-system setups (warehouses, data lakes, operational DBs, APIs) with large-scale, semi-structured, or dirty synthetic data. \nNice to Have:\nDomain knowledge in Healthcare/Health Insurance, Finance/FP&A, Product Analytics, or Supply Chain. \nExperience with LLM evaluation frameworks (e.g., SWE-bench, Cursor-style harnesses) or streaming/data-lake stacks (S3, Parquet, dbt, Airflow, Spark, DuckDB).\nCandidate Data & Privacy Notice\nBy submitting your application to Expedock, you acknowledge and consent to the collection, use, and processing of your personal information for recruitment and hiring purposes. Your information will be used to:\nEvaluate your qualifications and suitability for current and future roles\nCommunicate with you throughout the recruitment process Improve our hiring processes and overall candidate experience\nMaintain talent pools for future opportunities, where permitted by law\nWe handle candidate data with care and in accordance with applicable data protection and privacy regulations. Your information will only be accessed by authorized team members and will not be shared with third parties without your consent, unless required by law.","description_format":"text","description_chars":3485,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Health insurance"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Logistics","Recruiting & Staffing"],"lifecycle":[{"event":"open","at":"2026-10-09T17:42:09Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":33,"reasons":["conf:2","win:early"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/expedock-benchmark-evaluation-engineer","json_url":"https://alion.io/job/expedock-benchmark-evaluation-engineer.json","meta":{"generated_at":"2026-10-11T21:47:04Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler_verified","counted_by":"address","units_charged":1,"used_today":10922,"day_limit":null,"remaining_today":null,"minute_limit":300,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3785151},"rest":"https://alion.io/mcp/rest/get_company?id=3785151"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fexpedock-benchmark-evaluation-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fexpedock-benchmark-evaluation-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fexpedock-benchmark-evaluation-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/expedock-benchmark-evaluation-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fexpedock-benchmark-evaluation-engineer"}]}