{"id":1240378,"url":"https://alion.io/job/crustdata-ml-engineer-intern-summer-2026","title":"ML Engineer Intern | Summer 2026","company":{"id":2397442,"name":"Crustdata","domain":"crustdata.com","url":"https://alion.io/company/crustdata","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Work at a Startup","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"intern","employment_type":"internship","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":8000,"max":14000,"currency":"USD","period":"month","gross":null,"usd_annual":168000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Contrastive Learning","optional":false},{"name":"Machine Learning","optional":false},{"name":"NLP","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false}],"status":"live","first_seen_at":"2026-09-25T16:39:24Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-25T23:21:54Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"About the role\nSkills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification\nWe're building the gateway to the internet for AI agents. Our APIs already power hundreds of customers - and we went from 0 to $10M+ ARR in our first 18 months. Now we need someone who can push the boundaries of what our ML systems can do.\nWe're hiring an ML Engineer Intern to work directly with our founding team on the research and engineering behind our core intelligence layer. Our platform indexes hundreds of millions of professional profiles and company records from across the web. Making that data searchable, matchable, and enriched is an ML problem at its core.\nThis is a 12-week summer internship (June-August 2026). You will not be fetching coffee or watching from the sidelines. You will be researching, training, and shipping models - from paper to prototype to production. Previous interns' work has shipped to customers within weeks.\nWho you are\nCurrently pursuing a Master's or PhD in Computer Science, Machine Learning, NLP, or a related field\nStrong fundamentals in NLP, information retrieval, or entity resolution - through coursework, research, or side projects\nFamiliar with transformer architectures - you've trained or fine-tuned encoder models, not just called APIs\nExperience building retrieval systems, classifiers, or embedding models (in academic or personal projects)\nExposure to contrastive learning, metric learning, or representation learning\nHave used LLMs for structured extraction, classification, or data generation\nStrong Python and PyTorch\nA true grinder - we work very hard\nFounder mentality - someone who wants to build a company someday\nWhat you'll be doing\nYou'll own real ML problems that turn messy, multilingual, web-scale data into structured intelligence. Some example problems:\nA customer searches for \"RevOps professionals\" - you need to return people titled \"Head of Revenue Department,\" \"Revenue Operations Manager,\" and \"VP Sales Operations,\" across English, French, and German\nThree different data sources list what looks like three different companies - but it's actually one. You figure out how to resolve that automatically across millions of records\nGiven raw people data, infer the org chart - who reports to whom, what the team structure looks like, how the engineering org differs from sales\nDetect what technologies a company uses from unstructured signals scattered across the web\nClassify whether a job change was a promotion, lateral move, demotion, or just a title edit - and do it for millions of transitions\nMap raw job titles to canonical titles, seniority levels, and job functions - across dozens of languages and naming conventions\nNice to haves\nPublished research or conference papers (NeurIPS, ICML, ICLR, ACL, EMNLP, etc.)\nExperience with entity resolution or record linkage at scale\nBuilt taxonomy or ontology systems over messy real-world data\nBackground in multilingual NLP or cross-lingual transfer\nOpen-source contributions in NLP/IR\nExperience with distributed training on GPU clusters\nCompensation & perks\n$8,000-$14,000/month (above market rate for SF internships)\n\nHousing stipend for those relocating to SF\n\nDirect mentorship from the founding team - no layers between you and the CEO\nYour work ships to production and reaches real customers\nInterview Process\n30 min video call with the CEO\n45 min technical deep-dive (ML systems design + paper discussion)\nPaid work trial on a real problem (~1-2 weeks)\nOffer","description_format":"text","description_chars":3509,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"master","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","AI Agents"],"lifecycle":[{"event":"open","at":"2026-09-25T16:39:24Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":15,"reasons":["conf:2","win:early","comp:junior"],"computed_at":"2026-09-26T02:08:44Z"},"pay":{"stated_usd_annual":168000,"is_top_pay":false},"html_url":"https://alion.io/job/crustdata-ml-engineer-intern-summer-2026","json_url":"https://alion.io/job/crustdata-ml-engineer-intern-summer-2026.json","meta":{"generated_at":"2026-09-26T02:08:44Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2151,"day_limit":5000,"remaining_today":2849,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}