{"id":679662,"url":"https://alion.io/job/articul8-machine-learning-engineer","title":"Machine Learning Engineer","company":{"id":669038,"name":"Articul8","domain":"articul8.ai","url":"https://alion.io/company/articul8","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"D","score":40,"open_postings":20,"ghost_share":1,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-30T05:45:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Dublin, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":136000,"max_usd":296000,"period":"year","method":"role_country_seniority_unknown","sample_n":2986},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Docker","optional":false},{"name":"GitHub","optional":false},{"name":"Hadoop","optional":false},{"name":"Kubernetes","optional":false},{"name":"Machine Learning","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"Ray","optional":false},{"name":"Scrapy","optional":false},{"name":"Selenium","optional":false},{"name":"Synthetic Data","optional":false}],"status":"live","first_seen_at":"2026-06-16T15:58:58Z","employer_posted_date":"2026-06-16","last_verified_at":"2026-09-30T06:29:13Z","board_verified":true,"closed_at":null,"days_open":105,"trust":{"level":"ghost","repost_count":0,"flags":["stale","company_stale"],"days_open":105},"description":"About us:\nAt Articul8 AI, we relentlessly pursue excellence and create exceptional AI products that exceed customer expectations. We are a team of dedicated individuals who take pride in our work and strive for greatness in every aspect of our business. We believe in using our advantages to make a positive impact on the world and inspiring others to do the same.\nJob Description:\nWe are seeking machine learning engineers to join our team full-time. As part of your role, you will help us build pipelines of data collection, data extraction, data filtering/synthetic data generation and data analysis. You will own all work related to acquiring high-quality data to power the training of our domain-specific models end to end. You will work closely with other researchers and engineers to empower our next generation of domain-specific models. We value rapid prototyping, iterating, and shipping new systems quickly.\nRequired Qualifications:\nBS/MS/PhD in Computer Science or a related field.\n\nProficiency in at least one deep learning framework, such as PyTorch.\n\nExperience in machine learning projects in text or vision, e.g., has trained machine learning models to tackle a specific problem.\n\nStrong expertise in large stateful distributed systems and data processing.\n\nStrong proficiency in building large-scale data processing pipelines, familiar with distributed workload (e.g., multiprocessing, Ray, Docker, Kubernetes).\n\nProficiency in at least one programming language commonly used in machine learning, such as Python and ability to write clean, maintainable code.\n\nExcellent problem-solving skills and attention to detail, especially when handling data anomalies and biases to further improve data quality.\n\nKey Competencies\nActive Github contributions are a big plus.\n\nExperience in building large-scale datasets.\n\nFamiliar with at least one of the following tools for data crawling (e.g. Scrapy), data collection (e.g., VPNs, Selenium), data processing (e.g., Hadoop, Datasketch).\n\nBuilding bespoke data processing libraries from scratch.\n\nKeeping up with state-of-the-art techniques for preparing AI training data.\n\nOrganizing and meticulously bookkeeping data across multiple clouds, of multiple modalities, and from many sources.\n\nMultilingual which contributes to enriching the language diversity crucial for robust model training.\n\nResponsibilities:\nDesign and develop data processing pipelines, including data extraction, data filtering, data labeling, etc.\n\nImplement machine learning models to improve the quality and diversity of data (especially in the data extraction stage), e.g., quality classifier, document layout model, code verification model, etc.\n\nOwn and lead engineering projects in the area of data acquisition, including web crawling, data ingestion, and processing.\n\nCollaborate with our Applied Research, Technology, and Architecture teams to ensure smooth data flow and system operability.\n\nDevelop and deploy highly scalable distributed systems capable of handling terrabytes of data.\n\nArchitect and implement algorithms for data indexing and search capabilities.\n\nBuild and maintain backend services for data storage, including work with key-value databases and synchronization.\n\nDeploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks.\n\nBy joining our team, you become part of a community that embraces diversity, inclusiveness, and lifelong learning. We nurture curiosity and creativity, encouraging exploration beyond conventional wisdom. Through mentorship, knowledge exchange, and constructive feedback, we cultivate an environment that supports both personal and professional development.\nYour future experience at Articul8 will include continuous learning and growth opportunities as we embark on an exciting journey to disrupt the status quo. If you're excited about joining a team that's passionate about making a difference, we want to hear from you.\nIf you're ready to join a team that's changing the game, apply now to become a part of the Articul8 team. Join us on this adventure and help shape the future of Generative AI in the enterprise.","description_format":"text","description_chars":4143,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Continuous learning","Growth opportunities","Professional development"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["LLM & Generative AI","Foundation Models"],"lifecycle":[{"event":"open","at":"2026-09-10T23:22:45Z"}],"liveness":{"score":5,"band":"cold","label":"Long shot","p_open":1,"p_active":0.186,"p_room":0.28,"age_days":105,"expected_fill_days":30,"reasons":["conf:19","stale_co","ghost","win:tail","crowd:"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/articul8-machine-learning-engineer","json_url":"https://alion.io/job/articul8-machine-learning-engineer.json","meta":{"generated_at":"2026-09-30T06:36:38Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4637,"day_limit":5000,"remaining_today":363,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}