{"id":1975787,"url":"https://alion.io/job/turing-machine-learning-data-engineer","title":"Machine Learning / Data Engineer","company":{"id":13199,"name":"Turing","domain":"turing.com","url":"https://alion.io/company/turing","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":true,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"A","score":89,"open_postings":9,"ghost_share":0,"stale_share":0.444,"repost_share":0,"time_to_fill_p50_days":26,"computed_at":"2026-10-08T05:49:30Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"middle","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"posting_text","remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["São Paulo, Brazil"],"countries":["BR"],"hiring_countries":["BR"],"hiring_countries_total":1,"salary":null,"salary_estimate":null,"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"CI/CD","optional":false},{"name":"GDPR","optional":false},{"name":"Great Expectations","optional":false},{"name":"Harbor","optional":false},{"name":"HIPAA","optional":false},{"name":"Human-in-the-Loop","optional":false},{"name":"Machine Learning","optional":false},{"name":"NER","optional":false},{"name":"OCR","optional":false},{"name":"Pandera","optional":false},{"name":"PCI DSS","optional":false},{"name":"Pytest","optional":false},{"name":"Python","optional":false},{"name":"Reinforcement Learning","optional":false},{"name":"SQL","optional":false},{"name":"Agile","optional":true},{"name":"BigQuery","optional":true},{"name":"Computer Vision","optional":true},{"name":"Edge AI","optional":true},{"name":"GCP","optional":true},{"name":"Google BigQuery","optional":true},{"name":"Google Cloud Run","optional":true},{"name":"LLM","optional":true},{"name":"VLM","optional":true}],"status":"live","first_seen_at":"2026-10-06T19:31:00Z","employer_posted_date":"2026-10-06","last_verified_at":"2026-10-08T21:24:41Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"About Turing\nTuring’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com. \nSenior ML & Data Engineer - Data Quality & Sensitive Data Compliance\nThis is a full-time remote role based in Brazil or Colombia.\nAbout the role\nEnterprise data flows through our connectors, gets processed, and passes through a sanitization layer before anything downstream touches it. Two things have to be true at every step: the data is what we think it is, and no sensitive information - PII, PHI, company identifiable information (CII), or financial data - gets through. You'll own both. You'll do this primarily by building the machine learning that detects sensitive entities in text and image data and replaces them consistently at scale.\nThis is a hands-on IC engineering role with a QA mindset. You'll build the detection models, validation infrastructure, adversarial test sets, and audit processes that let us make strong claims about data quality and de-identification performance - and back them up with evidence. You'll work closely with a senior ML lead, with no client-facing responsibilities.\nWhat you'll do\nData Quality\nRun deep dives into enterprise data to assess quality: topic coherence across connectors, domain depth within connectors, completeness, and consistency\nDesign and automate validation suites for data pipelines - schema checks, completeness, drift detection, and reconciliation across raw → processed → sanitized stages\nSurface and characterize quality issues in ways that engineering and product can act on\nSensitive data compliance (PII / PHI / CII / financial)\nDesign, train, and evaluate ML models (NER and other approaches) that detect sensitive entities across text and image-based documents such as scans, invoices, and presentations\nBuild replacement pipelines that substitute detected entities with coherent alternatives, so the same entity always maps to the same replacement across every file in a corpus and the data stays useful\nRun these algorithms over large volumes of data to prepare it for downstream agentic task building\nBuild adversarial test sets for de-identification across all sensitive data classes: edge cases, obfuscated identifiers, multilingual entities, OCR noise, and formats designed to slip past detectors\nCover company identifiable information specifically - organization names and aliases, domains and email patterns, internal project and system names, org charts, vendor and partner relationships, contract terms, and any combination of details that could re-identify the source enterprise\nCover financial data - account and routing numbers, card numbers, revenue and pricing figures, transaction records, tax IDs, and financial statements\nMeasure and report de-identification performance by data class - entity-level precision and recall, leak rates, false-negative audits, and replacement consistency\nImplement regression gates in CI/CD so no pipeline change ships without passing data quality and sensitive-data checks\nRun sampling-based human-in-the-loop audits and maintain the audit trail as compliance evidence\nPartner with engineering on root-cause analysis when inconsistencies or leaks are found, and drive fixes to closure\nWhat we're looking for\nAbout 4 to 5 years of hands-on machine learning experience, with ML as your primary background\nStrong Python for ML development and data validation (pytest, Great Expectations, Pandera, or similar)\nSolid SQL and experience validating data across pipeline stages\nFamiliarity with sensitive data categories and the relevant standards - HIPAA Safe Harbor for PHI, GDPR/LGPD for PII, PCI DSS for cardholder data, and confidentiality/NDA obligations for company information\nExperience building and testing NER or other ML-based detection systems: building labeled eval sets, computing precision/recall, handling non-determinism\nUnderstanding of re-identification risk - how seemingly innocuous details combine to reveal an organization or individual\nComfort with ambiguity and a fast-moving environment\nA skeptical, detail-oriented approach - you assume things are broken until you've proven otherwise\nNice to have\nComputer vision and OCR experience, especially building, scaling, and evaluating document pipelines for contracts, statements, invoices, presentations, and internal documents\nHands-on experience with financial or healthcare data, including the privacy requirements specific to those industries\nStartup experience\nAuditing LLM or VLM outputs\nSynthetic sensitive-data generation (PII, PHI, company and financial records)\nFamiliarity with the GCP data stack (BigQuery, GCS, Cloud Run jobs) and CI/CD integration\nExperience handling multi-tenant enterprise data with strict customer confidentiality requirements\nCompliance reporting or working with auditors\nWhy this role matters\nOur enterprise customers trust us with their data on the condition that it can never be traced back to them. Every downstream model, dashboard, and customer commitment depends on the data being clean and the sanitization layer being airtight. When you find a leak, you've prevented an incident. When you prove there isn't one, you've earned the trust that lets the rest of the company move fast.\nValues\nWe are client first: We put our clients at the center of everything we do, because their success is the ultimate measure of our value.\nWe work at Start-Up Speed: We move fast, stay agile and favor action because momentum is the foundation of perfection\nWe are AI forward: We help our clients build the future of Al and implement it in our own roles and workflow to amplify productivity.\nAdvantages of joining Turing\nWork at the frontier of AI, helping the world’s leading AI labs improve their most advanced models by building expert datasets, RL environments, and first-of-a-kind benchmarks.\nContribute to leading-edge AI research and showcase your work at top conferences such as ICLR, ICML, and NeurIPS.\nBring frontier AI innovation to the enterprise, applying lessons learned from leading AI labs to solve real-world business challenges.\nCollaborate with and learn from exceptional colleagues with deep AI experience from Google, Meta, Amazon, and other leading technology companies.\nMove at the pace of AI innovation, with the speed, ownership, and impact of a startup.\nTuring is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender identity, sexual orientation, age, marital status, disability, protected veteran status, or any other legally protected characteristics. At Turing we are dedicated to building a diverse, inclusive and authentic workplace and celebrate authenticity, so if you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles.\nFor applicants from the European Union, please review Turing's GDPR notice here.","description_format":"text","description_chars":7826,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Brazil","iso":"BR","kind":"country"},{"name":"Colombia","iso":null,"kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Recruiting & Staffing","Freelance & Talent Marketplaces","AI Training Data & Annotation"],"lifecycle":[{"event":"open","at":"2026-10-06T21:14:48Z"}],"visa":[],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":1,"expected_fill_days":26,"reasons":["conf:1","velocity","win:early"],"computed_at":"2026-10-08T05:49:30Z"},"pay":null,"html_url":"https://alion.io/job/turing-machine-learning-data-engineer","json_url":"https://alion.io/job/turing-machine-learning-data-engineer.json","meta":{"generated_at":"2026-10-09T01:07:16Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1766,"day_limit":5000,"remaining_today":3234,"minute_limit":60,"resets_at":"2026-10-10T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":13199},"rest":"https://alion.io/mcp/rest/get_company?id=13199"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fturing-machine-learning-data-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fturing-machine-learning-data-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fturing-machine-learning-data-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/turing-machine-learning-data-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fturing-machine-learning-data-engineer"}]}