{"id":1425907,"url":"https://alion.io/job/visasq-data-platform-engineer","title":"Data Platform Engineer","company":{"id":9641,"name":"visasQ","domain":"visasq.co.jp","url":"https://alion.io/company/visasq","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Tokyo, Japan"],"countries":["JP"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":7000000,"max":12000000,"currency":"JPY","period":"year","gross":null,"usd_annual":77220},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Anomaly Detection","optional":false},{"name":"Azure","optional":false},{"name":"Azure Data Factory","optional":false},{"name":"BigQuery","optional":false},{"name":"ChatGPT","optional":false},{"name":"Claude","optional":false},{"name":"Claude Code","optional":false},{"name":"Confluence","optional":false},{"name":"Devin","optional":false},{"name":"GCP","optional":false},{"name":"Google BigQuery","optional":false},{"name":"Human-in-the-Loop","optional":false},{"name":"Jira","optional":false},{"name":"Power BI","optional":false},{"name":"Slack","optional":false},{"name":"SQL","optional":false},{"name":"Anthropic","optional":true},{"name":"Azure AKS","optional":true},{"name":"Azure Cosmos DB","optional":true},{"name":"Azure DevOps","optional":true},{"name":"Docker","optional":true},{"name":"ElasticSearch","optional":true},{"name":"FastAPI","optional":true},{"name":"Gemini","optional":true},{"name":"Hugging Face","optional":true},{"name":"Kubernetes","optional":true},{"name":"LangChain","optional":true},{"name":"LangGraph","optional":true},{"name":"LLM","optional":true},{"name":"Master Data Management","optional":true},{"name":"OpenAI","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Python","optional":true},{"name":"RAG","optional":true},{"name":"Redis","optional":true},{"name":"Scikit-learn","optional":true},{"name":"SQLAlchemy","optional":true},{"name":"Streamlit","optional":true},{"name":"Transformers","optional":true}],"status":"live","first_seen_at":"2026-09-28T23:27:18Z","employer_posted_date":null,"last_verified_at":"2026-09-28T23:27:18Z","board_verified":false,"closed_at":null,"days_open":1,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":1},"description":"VISASQ operates a knowledge platform dedicated to its mission: “We make insightful connections possible.” Connecting businesses with precisely the right expertise across 190+ countries and a network of over 800,000 experts, our Global Expert Network Service (ENS) provides high-precision matching across region and language barriers for professional clients, including management consulting firms and financial institutions.\nIn an era where generative AI has made publicly available web information easily accessible, the value of unstructured first-hand human experience and knowledge has grown exponentially. Currently, vast amounts of data continue to accumulate across multiple products and multi-cloud environments (Azure and GCP).\nHowever, the precision of AI agents relies fundamentally on the quality of the underlying data. Core elements-such as company master identity resolution, employment history accuracy, and compliance verification reliability-directly impact the quality of AI agent decision-making.\nWhile we are incrementally improving data quality within our current structures, we are looking toward building a unified master data platform that consolidates multiple data sources. We are seeking a Data Platform Engineer who can design the data platform architecture that fuels our AI products and lead complex, data-driven design decisions based on empirical evaluation.\nTechnical Environment\nLanguages: Python, SQL\nAI/ML & Orchestration: Azure OpenAI, Gemini, Anthropic, LangChain, LangGraph, Scikit-learn, Hugging Face Transformers\nBackend & Web: FastAPI, Streamlit, SQLAlchemy\nData & Search: Azure Cognitive Search (AI Search), Elasticsearch, Redis, Azure Cosmos DB, PostgreSQL\nInfrastructure & DevOps: Docker, Kubernetes, Azure Functions, Azure DevOps\nProject Management & Collaboration: Slack, Google Meet, Jira, Confluence, esa.io\nAI Coding Assistants / Developer Tools: Claude Code, Devin, ChatGPT\nData Platform Engineer specific technologies\nAzure Data Factory\nMicrosoft Power BI\nBigQuery\nHighlights of the Role\nFundamentally Elevate AI Product Precision. Drive the foundational quality of the data architecture that directly controls the accuracy limits of autonomous AI research agents and matching engines.\n\nEnd-to-End Ownership from Investigation to CTO Decision-Making. Work directly with complex, real-world legacy and multi-cloud data setups. You will own the full lifecycle-from data investigation and architecture design to CTO alignment and implementation.\n\nDirect Business Impact. Quality improvements in company entity resolution and career history directly drive core matching rates (revenue) while minimizing critical compliance risks.\n\nNavigate Complex Post-Acquisition Environments. Gain rare experience solving large-scale data engineering problems arising from combining distinct tech stacks, data models, and cross-border operations following our 2021 global acquisition of Coleman.\n\nResponsibilities\nLead the entire technical lifecycle of our data foundation-from architectural design supporting search, analytics, and compliance, to entity resolution (deduplication), data cleansing, and building mechanisms for continuous data quality assurance.\nData Platform Architectural Design:Design data models across multiple products and multi-cloud environments (Azure/GCP) to ensure reliable data supply for AI products, search engines, analytics, and compliance checks.\nDefine Single Source of Truth (SSOT) strategies and integrate confidence scoring into data models based on empirical findings.\n\nEntity Resolution & Data Cleansing Pipeline:Design and implement multi-stage entity resolution pipelines combining deterministic matching, scoring models, and human-in-the-loop review workflows.\nBuild robust data cleansing processes that enforce idempotency and state rollback capabilities to maintain high data integrity over time.\nEstablish monitoring and anomaly detection mechanisms to ensure stable pipeline operations.\n\nData Quality Research & Problem Solving:Investigate data quality issues, identify root causes within data/codebases, and formulate impactful technical solutions based on business metrics.\n\nTechnical Decision-Making & Stakeholder Alignment:Document architectural designs, research findings, and Architectural Decision Records (ADRs).\nDrive consensus with the CTO, local product teams, and global engineering leaders.\n\nRequirements\nData Engineering & DWH Experience: Proven track record of designing, building, and operating production-grade data pipelines and data warehouse architecture.\nRelational Database Expertise: Deep hands-on experience with SQL execution plan analysis, index optimization, and high-volume batch processing.\nRoot-Cause Data Investigation: Experience cross-analyzing code logic and raw data to resolve data inconsistencies, pipeline outages, and performance bottlenecks.\nTechnical Documentation & Consensus Building: Strong ability to document architectural choices (design docs, research reports, ADRs) and align technical direction with stakeholders.\nBackend Development: Hands-on experience developing backend applications in modern programming languages.\nLanguage Proficiency: Native or bilingual level Japanese proficiency (essential for analyzing complex local data schemas and collaborating with Japanese domain teams).\nNice to haves\nWhile not specifically required, tell us if you have any of the following.\nAI/Search Data Pipelines: Experience providing data foundations for AI/LLM products, RAG systems, or enterprise search platforms.\nMaster Data Management (MDM): Experience with record linkage, entity resolution, deduplication algorithms, and continuous data governance.\nMaster Data Migration: Experience leading multi-source database consolidation, unified schema design, ID mapping, and zero-downtime strangler-pattern migrations.\nDistributed System Data Syncing: Debugging and development experience in hybrid environments with event-driven and batch-oriented data syncs.\nHigh-Precision Domains: Engineering experience in domains requiring extreme data precision (e.g., finance, healthcare, authentication, compliance).\nCloud & Data Stack: Hands-on experience with SQL Server, Azure (AKS, Data Factory), or GCP (BigQuery, Cloud SQL).","description_format":"text","description_chars":6242,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"Japanese","level":"Advanced (C1)","optional":false}]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Professional Services","Sales & Marketing","Business Development","Market Research"],"lifecycle":[{"event":"open","at":"2026-09-28T23:40:27Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":13,"reasons":["seen:0","win:early"],"computed_at":"2026-09-29T05:45:00Z"},"pay":{"stated_usd_annual":77220,"is_top_pay":false},"html_url":"https://alion.io/job/visasq-data-platform-engineer","json_url":"https://alion.io/job/visasq-data-platform-engineer.json","meta":{"generated_at":"2026-09-30T00:55:53Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":604,"day_limit":5000,"remaining_today":4396,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}