371,660open jobs
9,621companies
49,388added this week
Browse all
Location
In office (Cape Town)
Seniority
Senior · 5+ years exp
Employment
Internship
Overview
Company
Impact
Profile match
impact.com (Impact Tech, Inc.) is an enterprise partnership management platform headquartered in Santa Barbara, California. Founded in 2008 (originally as Impact Radius) by Per Pettersen, Wade Crang, Roger Kjensrud, Lisa Riolo, Marc Diana, and Todd Crawford, the company provides a global SaaS platform designed to automate and scale the full partnership lifecycle. Its software streamlines partner discovery, contract management, tracking, fraud protection, and automated cross-border payments across diverse channel types - including affiliate marketing, influencer management, brand-to-brand partnerships, mobile app referrals, and publisher commerce content.

About impact.com

impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey. From affiliates and influencers to content publishers, brand ambassadors, and customer advocates, impact.com empowers brands to drive trusted, performance-based growth through authentic relationships. Its award-winning products- Performance (affiliate), Creator (influencer), and Advocate (customer referral)-unify every type of partner into one integrated platform. As consumers increasingly rely on recommendations from people and communities they trust, impact.com helps brands show up where it matters most. Today, over 5,000 global brands, including Walmart, Uber, Shopify, Lenovo, L’Oréal, and Fanatics, rely on impact.com to power more than 225,000 partnerships that deliver measurable business results.

About the Role

We're seeking a Senior Data Scientist specializing in Product Data Quality to join our Cape Town Data Science team. In this role, you'll own the analytical and technical foundation of product data quality across our ecosystem-spanning catalog hygiene, transaction matching, classification modeling, deduplication, and global product identity. You'll work across both the structured catalog universe and the messier, larger-scale sales transaction universe, building models and infrastructure that power search, recommendations, and business intelligence. This is a high-impact role that demands both analytical depth and strong engineering capabilities: you'll take models from research to production, build scalable data pipelines, and create the monitoring infrastructure that makes our product data foundation trustworthy and continuously improving. Your work will directly influence search relevance, recommendation quality, match rates, and reporting accuracy across the business.

Core Responsibilities

Product classification & taxonomy modeling

  • Develop, deploy, and maintain ML models for automated product categorization and taxonomy assignment across hierarchical category structures.
  • Improve classification accuracy through feature engineering (text, attributes, embeddings), model iteration, and robust evaluation on both catalog and sales transaction data.
  • Monitor production model performance; identify and remediate misclassification patterns that impact search, recommendations, and reporting.
  • Collaborate with category experts and Product teams to refine taxonomy definitions, handle edge cases, and adapt to new product types.

Catalog & sales universe data quality

  • Conduct deep-dive analyses into catalog completeness, consistency, and correctness across retailers, categories, and product attributes.
  • Own data quality analytics for the sales transaction universe -a larger, messier dataset than catalog-measuring match rates, diagnosing gaps (unmatched transactions, misattributed products), and identifying systematic failures.
  • Define and track catalog and transaction health KPIs (attribute coverage, schema compliance, match rates, GPID coverage, freshness); identify root causes and drive remediation.
  • Build monitoring systems and dashboards to track data quality trends across retailers, categories, and time periods.

Global Product ID (GPID) coverage & matching

  • Assess GPID (GTIN/UPC/EAN) coverage and accuracy across both catalog and sales transaction data; identify gaps by category, retailer, and brand.
  • Build and improve matching algorithms to link sales transactions to catalog products, handling missing GPIDs, naming inconsistencies, and category misclassification.
  • Quantify the impact of GPID enrichment and matching improvements on search, deduplication, and reporting accuracy.
  • Partner with external data providers and brands to improve GPID coverage and resolve identifier conflicts.

Deduplication & entity resolution

  • Identify product variants (size, color, packaging) and duplicates within and across retailer catalogs using clustering, entity resolution, embeddings, and similarity-based techniques.
  • Build scalable deduplication pipelines that handle catalog and transaction data at scale; define patterns, heuristics, and ML-based approaches for variant grouping.
  • Measure the impact of deduplication on search quality, recommendation accuracy, and reporting; iterate on models to reduce false positives and improve precision.
  • Support Data Engineering and Platform teams in productionizing deduplication and entity linking infrastructure.

Manufacturer data quality & brand engagement

  • Evaluate the consistency and accuracy of manufacturer-level attributes (brand name, MPN, manufacturer identifiers) across catalogs and transactions.
  • Detect systemic issues at the brand and retailer level; build scorecards and engage brands (via the Tiger Team) to drive data quality improvements.
  • Create feedback loops to measure manufacturer data quality and track progress on remediation initiatives.

Product search & retrieval infrastructure

  • Research and prototype improvements to product search and retrieval pipelines, including vector search, semantic similarity, and embedding-based matching.
  • Explore and implement vector database infrastructure (e.g., FAISS, Pinecone, Weaviate) to support fast, scalable product retrieval and similarity search.
  • Contribute to the design and optimization of retrieval pipelines that combine text, attributes, and embeddings for search and recommendations.
  • Evaluate search relevance and ranking quality; iterate on indexing strategies, query preprocessing, and re-ranking models.

Product graph & relational modeling

  • Build and maintain product graph infrastructure that captures relationships between products, variants, brands, categories, retailers, and transactions.
  • Use graph-based techniques (community detection, link analysis, centrality) to identify product families, detect duplicates, and surface insights on product hierarchies.
  • Partner with Data Platform teams to design scalable graph storage and query patterns (e.g., Neo4j, graph extensions in BigQuery).

Insights, monitoring & reporting

  • Systematically identify, classify, and prioritize product data quality issues; create clear summaries, visualizations, and actionable recommendations for stakeholders.
  • Build and maintain dashboards and recurring reports for key product data KPIs (match rates, GPID coverage, duplicate rates, classification accuracy, attribute completeness).
  • Establish alerting and anomaly detection systems to proactively surface data quality degradation and model performance issues.

Engineering & production deployment

  • Take models and analytics prototypes from POC to production, with or without engineering partnership-owning deployment, testing, monitoring, and iteration.
  • Build robust, scalable data pipelines and ML workflows using production-grade tools and best practices (versioning, CI/CD, testing, observability).
  • Collaborate with MLOps and Data Engineering teams to ensure production readiness: reliability, latency, drift monitoring, and SLOs.

Qualifications

Required

  • Experience: 5+ years in data science, ML engineering, or analytics engineering, with at least 2+ years focused on product data, catalog quality, entity resolution, search/retrieval, or e-commerce/marketplace analytics.
  • Engineering strength: Proven ability to build production-grade data pipelines and deploy ML models independently; strong software engineering fundamentals (code quality, testing, version control, CI/CD).
  • Data quality expertise: Demonstrated experience analyzing and improving large-scale structured data quality (completeness, consistency, accuracy, deduplication, entity resolution).
  • ML & classification experience: Track record building and deploying classification models, ranking systems, or search/retrieval pipelines in production.
  • Technical skills:
    • Strong Python and SQL; proficiency with ML libraries (scikit-learn, XGBoost, LightGBM, PyTorch/TensorFlow) and data manipulation tools (pandas, PySpark).
    • Experience with entity resolution, fuzzy matching, clustering, embeddings, and similarity-based techniques (Levenshtein distance, cosine similarity, nearest-neighbor search).
    • Familiarity with production ML workflows (model versioning, monitoring, evaluation, retraining, A/B testing).
    • Experience with data profiling, anomaly detection, and exploratory analysis at scale.
  • Analytical rigor: Strong foundation in statistics and ML; ability to design experiments, validate models, interpret results, and communicate insights with business context.
  • Stakeholder collaboration: Experience working cross-functionally with Product, Engineering, and business teams; ability to translate technical work into actionable recommendations.
  • Education: Bachelor's in a quantitative field (CS, Statistics, Math, Engineering, or similar); Master's/PhD preferred.

Preferred / Nice to have

  • Experience with vector search and embeddings (sentence transformers, OpenAI embeddings, BERT-based models) and vector databases (FAISS, Pinecone, Weaviate, Milvus, pgvector).
  • Familiarity with search and retrieval systems (Elasticsearch, Solr, semantic search, BM25, hybrid ranking) and understanding how data quality impacts relevance.
  • Experience with graph databases and graph analytics (Neo4j, NetworkX, graph algorithms for clustering and link prediction).
  • Knowledge of NLP techniques for product data (text classification, named entity recognition, attribute extraction, title/description parsing, semantic similarity).
  • Experience with multimodal modeling (combining text, images, and structured attributes for classification or retrieval).
  • Familiarity with global product identifiers (GTIN/UPC/EAN, MPN, SKU hierarchies) and standards organizations (GS1, GDSN).
  • Experience with deduplication and record linkage at scale (blocking strategies, probabilistic matching, hierarchical clustering).
  • Familiarity with GCP tools (BigQuery, Vertex AI, Dataflow, Cloud Run, Looker) and/or Databricks/Spark for large-scale processing and deployment.
  • Exposure to master data management (MDM) or data governance practices in product or catalog contexts.
  • Experience with recommendation systems or understanding how product data quality impacts personalization and ranking.

What sets you apart

  • Product data obsession: You care deeply about data quality and understand how poor catalog hygiene cascades into user experience, business reporting, and operational inefficiencies.
  • Engineering mindset: You don't just build prototypes-you ship them. You write clean, tested, production-ready code and can own the full lifecycle from research to deployment.
  • Detective instincts: You love digging into messy data, finding patterns, and uncovering root causes-whether it's a systematic retailer issue, a subtle duplicate cluster, or a classification edge case.
  • Pragmatic prioritization: You balance comprehensiveness with impact, focusing on the 20% of issues that drive 80% of quality problems and business value.
  • Search & retrieval intuition: You understand how product data powers search and recommendations, and you know how to build infrastructure (embeddings, vector DBs, graphs) that makes these systems work at scale.
  • Stakeholder fluency: You translate messy data findings into clear, actionable recommendations and build trust with brands, retailers, Product, and Engineering teams.
  • Comfort with ambiguity: You thrive in evolving data ecosystems, defining your own quality metrics and technical roadmaps when the problem space is still being shaped.

Benefits and Perks: 

At impact.com, we believe that when you’re happy and fulfilled, you do your best work. That’s why we’ve built a benefits package that supports your well-being, growth, and work-life balance.

  • Flexible Working: OurResponsible PTO policy means you can take the time off you need to rest and recharge. We're committed to a positive work-life balance and provide a flexible environment that allows you to be happy and fulfilled in both your career and your personal life.
  • Health and Wellness: Your well-being is a priority. Our mental health and wellness benefit includes up to12 fully covered therapy/coaching sessions per year, with additional dependent coverage. We also offer a monthly gym reimbursement policy to support your physical health.
  • A Stake in Our Growth: We offer Restricted Stock Units (RSUs) as part of our total compensation, giving you a stake in the company's growth with a 3-year vesting schedule, pending Board approval.
  • Investing in Your Growth: We’re committed to your continuous learning. Take advantage of our free Coursera subscription and our PXA courses.
  • Parental Support: We offer a generous parental leave policy, 26 weeks of fully paid leave for the primary caregiver and 13 weeks fully paid leave for the secondary caregiver.
  • Technology Financial Support: We provide a technology stipend to help you set up your home office and a monthly allowance to cover your internet expenses

impact.com is proud to be an equal opportunity workplace. All employees and applicants for employment shall be given fair treatment and equal employment opportunity regardless of their race, ethnicity or ancestry, color or caste, religion or belief, age, sex (including gender identity, gender reassignment, sexual orientation, pregnancy/maternity), national origin, weight, neurodivergence, disability, marital and civil partnership status, caregiving status, veteran status, genetic information, political affiliation, or other prohibited non-merit factors.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
371,660 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Cape Town
$117k – $177k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Chicago • Dallas
Apex
C#
Java
Python
Apex
Salesforce Data Cloud
AI/ML
Agentforce
AI Agents
Chain-of-Thought
Embeddings
Few-Shot Learning
Fine-tuning
Hallucination
LLM Guardrails
NLP
Prompt Engineering
RAG
RLHF
Semantic Search
Semantic Search
Frontend
GraphQL
DevOps
CI/CD
Git
Analytics
A/B Testing
Marketing
Salesforce
Apply
Senior ML Engineer 1 day ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Full-Time • PhD • Guadalajara
Python
SQL
Databases
Amazon Redshift
Snowflake
Trino
AI/ML
dbt
DevOps
AWS
AWS Lambda
CI/CD
Git
GitHub Actions
Terraform
GitHub
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
$129k – $232k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Huntsville
Ada
C++
Fortran
MATLAB
Python
Cython
Cython
Meson
AI/ML
NumPy
DevOps
CI/CD
GitLab
Analytics
Matplotlib
Apply
$110k – $130k per year • Equity • In office • Internship • 1+ year exp • Seattle
Java
JavaScript
Java
Hibernate
Spring Boot
Databases
Apache Kafka
MySQL
Frontend
Vue.js
DevOps
Docker
GCP
Git
Apply
Senior QA Engineer 12 days ago
$77k – $169k per year (Estimated) • Equity • In office • 5+ years exp • Santa Barbara
Databases
MySQL
DevOps
Git
QA
Cypress
Playwright
Selenium
Apply
$150k – $170k per year • Equity • In office • Internship • 4+ years exp • Bachelor's Degree • Columbus
Python
Databases
Amazon Aurora
ClickHouse
MySQL
PostgreSQL
DevOps
AWS
Terraform
Apply
$140k – $200k per year • Equity • Remote/Hybrid • Columbus
JavaScript
Java
Java
Spring Framework
Databases
Apache Kafka
AI/ML
Claude
Claude Code
DevOps
GCP
Vector
Apply
Equity • In office • Internship • 2+ years exp • Bachelor's Degree • Cape Town
Java
Apply
Lead Data Engineer 6 hours ago
In office • Cape Town
Python
Scala
SQL
Python
pySpark
Databases
Amazon Redshift
Databricks
Snowflake
AI/ML
RAG
Spark
DevOps
Amazon S3
AWS
Azure
Apply
up to $60k per year • Remote • Freelance • Cape Town
Design
Adobe Photoshop
Figma
Apply
Remote • Full-Time • 3+ years exp • Bachelor's Degree • Cape Town
JavaScript
Python
SQL
TypeScript
AI/ML
AI Agents
DevOps
Rest API
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Cape Town • Pune
Apex
Marketing
Salesforce
Apply
Lead Developer 2 days ago
In office • Full-Time • 8+ years exp • Cape Town
PHP
JavaScript
PHP
Drupal
WordPress
Frontend
React.js
Vue.js
DevOps
CI/CD
Apply
See all jobs
This is one of many
371,660 more open roles from verified company boards, updated every day.