409,719open jobs
14,193companies
73,646added this week
Browse all
Salary
$180k – $230k per year
Location
In office (San Francisco)
Seniority
Senior · 5+ years exp
Employment
Internship
Overview
Company
Impact
Profile match
Samba TV measures television viewing through software embedded in smart TVs from major manufacturers. Advertisers use the resulting panel to see which households saw a campaign and to target follow-up advertising on other devices. The company operates across dozens of countries and publishes regular viewership research.

As Senior Ontologist on Samba TV's Knowledge Graph & Identity team, you will own the design, development, and governance of the semantic data models and ontological frameworks that sit at the foundation of Samba's knowledge graph. You are the domain authority for how Samba represents and relates the entities that matter most to our business - and you ensure that representation is rigorous, scalable, and aligned with industry standards.

This is a hands-on technical role. You will spend the majority of your time designing ontologies, writing SPARQL, building knowledge graph pipelines, and working closely with data engineering and data science peers to put your models into production. You bring enough breadth in ML and AI to leverage embedding-based and LLM-augmented approaches where they strengthen the graph, and you contribute meaningfully to entity resolution and identity linking work that depends on the semantic layer you define.

This role reports to the Data Science Manager, Knowledge Graph & Identity.

What You'll Do:

    Ontology Design & Governance

    • Own the end-to-end design, development, and versioning of Samba TV's core ontologies in RDF/RDFS/OWL - defining entity classes, properties, hierarchies, and constraints that accurately model Samba's data domain at scale

    • Author and maintain SHACL shapes for post-load graph validation, consistency checking, and data quality enforcement

    • Define and document derived-attribute schemas - genre affinity, brand affinity, topic affinity, lifecycle signals, and viewing summaries - and own the logical definitions that govern how raw events become durable graph attributes

    • Establish ontology design standards, change management processes, and versioning practices; evaluate alignment with W3C standards and relevant industry schemas (Schema.org, EIDR, DDEX, W3C PROV)

    • Lead ontology design reviews with product, data engineering, and data science stakeholders - articulating trade-offs between expressivity, scalability, and query performance clearly

    • Event-to-Ontology Derivation

      • Define the aggregation and scoring logic that transforms raw TV viewership and web activity events into the durable affinities, summaries, and inferred signals that live in the graph

      • Co-own derivation pipeline design with data engineering - specifying transformation logic, intermediate schemas, and validation checkpoints for Databricks/Spark pipelines that feed the materialized graph substrate

      • Reason carefully about what belongs in the graph vs. what should remain virtualized in the data lake - balancing query performance against storage and refresh cost

      • Knowledge Graph Development & AI Integration

        • Build and maintain production-quality knowledge graph pipelines in Python and SPARQL - well-tested, documented, and scalable to Samba's data volumes

        • Design and implement entity resolution and record linkage pipelines that map real-world entities (content titles, devices, audiences, advertisers) to canonical knowledge graph nodes

        • Develop enrichment workflows that integrate third-party data sources (metadata providers, identity vendors, web sources) into Samba's knowledge graph in a consistent, governed way

        • Apply embedding-based and LLM-augmented approaches to ontology mapping, entity disambiguation, and semantic similarity problems

        • Support content and semantic embedding pipelines that feed into the vector store and underpin GraphRAG-based AI solutions

        • Cross-functional Collaboration & Mentorship

          • Partner with data engineering and platform teams to ensure the knowledge graph is integrated, queryable, and production-ready at scale

          • Collaborate with product to translate business requirements into ontological and graph data model decisions

          • Formally mentor Ontology Engineers and junior data scientists on semantic modeling, SHACL design patterns, and graph best practices

          • Lead internal technical talks and workshops on ontology, knowledge graph, and semantic web topics

Who You Are:

    Must-Haves

    • 5-8 years of hands-on experience in ontology engineering, semantic data modeling, or knowledge graph development - with a demonstrable track record of production ontologies at scale

    • Deep expertise in W3C semantic web standards: RDF, RDFS, OWL, SPARQL 1.1, and SHACL - with hands-on experience building and validating graph schemas in a production triplestore (Amazon Neptune, Stardog, GraphDB, Jena, or equivalent)

    • Strong Python - production-quality, well-tested code; comfortable building data pipelines and graph processing workflows

    • First-principles understanding of description logics, ontology design patterns, and the practical trade-offs between OWL expressivity and triplestore scalability

    • Hands-on experience with entity resolution, record linkage, or deduplication at scale - mapping messy, multi-source real-world data to clean ontological representations

    • Bachelor's degree required in Computer Science, Information Science, Computational Linguistics, Mathematics, or a related field; Master's or PhD strongly preferred

    • Strong communicator - able to defend ontological modeling decisions in design reviews and explain trade-offs to non-specialist stakeholders

    • Strongly Preferred

      • Hands-on experience with Amazon Neptune or Stardog - including data virtualization (Neptune Orion or Stardog Virtual Graphs) over data lake sources

      • Experience designing aggregation and derivation logic that converts raw behavioral event data into durable, graph-resident derived attributes

      • Domain knowledge in media, entertainment, or ad tech - TV viewership (ACR/STB), digital audience modeling (device graphs, identity resolution), or ad exposure data

      • Familiarity with industry content and identity schemas: EIDR, Schema.org VideoObject, DDEX, or equivalent

      • Experience with embedding models, vector databases (Milvus, Pinecone, Weaviate), and GraphRAG architectures (LangChain/LlamaIndex)

      • Familiarity with GNN-based approaches to knowledge graph reasoning or entity resolution a plus

      • Working knowledge of PySpark and Databricks for large-scale transformation pipelines

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
409,719 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
In office • 7+ years exp • Bachelor's Degree
Databases
Databricks
AI/ML
AI Agents
LLM
LLM Guardrails
Model Context Protocol
DevOps
AWS
Azure
Platform Engineering
Cybersecurity
Crowdstrike
Wiz
Apply
In office • 6+ years exp
Python
Databases
Amazon Redshift
BigQuery
Databricks
Google BigQuery
Microsoft Fabric
Snowflake
AI/ML
dbt
DevOps
Amazon S3
AWS
AWS Lambda
CI/CD
IAM
Prometheus
Analytics
ETL/ELT
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hong Kong
Python
SQL
JavaScript
Python
FastAPI
Databases
Databricks
AI/ML
Agentic Workflows
AI Agents
AutoGen
CrewAI
Embeddings
LangChain
LangGraph
Multi-Agent Systems
OpenAI
Prompt Engineering
RAG
Frontend
Next.js
React.js
DevOps
Azure
CI/CD
Docker
Kubernetes
OpenShift
Rest API
Apply
In office • 10+ years exp • Bachelor's Degree
Databases
Databricks
Snowflake
AI/ML
RAG
DevOps
AWS
Azure
Cybersecurity
Least Privilege
Analytics
ETL/ELT
Apply
In office • 8+ years exp • PhD
Databases
Databricks
Snowflake
Apply
$156k – $184k per year • Remote/Hybrid • Internship • 12+ years exp • Bachelor's Degree • Warsaw
Databases
Apache Kafka
BigQuery
Databricks
Google BigQuery
Kafka
Snowflake
AI/ML
Amazon SageMaker
Embeddings
Flink
Kubeflow
MLFlow
Multimodal AI
Semantic Search
Semantic Search
DevOps
Amazon EKS
CloudFormation
Google GKE
Kubernetes
Terraform
Vector
AWS
GCP
Cybersecurity
GDPR
Apply
$67k – $102k per year • Remote/Hybrid • Internship • 5+ years exp • Bachelor's Degree • Warsaw
Python
Python
pySpark
Databases
Amazon Neptune
Databricks
GraphDB
Milvus
Pinecone
Weaviate
AI/ML
GNN
GraphRAG
Knowledge Graph
LangChain
LlamaIndex
LLM
Spark
DevOps
Vector
Apply
Data Scientist 13 days ago
$49k – $138k per year (Estimated) • In office • Internship • 2+ years exp • Bachelor's Degree • Amsterdam
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Claude
LLM
RAG
Semantic Search
Semantic Search
Spark
DevOps
AWS
GCP
Vector
Apply
Data Scientist 17 days ago
$48k – $89k per year • In office • Internship • 3+ years exp • Bachelor's Degree • Warsaw
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Knowledge Graph
LLM
RAG
Semantic Search
Semantic Search
Spark
DevOps
AWS
GCP
Vector
Analytics
A/B Testing
Apply
$29k – $41k per year • Remote/Hybrid • Internship • Bachelor's Degree • Porto
C#
C++
Java
PowerShell
Python
DevOps
CI/CD
Git
Apply
$85k – $105k per year • Equity 0–0.1% • In office • Full-Time • San Francisco
Management
Slack
Marketing
HubSpot
LinkedIn
Apply
$260k – $310k per year • Equity 0.1–0.4% • In office • Full-Time • 3+ years exp • San Francisco
Management
Slack
Marketing
HubSpot
Apply
In office • Internship • San Francisco
AI/ML
LLM
Management
Slack
Apply
$65k – $100k per year • Remote • Contractor • San Francisco
AI/ML
Claude
Apply
Coordinator: Docket 10 hours ago
$71k – $95k per year • In office • 2+ years exp • Bachelor's Degree • San Francisco
Apply
See all jobs
This is one of many
409,719 more open roles from verified company boards, updated every day.