394,870open jobs
13,845companies
76,904added this week
Browse all
Salary
$150k – $180k per year
Location
In office (San Francisco)
Seniority
Junior · 2+ years exp
Employment
Internship
Overview
Company
Impact
Profile match
Samba TV measures television viewing through software embedded in smart TVs from major manufacturers. Advertisers use the resulting panel to see which households saw a campaign and to target follow-up advertising on other devices. The company operates across dozens of countries and publishes regular viewership research.

As an Ontology Engineer on Samba TV's Knowledge Graph & Identity team, you will build, maintain, and extend the knowledge graph schemas, derivation pipelines, and graph data models that underpin Samba's measurement and audience intelligence products. Working closely with the Senior Ontologist and peer data scientists, you will implement ontological frameworks in production, contribute to entity resolution and data enrichment pipelines, and help ensure the graph layer remains accurate, consistent, and production-ready.

This is a hands-on technical role. You are expected to write clean, production-quality Python and SPARQL, take ownership of well-scoped graph work streams, and grow your depth in semantic modeling under the guidance of senior team members.

This role reports to the Data Science Manager, Knowledge Graph & Identity.

What You'll Do:

    Ontology Implementation & Validation

    • Implement and extend Samba's RDF/RDFS/OWL ontology schemas in the graph database - adding entity classes, properties, and constraints in a consistent, governed way under the direction of the Senior Ontologist

    • Build and maintain SHACL validation shapes for post-load graph consistency checks; identify and triage data quality and schema violations

    • Support ontology versioning, change log documentation, and consistency checking across schema updates

    • Write efficient, well-structured SPARQL queries and graph traversals to support downstream data science and product use cases

    • Event-to-Ontology Derivation Pipelines

      • Contribute to the event-to-ontology transformation and derivation layer - building PySpark/Databricks pipelines that aggregate raw TV viewership and web activity events into durable graph attributes (genre affinity, brand affinity, topic affinity, viewing summaries, lifecycle signals)

      • Implement derivation logic specified by the Senior Ontologist and data science team; validate outputs against SHACL shapes before graph load

      • Support incremental refresh and update logic aligned with the graph's batch refresh cadence

      • Technical Contribution

        • Write production-quality Python - clean, well-tested, documented, and reusable by teammates

        • Work with PySpark and Databricks to process and transform high-volume data as part of graph pipeline development

        • Apply embedding-based approaches (semantic similarity, vector search) to entity matching and ontology alignment tasks

        • Contribute to team tooling, documentation, and reusable components that improve knowledge graph development efficiency

        • Collaboration & Growth

          • Partner closely with data engineering on pipeline design, data quality, and incremental ingestion patterns feeding the materialized graph substrate

          • Participate in ontology design reviews and cross-functional working groups

          • Work with product and operations teams to understand use case requirements and translate them into graph schema updates

          • Actively develop expertise in W3C semantic web standards, RDF-native graph databases, and entity resolution under the guidance of the Senior Ontologist

Who You Are:

    Must-Haves

    • 2-4 years of hands-on experience in knowledge graph development, semantic data modeling, ontology engineering, or a closely related field

    • Working knowledge of W3C semantic web standards: RDF, RDFS, OWL, and SPARQL - with practical experience querying or building in at least one triplestore or graph database

    • Familiarity with SHACL or equivalent constraint and validation frameworks for graph data quality

    • Strong Python skills - clean, readable, production-quality code with testing and documentation

    • Solid understanding of data modeling fundamentals - entity-relationship design, taxonomies, hierarchies, and how to represent complex real-world relationships in structured form

    • Familiarity with entity resolution or data matching concepts - understanding of why the same real-world entity appears under different identifiers across data sources

    • Bachelor's degree required in Computer Science, Information Science, Mathematics, or a related field; Master's preferred

    • Detail-oriented and proactive about flagging data quality issues and schema inconsistencies

    • Strongly Preferred

      • Hands-on experience with Amazon Neptune or Stardog - or equivalent RDF-native triplestore; exposure to data virtualization (Neptune Orion or Stardog Virtual Graphs) a plus

      • Working knowledge of PySpark and Databricks - particularly for large-scale event aggregation and transformation pipelines

      • Familiarity with embedding models, vector search, or semantic similarity - applied to entity matching, ontology alignment, or knowledge graph enrichment

      • Experience with LLM APIs or RAG-based approaches applied to information extraction, entity disambiguation, or schema mapping

      • Domain knowledge in media, entertainment, or ad tech - content metadata, advertising entities, TV viewership data, or audience/identity data

      • Exposure to identity resolution, probabilistic record linkage, or device graph approaches

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
394,870 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$130k – $180k per year • Remote • 10+ years exp • Master's Degree
Python
AI/ML
AI Agents
DeepSpeed
DPO
Fine-tuning
FSDP
Hallucination
Knowledge Graph
LangChain
LangGraph
LlamaIndex
LLM
LoRA
Multimodal AI
NLP
PEFT
PPO
PyTorch
QLoRA
RAG
Ray
RLHF
SFT
Synthetic Data
Transformers
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Vector
Apply
$74k – $98k per year • Remote • 7+ years exp • Bachelor's Degree
C++
Python
Rust
AI/ML
Knowledge Distillation
LLM
Model Distillation
Quantization
TensorRT
TensorRT-LLM
vLLM
DevOps
FinOps
Kubernetes
Platform Engineering
Apply
$72k – $100k per year • Remote • 8+ years exp • Master's Degree
Python
AI/ML
DPO
Fine-tuning
FSDP
Knowledge Distillation
LLM
Model Distillation
Multimodal AI
PyTorch
Reinforcement Learning
RLHF
Synthetic Data
Apply
$86k – $103k per year • Remote • 6+ years exp • Bachelor's Degree
Java
Python
Scala
SQL
Databases
Apache Hudi
Apache Iceberg
Apache Kafka
Databricks
HBase
Kafka
Trino
AI/ML
Airflow
Flink
Hadoop
Spark
DevOps
AWS
Azure
CI/CD
Kubernetes
Apply
$140k – $180k per year • Remote • 15+ years exp • Bachelor's Degree
Databases
Amazon Redshift
Apache Kafka
BigQuery
Databricks
Google BigQuery
Kafka
LookML
Snowflake
AI/ML
dbt
Flink
DevOps
AWS
Azure
GCP
Apply
$156k – $184k per year • Remote/Hybrid • Internship • 12+ years exp • Bachelor's Degree • Warsaw
Databases
Apache Kafka
BigQuery
Databricks
Google BigQuery
Kafka
Snowflake
AI/ML
Amazon SageMaker
Embeddings
Flink
Kubeflow
MLFlow
Multimodal AI
Semantic Search
Semantic Search
DevOps
Amazon EKS
CloudFormation
Google GKE
Kubernetes
Terraform
Vector
AWS
GCP
Cybersecurity
GDPR
Apply
$67k – $102k per year • Remote/Hybrid • Internship • 5+ years exp • Bachelor's Degree • Warsaw
Python
Python
pySpark
Databases
Amazon Neptune
Databricks
GraphDB
Milvus
Pinecone
Weaviate
AI/ML
GNN
GraphRAG
Knowledge Graph
LangChain
LlamaIndex
LLM
Spark
DevOps
Vector
Apply
Data Scientist 12 days ago
$49k – $138k per year (Estimated) • In office • Internship • 2+ years exp • Bachelor's Degree • Amsterdam
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Claude
LLM
RAG
Semantic Search
Semantic Search
Spark
DevOps
AWS
GCP
Vector
Apply
Data Scientist 16 days ago
$48k – $89k per year • In office • Internship • 3+ years exp • Bachelor's Degree • Warsaw
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Knowledge Graph
LLM
RAG
Semantic Search
Semantic Search
Spark
DevOps
AWS
GCP
Vector
Analytics
A/B Testing
Apply
$29k – $41k per year • Remote/Hybrid • Internship • Bachelor's Degree • Porto
C#
C++
Java
PowerShell
Python
DevOps
CI/CD
Git
Apply
Founding AE 1 day ago
$130k – $180k per year • Equity 0.2–1% • Remote • Full-Time • 3+ years exp • San Francisco
DevOps
GitHub
Apply
$72k – $120k per year • In office • Internship • San Francisco
Apply
AI/SWE Intern 1 day ago
$36k – $120k per year • In office • Internship • San Francisco
JavaScript
Python
TypeScript
AI/ML
Computer Vision
Frontend
React.js
Apply
$120k – $180k per year • Equity 0.1–0.4% • In office • Full-Time • San Francisco
Python
JavaScript
Frontend
React.js
Apply
$120k – $180k per year • Equity 0.1–0.4% • In office • Full-Time • San Francisco
AI/ML
Computer Vision
LLM
Multimodal AI
Apply
See all jobs
This is one of many
394,870 more open roles from verified company boards, updated every day.