Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs

Robusta

Robusta is an innovative technology company that specializes in building open-source automation and observability solutions tailored for Kubernetes clusters. Founded in 2021, the company provides a unified platform that automatically gathers context, logs, and graphs during system alerts to help software engineers troubleshoot infrastructure issues faster. By integrating seamlessly with existing tools like Prometheus, Slack, and PagerDuty, the organization streamlines incident response and lowers downtime for modern cloud-native enterprises.

Robusta assists organizations in transitioning to a digital-first approach, crafting unforgettable experiences for their customers. We provide strategy, design, product, and technology services to prominent businesses and brands, utilizing our go-to-market expertise to facilitate seamless customer experiences and enhance conversion rates.

About the Role

We are seeking a highly experienced Senior Data Engineer to lead the technical design, implementation, and delivery of an enterprise-grade, AI-ready Data Lakehouse platform. This role is critical in building the foundational data layer for a large-scale digital transformation initiative that will support AI agents, digital workers, and knowledge graph (ontology) systems.

The ideal candidate will have strong software engineering experience with a focus on data pipeline development, data architecture, and scalable distributed systems. You will play a key role in designing and maintaining robust data infrastructure that enables advanced analytics and AI capabilities.

This position also involves technical leadership, mentoring engineering teams in a collaborative co-building model, and ensuring long-term operational ownership.

Key Responsibilities

  • Lakehouse Architecture & Implementation: Design and deploy a unified Data Lakehouse utilizing the Medallion architecture (Bronze, Silver, Gold) and open table formats (e.g., Delta Lake, Apache Iceberg) on cloud infrastructure hosted within Saudi Arabia.
  • Data Ingestion & Pipeline Engineering: Build reusable, automated ingestion frameworks (batch and streaming) capable of processing both structured data (RDBMS, APIs) and unstructured data (PDFs, policy documents) to feed downstream AI models and semantic reasoning engines.
  • Data Quality & Governance: Implement automated data quality "circuit breakers" (completeness, uniqueness, referential integrity) and end-to-end data lineage tracking frameworks.
  • Optimization: Optimize data processing workflows for performance, scalability, and cost-efficiency.
  • System Monitoring and Maintenance: Monitor and maintain data systems, responding to SEVs or other urgent issues to ensure continuous operations.
  • Security & Compliance: Ensure the platform adheres strictly to NCA (National Cybersecurity Authority) and NDMO (National Data Management Office) standards. Implement AES-256 encryption at rest, TLS 1.2+ in transit, robust Key Management Systems (KMS), and centralized audit logging.
  • Access Control Integration: Design and deploy granular Role-Based Access Control (RBAC) and Attribute-Based Access Control (ABAC), integrating seamlessly with existing enterprise Identity Providers (e.g., Active Directory).
  • Capability Building & Handover: Lead hands-on knowledge transfer sessions, pair-programming with client engineers, creating operational runbooks, and conducting "Game Day" failure simulations to ensure the client’s team is fully ready to operate the platform independently.

Requirements

  • Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • Experience: 5+ years of proven experience in Data Engineering, Distributed Systems, or Big Data Architecture, with at least 2+ years specifically leading Data Lakehouse or Cloud Data Platform implementations.

Technical Skills & Core Technologies:

  • Programming Languages: Proficiency in programming languages such as Python, Java, or Scala.
  • Data Architecture & System Design: Strong expertise in designing data-intensive applications, complex data modeling, and schema design for enterprise environments.
  • Distributed Systems & Lakehouse Technologies: Deep, hands-on experience with distributed processing engines (e.g., Apache Spark, Kafka, Hadoop) and modern open table formats (e.g., Delta Lake, Apache Iceberg, Apache Hudi).
  • ETL/ELT & Orchestration: Experience designing and building robust data pipelines using modern transformation and orchestration tools (e.g., Apache Airflow, Prefect, dbt).
  • Database Ecosystems: Proven track record with relational databases (e.g., PostgreSQL, MySQL), NoSQL platforms (e.g., MongoDB, Cassandra), and distributed SQL query engines like Hive and Trino.
  • Cloud Infrastructure: Proven experience deploying enterprise data solutions on major cloud providers, specifically within localized Saudi cloud regions. Expertise in Oracle Cloud Infrastructure (OCI) or Google Cloud Platform (GCP) is highly preferred, though experience with AWS or Azure is acceptable.
  • Analytical Skills: Strong problem-solving skills with a keen eye for detail and a passion for data.
  • AI/Data Science Enablement: Prior experience building data pipelines optimized for Machine Learning, Natural Language Processing (NLP), vector embeddings, or Knowledge Graphs/Ontologies is highly desirable.
  • Security & Networking: Strong understanding of enterprise network security, Private Endpoints, Identity & Access Management (IAM), and cryptographic key management.
  • Communication: Excellent written and verbal communication skills, with the ability to articulate complex technical concepts to non-technical stakeholders.
  • Leadership Skills: Demonstrated ability to lead technical teams, manage stakeholder expectations, and successfully transition complex systems to internal IT/Data teams.
  • Regulatory Knowledge: Familiarity with Saudi Arabian data compliance frameworks (NCA CCC, NDMO, SDAIA) is highly preferred.

Recommended for you based on this role

Similar stack
Same company
In your city
DevOps
Ansible
ArgoCD
CI/CD
GCP
GitHub Actions
GitOps
Google GKE
Grafana
Kubernetes
Prometheus
Splunk
Terraform
Cybersecurity
SonarQube
Apply
Remote/Hybrid • 10+ year exp • Bachelor's Degree
AI/ML
AI Agents
LLM
DevOps
AWS
Azure
GCP
Apply
Cairo
JavaScript
TypeScript
Frontend
Bootstrap
GraphQL
Next.js
React.js
Redux
Redux Toolkit
Tailwind CSS
Mobile
State Management
DevOps
AWS
CI/CD
Docker
Git
Rest API
QA
Jest
Apply
Cairo
C
Apex
C
GCC
Apex
MuleSoft
DevOps
Octopus Deploy
Marketing
Salesforce
Apply
Cairo
DevOps
AWS
Azure
CI/CD
GCP
Git
Apply
Contractor
C#
SQL
C#
ASP.NET Core
DevOps
Git
Rest API
Apply
Remote
Apex
Apex
Lightning Web Components
Salesforce Data Cloud
Salesforce Flow
AI/ML
AI Agents
Model Context Protocol
RAG
DevOps
Vector
Marketing
Salesforce
Apply
Remote • Contractor • Cairo
Apex
Apex
Lightning Web Components
MuleSoft
Visualforce
DevOps
CI/CD
Git
Rest API
Marketing
Salesforce
Apply
Remote • Contractor • Cairo
Apex
Apex
Lightning Web Components
Visualforce
DevOps
Azure
Azure DevOps
CI/CD
Git
Marketing
Salesforce
Apply
Cairo
Node JS
PHP
Python
JavaScript
PHP
Laravel
Databases
MySQL
RabbitMQ
Redis
Frontend
GraphQL
DevOps
AWS
CI/CD
Docker
Docker Compose
Grafana
Jenkins
Kubernetes
New Relic
Prometheus
WebSockets
Cybersecurity
SOC 2
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
Jeddah