430,068open jobs
14,618companies
59,434added this week
Browse all
Salary
$18k – $44k per year (Estimated)
Location
Remote/Hybrid (India)
Seniority
Middle
Employment
Full-Time
Overview
Company
Impact
Profile match
irth Solutions is a provider of cloud-based software for 811 ticket management, asset protection, mobile workforce management, and no-code app creation. irth Solutions software is used by industries such as construction, telecommunications, utilit...

About Irth Solutions

Irth Solutions is a leading provider of cloud-based SaaS software for damage prevention, asset integrity, stakeholder engagement and land management, helping energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. With nearly three decades of industry experience, Irth serves customers across North America and continues to expand its platform with new data-driven and AI-powered capabilities.

Data Engineer - Insights (AI/ML)

Location: Remote - India

Department: Insights (AI/ML)

Reports to: Data Platform & Analytics Manager

About the Role

Irth is building a modern, multi-cloud, enterprise-grade data estate-a unified Databricks-based data platform that centralizes data across Irth’s products and cloud environments, including AWS, Azure, and GCP.

As a Data Engineer, you will play a hands-on implementation role, working closely with the Senior Data Architect to bring the enterprise data platform vision to life.

You will design and develop data pipelines based on established architectural patterns, implement data quality and governance controls, build Delta Lake and medallion architecture solutions, and help operationalize the new data platform.

This is an excellent opportunity for a mid-level Data Engineer looking to deepen their expertise in Databricks, Apache Spark, cloud data engineering, and modern lakehouse architecture while working in a multi-cloud enterprise environment.

Key Responsibilities

1. Data Pipeline Development - Primary Responsibility

  • Build, maintain, and enhance data ingestion pipelines across AWS, Azure, and GCP, following architecture and engineering patterns established by the Senior Data Architect.
  • Develop both batch and streaming pipelines using:
    • Databricks Workflows
    • Apache Spark / PySpark
    • SQL
    • Delta Live Tables
    • Databricks Lakeflow components
  • Implement Bronze → Silver → Gold medallion architecture patterns for ingestion, transformation, cleansing, and standardization.
  • Implement Change Data Capture (CDC) and Slowly Changing Dimensions (SCD Type 1 and Type 2).
  • Handle schema evolution and changing source-system structures.
  • Implement data validation, reconciliation, and quality rules as part of pipeline processing.
  • Build reusable and maintainable pipeline components following established engineering standards.

2. Platform & Storage Implementation

  • Configure and maintain Delta Lake storage structures, tables, schemas, partitions, and optimization routines.
  • Apply Delta Lake performance and maintenance practices, including:
    • OPTIMIZE
    • Z-ORDER
    • VACUUM
    • Appropriate partitioning and file-management strategies
  • Assist with implementation of metadata, cataloging, and lineage standards using Unity Catalog.
  • Support integration between cloud storage platforms and Databricks, including:
    • Amazon S3 → Databricks
    • Azure Storage → Databricks
    • Google Cloud Storage → Databricks
  • Assist with implementation of scalable storage and processing patterns defined by the Data Architect.

3. Data Governance, Quality & Compliance Enablement

  • Implement automated data-quality checks, profiling, validation, and monitoring in accordance with enterprise governance standards.
  • Apply data-quality rules at appropriate stages of the Bronze, Silver, and Gold layers.
  • Implement RBAC policies, security controls, and data-classification tags defined by the enterprise governance model.
  • Support implementation of metadata and lineage mapping across Unity Catalog and Microsoft Purview.
  • Help ensure datasets are properly documented, classified, governed, and discoverable.
  • Support remediation of data-quality and governance issues identified through monitoring or reviews.

4. Orchestration, Automation & Operational Support

  • Build, schedule, monitor, and maintain production workflows using:
    • Databricks Workflows
    • Delta Live Tables
    • Azure Data Factory (ADF)
    • Other approved orchestration tools
  • Contribute to CI/CD pipelines for data-engineering code, including source control, automated testing, deployment, and environment management.
  • Support DEV → QA → PROD promotion processes.
  • Monitor production pipelines and respond to failures and data-quality issues.
  • Troubleshoot failed jobs, investigate root causes, and support pipeline recovery.
  • Perform performance tuning across Spark jobs, SQL workloads, Delta tables, and data pipelines.
  • Participate in operational improvements that increase pipeline reliability, scalability, and cost efficiency.

5. Collaboration & Documentation

  • Work directly with the Senior Data Architect to translate architecture designs and technical standards into actionable implementation tasks.
  • Participate in architecture reviews, technical design discussions, coding reviews, and engineering standards meetings.
  • Collaborate with Data Scientists, ML Engineers, Analysts, Product teams, and other engineering stakeholders to understand data requirements.
  • Document:
    • Data pipelines
    • Data flows
    • Data dictionaries
    • Transformation logic
    • Data-quality rules
    • Test cases
    • Job schedules
    • Operational procedures
  • Maintain clear and accurate technical documentation to support platform adoption, troubleshooting, and future development.
  • Provide implementation feedback to the Data Architect and identify opportunities to improve platform patterns, tooling, and developer experience.

Role Scope

This is primarily an implementation-focused Data Engineering role. The Senior Data Architect will establish the overall platform architecture, standards, and design patterns; the Data Engineer will translate those patterns into reliable, production-ready pipelines and platform capabilities.

The role provides an opportunity to gain deeper hands-on experience with Databricks, Spark, Delta Lake, Unity Catalog, cloud data platforms, data governance, and multi-cloud lakehouse engineering while contributing to a strategic enterprise data platform.

Requirements

Qualifications

Required Qualifications

  • 3-5 years of experience in Data Engineering, ETL development, or cloud data platform engineering.
  • Hands-on experience with Databricks, Apache Spark, PySpark, or other distributed data-processing technologies.
  • Strong proficiency in SQL, including structured data transformation, joins, aggregations, and performance-aware query development.
  • Experience working with at least one major cloud platform, with Microsoft Azure preferred; AWS and/or GCP experience is also valuable.
  • Understanding of core data-engineering concepts, including:
    • Data modeling
    • Data quality
    • Schema evolution
    • Data validation
    • Pipeline monitoring and troubleshooting
  • Basic understanding of data-security practices, including:
    • Role-Based Access Control (RBAC)
    • Encryption
    • Credential and secret management
    • Secure access to cloud and data-platform resources

Preferred Qualifications

  • Hands-on or working knowledge of Delta Lake, medallion architecture, and modern lakehouse best practices.
  • Experience with metadata, cataloging, and governance platforms such as:
    • Unity Catalog
    • Microsoft Purview
    • AWS Glue Data Catalog
    • Similar enterprise metadata and data-governance tools
  • Experience with workflow orchestration and scheduling technologies, such as:
    • Azure Data Factory (ADF)
    • Databricks Workflows
    • Apache Airflow
    • Databricks Jobs / DBX
    • Similar orchestration frameworks
  • Experience with Git-based development, CI/CD, and DevOps practices.
  • Knowledge or experience in one or more of the following areas:
    • Geospatial/GIS data
    • BI semantic layers, particularly Power BI
    • Data preparation for AI/ML workloads
  • Relevant cloud or Databricks certifications, such as:
    • Databricks Data Engineer Associate
    • Microsoft Azure Data Engineer Associate (DP-203)
    • Equivalent cloud or data-engineering certifications

Nice-to-Have Qualifications

  • Understanding of asset integrity management concepts, including inspection data, risk scoring, corrosion tracking, defect management, and maintenance data as applied to pipeline or utility operations.
  • Previous experience working with or integrating oil & gas, utility, infrastructure, or pipeline asset data into enterprise data platforms.
  • Experience working with:
    • Pipeline and facility data
    • GIS/geospatial asset data
    • Inspection and maintenance records
    • Asset-risk datasets
  • Familiarity with regulatory, compliance, and audit-reporting requirements associated with pipeline, utility, or asset-integrity data.

Success Metrics

Success in this role will be measured by the engineer’s ability to reliably implement and operationalize the data-platform patterns established by the Data Architect.

Key measures include:

  • High-quality implementation of ingestion, transformation, data-quality, and governance patterns defined by the Data Architect.
  • Reliable and maintainable pipelines supporting consistent Bronze → Silver → Gold data flows.
  • Strong adherence to cataloging, metadata, lineage, security, and data-governance standards.
  • Reduction in pipeline failures and production incidents through improved monitoring, testing, troubleshooting, and operational practices.
  • Continuous improvements in pipeline performance, scalability, reliability, and maintainability.
  • Clear and complete technical documentation covering pipelines, transformations, data-quality rules, and operational procedures.
  • Effective collaboration with the Senior Data Architect, engineering teams, product teams, and business stakeholders.
  • Demonstrated ability to take architecture guidance and translate it into production-ready, scalable data-engineering solutions.

Benefits

Benefits

  • Competitive Salary - A competitive compensation package based on experience and qualifications.
  • Medical, Dental, and Vision Insurance - Comprehensive insurance coverage to support you and your family.
  • 401(k) Plan with Company Match.
  • Generous Paid Time Off (PTO) - Time off to support work-life balance and personal needs.
  • Company-Paid Holidays - Paid holidays throughout the year.
  • Flexible Work Options - Work-from-home opportunities are available, depending on role and business needs.
  • On-Call Compensation - Additional pay for eligible on-call shifts.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
430,068 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$96k – $130k per year • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Calgary • Edmonton • Houston
Python
SQL
Python
pySpark
Databases
Databricks
AI/ML
Spark
MLFlow
Computer Vision
Anomaly Detection
Time Series Forecasting
Apply
$149k – $202k per year • In office • Full-Time • New York
Python
JavaScript
Java
SQL
Java
Spring Boot
AI/ML
Copilot
Cursor
LangGraph
LangChain
Claude
Hadoop
Spark
Claude Code
Scikit-learn
AI Agents
TensorFlow
PyTorch
CrewAI
Google ADK
AWS Strands Agents
Frontend
React.js
DevOps
Git
AWS
Docker
Kubernetes
GitHub
Apply
$27k – $63k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru • Mumbai
Rust
SQL
Ruby
Rust
Actix Web
Diesel
Rocket
Ruby
Ruby on Rails
Databases
PostgreSQL
DevOps
Rest API
GCP
AWS
Docker
Kubernetes
Amazon EKS
Google GKE
Apply
$102k – $209k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • New York
DevOps
GCP
Azure
CI/CD
Jenkins
AWS
GitHub
Cybersecurity
ISO 27001
SOC 2
GDPR
Fortify
Mend
Analytics
Power BI
Management
Power Automate
Power Apps
Apply
Risk Modeling Intern 4 hours ago
$36k – $55k per year (Estimated) • Remote/Hybrid • Internship • Raleigh
Python
SAS
Databases
Snowflake
DevOps
AWS
Analytics
Microsoft Excel
Apply
Quality Analyst 1 day ago
$75k – $179k per year (Estimated) • Remote • Full-Time
Python
JavaScript
SQL
Databases
Databricks
AI/ML
Copilot
Cursor
Claude
Spark
Claude Code
AI Agents
DevOps
Rest API
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
GitHub
Analytics
ETL/ELT
QA
Selenium
JMeter
Cypress
Playwright
Postman
Rest-Assured
k6
Locust
Apply
$22k – $58k per year (Estimated) • Remote/Hybrid • Full-Time
Python
SQL
Databases
PostgreSQL
Databricks
Delta Lake
PostGIS
AI/ML
Spark
Quantization
Prompt Engineering
Knowledge Distillation
NER
Great Expectations
LLM
RAG
Hallucination
Anomaly Detection
LLMOps
LLM Evaluation
LLM Guardrails
Model Distillation
DevOps
GitHub Actions
Azure
CI/CD
AWS
Vector
FinOps
Incident Management
GitHub
Analytics
Power BI
Management
Jira
Apply
$21k – $57k per year (Estimated) • Remote/Hybrid • Full-Time
Python
SQL
Python
pySpark
Databases
PostgreSQL
Databricks
Delta Lake
PostGIS
DynamoDB
Microsoft Fabric
AI/ML
Spark
MLFlow
Prompt Engineering
NLP
LLM
RAG
Anomaly Detection
Time Series Forecasting
LLM Guardrails
DevOps
GitHub Actions
Azure
CI/CD
AWS
Vector
FinOps
SLI/SLO/SLA
GitHub
Amazon S3
Cybersecurity
ISO 27001
SOC 2
GDPR
Microsoft Entra ID
Analytics
Power BI
A/B Testing
Management
Jira
Apply
$88k – $168k per year (Estimated) • Remote • Full-Time
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Copilot
Cursor
Claude
Spark
Claude Code
AI Agents
Human-in-the-Loop
DevOps
Azure
CI/CD
Git
Platform Engineering
GitHub
Analytics
ETL/ELT
Apply
$155k – $265k per year (Estimated) • Remote • Full-Time
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Copilot
Cursor
Claude
Spark
Claude Code
AI Agents
Feature Store
DevOps
Azure
GitHub
Analytics
Dimensional Modeling
Apply
See all jobs
This is one of many
430,068 more open roles from verified company boards, updated every day.