430,068open jobs
14,618companies
59,434added this week
Browse all
Salary
$155k – $265k per year (Estimated)
Location
Remote (United States)
Seniority
Architect
Employment
Full-Time
Overview
Company
Impact
Profile match
irth Solutions is a provider of cloud-based software for 811 ticket management, asset protection, mobile workforce management, and no-code app creation. irth Solutions software is used by industries such as construction, telecommunications, utilit...

Data Architect - Insights (AI/ML)

Location: Remote (US)

Department: Insights (AI/ML)

Reports to: Engineering Manager

About the Role

Irth is building a new AI-driven threat and risk management platform for pipeline asset integrity. The platform brings together three capabilities that have historically been separate at Irth:

  • A governed, cross-product data platform built on Databricks and Azure
  • An AI-powered ingestion layer that normalizes, repairs, and enriches customer data without services-heavy onboarding
  • A reusable analytical layer that runs industry-standard, Irth-developed, and customer-built risk models against the data

We are seeking a mid-level Data Architect to design the schema and data model that powers this platform. You will create a unified structure spanning asset integrity, damage prevention, and land and stakeholder management, capable of supporting probabilistic risk models that require dozens of precisely defined input attributes per pipeline segment.

This is a hands-on design role. You will establish standards for data lineage, governance, and quality while working closely with the data engineering team to implement them. This role is ideal for someone with substantial experience in lakehouse and medallion architecture who is ready to own the data model for a platform rather than a single project.

Key Responsibilities

1. Data Modeling & Architecture

  • Design the schema and unified data model for the platform, spanning asset, inspection, incident, geospatial, and consequence data across multiple products.
  • Apply medallion architecture patterns (Bronze, Silver, Gold) and define the semantic layer consumed by downstream models, reporting, and applications.
  • Model the data required by probabilistic risk models, ensuring every required input attribute has a defined source, transformation path, and data-quality expectation.
  • Design cross-product entity resolution so records from separate products reconcile to shared assets and entities.
  • Define ingestion patterns for batch, streaming, and change-data-capture (CDC) sources, including schema evolution and slowly changing dimensions.
  • Produce architecture documentation, diagrams, data dictionaries, and decision records that engineers can implement without ambiguity.

2. Geospatial & Domain Data Design

  • Design the shared GIS layer as a first-class component of the data model, including pipeline centerlines, consequence-area polygons, right-of-way corridors, and parcel data.
  • Model linear referencing and dynamic segmentation so risk results align with centerline geometry and can be aggregated across differing segmentations.
  • Define data structures for inspection data, repair history, and external enrichment sources such as weather history, soil characteristics, and satellite-derived data.

3. Governance, Lineage & Data Quality

  • Define cataloging, classification, and metadata standards in Unity Catalog and establish expectations for lineage coverage across the data estate.
  • Establish data-quality rules, validation gates, and profiling standards, including processes for quarantining and remediating failures.
  • Define access-control patterns, including role-based and attribute-based access, tenant isolation, and appropriate handling of sensitive data.
  • Design for end-to-end traceability so that any risk output can be traced through every transformation back to its source record.

4. Implementation Partnership

  • Work directly with data engineers to translate architectural designs into production pipelines and review implementations for conformance.
  • Contribute hands-on to schema deployment, transformation logic, and optimization when it accelerates delivery.
  • Partner with application engineers and data scientists to determine how the data model is consumed through serving layers, feature stores, and APIs.
  • Support performance and cost optimization, including partitioning, clustering, file layout, and compute sizing.

5. Collaboration & Documentation

  • Participate in design reviews and architecture alignment sessions and defend design decisions using data and evidence.
  • Work with subject-matter experts to translate domain and regulatory requirements into concrete data structures and standards.
  • Maintain architecture documentation as the model evolves, treating stale or inaccurate documentation as a defect.

Requirements

Required Qualifications

  • 5-8 years of experience in data architecture, data modeling, or senior data engineering, including end-to-end ownership of a non-trivial data model.
  • Hands-on experience with Databricks, including Delta Lake, Unity Catalog, SQL Warehouses, and pipeline orchestration.
  • Strong command of medallion and lakehouse architecture patterns, dimensional modeling, and semantic layer design.
  • Advanced SQL skills and proficiency in Python or PySpark.
  • Experience with at least one major cloud platform; Azure experience preferred.
  • Practical experience with change data capture (CDC), slowly changing dimensions (SCD), schema evolution, and data-quality frameworks.
  • Working knowledge of data governance, including cataloging, lineage, classification, role-based access control (RBAC), and encryption.
  • Strong documentation and diagramming skills, with the ability to explain data models clearly to both technical and non-technical stakeholders.

Preferred Qualifications

  • Experience designing geospatial data models and working with GIS tooling, spatial joins, and linear referencing.
  • Experience with multi-tenant platforms, data residency requirements, or regulated environments.
  • Familiarity with metadata and lineage tools such as Unity Catalog, Microsoft Purview, or equivalent platforms.
  • Experience modeling data specifically for machine learning or probabilistic risk models.
  • Experience designing semantic layers for business intelligence (BI) consumption.
  • Databricks or cloud data certifications.
  • Experience using AI-assisted coding tools such as Cursor or GitHub Copilot and/or agentic coding tools such as Claude Code as part of a professional development workflow.

Nice to Have

  • Familiarity with pipeline or utility asset data, including inline inspection results, alignment sheets, facilities, and centerline geometry.
  • Understanding of asset integrity concepts such as corrosion growth, defect tracking, and consequence-of-failure modeling.
  • Awareness of regulatory reporting requirements applicable to pipeline integrity data.
  • Experience migrating data from legacy desktop or spreadsheet-based systems.

Success Metrics

Success in this role will be measured by:

  • A documented and implemented data model that supports every required risk-model input attribute with a clearly defined source and transformation path.
  • High lineage coverage and demonstrable traceability from risk outputs back to the originating source records.
  • Data engineers consistently able to build from the architecture and documentation without repeated clarification.
  • Reliable cross-product entity resolution across shared assets and entities.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
430,068 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Enterprise Architect 3 hours ago
$38k – $81k per year (Estimated) • In office • Full-Time • 20+ years exp • Bachelor's Degree • Mumbai
Python
JavaScript
TypeScript
C#
C#
ASP.NET Core
Databases
Apache Kafka
AI/ML
Hadoop
Spark
Prompt Engineering
AI Agents
Frontend
GraphQL
Angular
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Bicep
IAM
Apply
$51k – $101k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Toronto
Python
SQL
Analytics
Tableau
Management
Google Sheets
Apply
$87k – $176k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Bogotá
Python
SQL
SAS
Apply
Web Developer 3 hours ago
$25k – $59k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
Python
JavaScript
Frontend
Web Components
Apply
$15k – $35k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
SAS
Analytics
Tableau
Microsoft Excel
Apply
Quality Analyst 1 day ago
$75k – $179k per year (Estimated) • Remote • Full-Time
Python
JavaScript
SQL
Databases
Databricks
AI/ML
Copilot
Cursor
Claude
Spark
Claude Code
AI Agents
DevOps
Rest API
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
GitHub
Analytics
ETL/ELT
QA
Selenium
JMeter
Cypress
Playwright
Postman
Rest-Assured
k6
Locust
Apply
$22k – $58k per year (Estimated) • Remote/Hybrid • Full-Time
Python
SQL
Databases
PostgreSQL
Databricks
Delta Lake
PostGIS
AI/ML
Spark
Quantization
Prompt Engineering
Knowledge Distillation
NER
Great Expectations
LLM
RAG
Hallucination
Anomaly Detection
LLMOps
LLM Evaluation
LLM Guardrails
Model Distillation
DevOps
GitHub Actions
Azure
CI/CD
AWS
Vector
FinOps
Incident Management
GitHub
Analytics
Power BI
Management
Jira
Apply
$21k – $57k per year (Estimated) • Remote/Hybrid • Full-Time
Python
SQL
Python
pySpark
Databases
PostgreSQL
Databricks
Delta Lake
PostGIS
DynamoDB
Microsoft Fabric
AI/ML
Spark
MLFlow
Prompt Engineering
NLP
LLM
RAG
Anomaly Detection
Time Series Forecasting
LLM Guardrails
DevOps
GitHub Actions
Azure
CI/CD
AWS
Vector
FinOps
SLI/SLO/SLA
GitHub
Amazon S3
Cybersecurity
ISO 27001
SOC 2
GDPR
Microsoft Entra ID
Analytics
Power BI
A/B Testing
Management
Jira
Apply
$18k – $44k per year (Estimated) • Remote/Hybrid • Full-Time
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
Airflow
DevOps
GCP
Azure
CI/CD
Git
AWS
Platform Engineering
Amazon S3
Analytics
Power BI
ETL/ELT
Azure Data Factory
AWS Glue
Apply
$88k – $168k per year (Estimated) • Remote • Full-Time
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Copilot
Cursor
Claude
Spark
Claude Code
AI Agents
Human-in-the-Loop
DevOps
Azure
CI/CD
Git
Platform Engineering
GitHub
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
430,068 more open roles from verified company boards, updated every day.