First seen by Alion on Oct 7, 2026.
We are building a data lake/lakehouse for SCADA and industrial telemetry, storing years of high-frequency time-series data from RTUs, PLCs, IEDs, and historians, along with asset registers, maintenance records, and event logs. This role will own the data storage and database platform end-to-end, from lakehouse architecture, partitioning, and schema evolution to database performance, query optimization, governance, backup, and recovery. This is a hands-on platform architecture role, not a pure design/architecture position. You should be equally comfortable designing a lakehouse using Iceberg/Delta/Hudi and tuning queries on a busy PostgreSQL/time-series database.
The core responsibilities for the job include the following:
Data Lake / Lakehouse Architecture:
- Own lakehouse architecture, including storage layout, open table formats (Apache Iceberg, Delta Lake, or Hudi), catalog, and medallion/zone architecture.
- Design partitioning, clustering, and compaction strategies for large-scale time-series data.
- Solve the small-files problem created by streaming SCADA ingestion.
- Define file sizing, sorting, and maintenance standards.
- Design schema evolution and versioning for tag additions, renames, and vendor changes.
- Build rollups, downsampling, and retention tiers for raw, aggregated, and archival data.
- Own and optimize the query layer using Trino/Presto, Spark SQL, DuckDB, or equivalent.
Database Engineering:
- Own relational and time-series databases supporting the platform.
- Design schemas, indexes, and queries and perform query-plan-based optimization.
- Manage connection/resource utilization, capacity planning, HA/replication, backup, and recovery.
- Define upgrade and migration strategies.
- Decide what data should remain in operational/time-series databases versus the lake.
- Ensure reconciliation between the lake, operational stores, and source historians.
Governance and Platform Operations:
- Define data cataloging, lineage, access control, and audit standards.
- Manage storage growth, infrastructure capacity, and cloud/on-premise costs.
- Set technical standards and review architecture/designs.
- Mentor data engineering teams.
- Work with OT and network teams on data movement across segmented zones, the Purdue model, DMZ, and unidirectional gateways.
Requirements:
- We are looking for someone hands-on, technically strong, and architecture-oriented, with deep experience in data lakes/lakehouses, databases, and high-volume time-series data.
- Candidates with exposure to SCADA, industrial/OT, railway/metro, historians, or other large-scale telemetry environments will be strongly preferred.
- 5+ years in data engineering, database engineering, or data architecture.
- 3+ years of hands-on production experience with data lake/lakehouse platforms.
- Deep expertise in at least one open table format: Apache Iceberg, Delta Lake, or Apache Hudi.
- Strong understanding of table maintenance, snapshots, compaction, and failure scenarios.
- Strong object storage experience: S3, ADLS, GCS, MinIO, Ceph, or equivalent.
- Hands-on experience with distributed query engines such as Trino/Presto, Spark, or equivalent, including performance tuning.
- Expert SQL skills and strong RDBMS expertise (PostgreSQL, Oracle, SQL Server, or similar).
- Experience with query plans, indexing, and database performance optimization.
- Experience handling large-scale time-series data using TimescaleDB, InfluxDB, ClickHouse, industrial historians, or equivalent.
- Experience with metadata/catalog platforms such as Hive Metastore, AWS Glue, Unity Catalog, Nessie, Polaris, or similar.
- Strong Python and experience with a JVM language or equivalent platform-level programming depth.
- Strong infrastructure knowledge: Linux, containers, Kubernetes, IaC, and CI/CD.
- Proven experience defining technical standards that other engineers build against.
Good to Have:
- Industrial/OT data experience with AVEVA/OSIsoft PI, Wonderware, GE Proficy, Siemens WinCC, Ignition, or similar historians.
- Experience with SCADA, RTUs, PLCs, IEDs, tag namespaces, and asset hierarchies.
- Railway/Metro domain experience in traction power, signalling, interlocking, tunnel ventilation, ECS/BMS, rolling stock, AFC, or depot systems.
- Experience with air-gapped or on-premises deployments.
- Knowledge of OT security, IEC 62443, and segmented network architectures.
- Knowledge of ISA-95 or equivalent asset information modelling.
- Experience with dbt, Great Expectations, Soda, OpenLineage, DataHub, or Amundsen.
- Experience working in regulated engineering environments.
- Familiarity with railway assurance standards such as EN 50126 / EN 50128 / EN 50129.

