1,463,475open jobs
87,550companies
229,824added this week
Browse all
Salary
≈ $28k – $55k per year (Estimated)
Location
In office (Pune)
Seniority
Architect · 5+ years exp

First seen by Alion on Oct 7, 2026.

Overview
Company
Impact
Profile match
Stratzi.ai is an AI consulting firm building custom AI systems for businesses, offering Gen-AI, predictive analytics, automation and enterprise search solutions to help organizations optimize and scale operations.

We are building a data lake/lakehouse for SCADA and industrial telemetry, storing years of high-frequency time-series data from RTUs, PLCs, IEDs, and historians, along with asset registers, maintenance records, and event logs. This role will own the data storage and database platform end-to-end, from lakehouse architecture, partitioning, and schema evolution to database performance, query optimization, governance, backup, and recovery. This is a hands-on platform architecture role, not a pure design/architecture position. You should be equally comfortable designing a lakehouse using Iceberg/Delta/Hudi and tuning queries on a busy PostgreSQL/time-series database.

The core responsibilities for the job include the following:

Data Lake / Lakehouse Architecture:

  • Own lakehouse architecture, including storage layout, open table formats (Apache Iceberg, Delta Lake, or Hudi), catalog, and medallion/zone architecture.
  • Design partitioning, clustering, and compaction strategies for large-scale time-series data.
  • Solve the small-files problem created by streaming SCADA ingestion.
  • Define file sizing, sorting, and maintenance standards.
  • Design schema evolution and versioning for tag additions, renames, and vendor changes.
  • Build rollups, downsampling, and retention tiers for raw, aggregated, and archival data.
  • Own and optimize the query layer using Trino/Presto, Spark SQL, DuckDB, or equivalent.

Database Engineering:

  • Own relational and time-series databases supporting the platform.
  • Design schemas, indexes, and queries and perform query-plan-based optimization.
  • Manage connection/resource utilization, capacity planning, HA/replication, backup, and recovery.
  • Define upgrade and migration strategies.
  • Decide what data should remain in operational/time-series databases versus the lake.
  • Ensure reconciliation between the lake, operational stores, and source historians.

Governance and Platform Operations:

  • Define data cataloging, lineage, access control, and audit standards.
  • Manage storage growth, infrastructure capacity, and cloud/on-premise costs.
  • Set technical standards and review architecture/designs.
  • Mentor data engineering teams.
  • Work with OT and network teams on data movement across segmented zones, the Purdue model, DMZ, and unidirectional gateways.

Requirements:

  • We are looking for someone hands-on, technically strong, and architecture-oriented, with deep experience in data lakes/lakehouses, databases, and high-volume time-series data.
  • Candidates with exposure to SCADA, industrial/OT, railway/metro, historians, or other large-scale telemetry environments will be strongly preferred.
  • 5+ years in data engineering, database engineering, or data architecture.
  • 3+ years of hands-on production experience with data lake/lakehouse platforms.
  • Deep expertise in at least one open table format: Apache Iceberg, Delta Lake, or Apache Hudi.
  • Strong understanding of table maintenance, snapshots, compaction, and failure scenarios.
  • Strong object storage experience: S3, ADLS, GCS, MinIO, Ceph, or equivalent.
  • Hands-on experience with distributed query engines such as Trino/Presto, Spark, or equivalent, including performance tuning.
  • Expert SQL skills and strong RDBMS expertise (PostgreSQL, Oracle, SQL Server, or similar).
  • Experience with query plans, indexing, and database performance optimization.
  • Experience handling large-scale time-series data using TimescaleDB, InfluxDB, ClickHouse, industrial historians, or equivalent.
  • Experience with metadata/catalog platforms such as Hive Metastore, AWS Glue, Unity Catalog, Nessie, Polaris, or similar.
  • Strong Python and experience with a JVM language or equivalent platform-level programming depth.
  • Strong infrastructure knowledge: Linux, containers, Kubernetes, IaC, and CI/CD.
  • Proven experience defining technical standards that other engineers build against.

Good to Have:

  • Industrial/OT data experience with AVEVA/OSIsoft PI, Wonderware, GE Proficy, Siemens WinCC, Ignition, or similar historians.
  • Experience with SCADA, RTUs, PLCs, IEDs, tag namespaces, and asset hierarchies.
  • Railway/Metro domain experience in traction power, signalling, interlocking, tunnel ventilation, ECS/BMS, rolling stock, AFC, or depot systems.
  • Experience with air-gapped or on-premises deployments.
  • Knowledge of OT security, IEC 62443, and segmented network architectures.
  • Knowledge of ISA-95 or equivalent asset information modelling.
  • Experience with dbt, Great Expectations, Soda, OpenLineage, DataHub, or Amundsen.
  • Experience working in regulated engineering environments.
  • Familiarity with railway assurance standards such as EN 50126 / EN 50128 / EN 50129.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,463,475 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Pune
≈ $100k – $227k per year (Estimated) • In office • Copenhagen
Python
SQL
AI/ML
Knowledge Graph
Machine Learning
Apply
$97k – $162k per year • Remote (United States) • Full-Time • United States
Python
SQL
Apply
≈ $187k – $405k per year (Estimated) • Remote (United States) • Full-Time • United States
Python
SQL
MATLAB
SAS
AI/ML
Context Engineering
Machine Learning
Apply
$172k – $237k per year • Remote (United States) • Full-Time • United States
Python
SQL
Databases
Snowflake
Databricks
AI/ML
Arize Phoenix
Amazon SageMaker
Metaflow
Feature Store
Machine Learning
DevOps
AWS
Apply
≈ $93k – $170k per year (Estimated) • Remote (United States) • Full-Time • Reston
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
DevOps
CI/CD
Git
AWS
GitLab
Amazon S3
IAM
Amazon CloudWatch
Analytics
ETL/ELT
Management
Agile
Apply
≈ $44k – $78k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Warsaw
Java
SQL
Databases
Oracle
DevOps
Rest API
CI/CD
Management
Agile
Apply
≈ $32k – $65k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Warsaw
SQL
C#
C++
C#
.NET
Databases
Oracle
DevOps
SLI/SLO/SLA
Management
ITIL
Service Desk
QA
Postman
Apply
Data Analyst 1 day ago
$75k – $115k per year • In office • 3+ years exp • Bachelor's Degree • Pasadena
Python
SQL
SAS
Analytics
Tableau
Power BI
A/B Testing
Microsoft Excel
Apply
SDET II 1 day ago
≈ $9.5k – $25k per year (Estimated) • In office • Bengaluru
Python
Java
DevOps
Rest API
CI/CD
SOAP
Management
Agile
Scrum
QA
Appium
Rest-Assured
Apply
≈ $52k – $97k per year (Estimated) • Hybrid • Full-Time • Warsaw
Python
PowerShell
Bash
DevOps
Terraform
Ansible
GCP
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Platform Engineering
Windows
Management
ITIL
Apply
Senior Data Architect 11 days ago
≈ $28k – $55k per year (Estimated) • In office • 5+ years exp • Pune
Python
SQL
Databases
PostgreSQL
Oracle
ClickHouse
Apache Iceberg
Delta Lake
DuckDB
Presto
TimescaleDB
MinIO
InfluxDB
Apache Hudi
Trino
AI/ML
Spark
dbt
Great Expectations
Time Series Forecasting
DevOps
CI/CD
Kubernetes
Amazon S3
Amazon ECS
Linux
Apache HTTP Server
Cybersecurity
Polaris
Analytics
AWS Glue
Apply
Data Scientist 11 days ago
≈ $15k – $36k per year (Estimated) • In office • 4+ years exp • Pune
Python
SQL
AI/ML
Weights & Biases
MLFlow
Fine-tuning
Embeddings
Scikit-learn
Prompt Engineering
TensorFlow
NumPy
PyTorch
LLM
RAG
Time Series Forecasting
Structured Outputs
LLM Guardrails
Apply
Data Engineer 11 days ago
≈ $16k – $37k per year (Estimated) • In office • 3+ years exp • Pune
Python
SQL
Databases
PostgreSQL
ClickHouse
TimescaleDB
InfluxDB
Apache Kafka
AI/ML
Spark
Dagster
dbt
Prefect
Flink
Time Series Forecasting
DevOps
Terraform
GCP
Azure
CI/CD
Git
AWS
Kubernetes
Amazon ECS
Linux
Analytics
ETL/ELT
IoT
MQTT
OPC UA
Apply
≈ $16k – $37k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Pune
Python
Java
SQL
Databases
MySQL
PostgreSQL
AI/ML
LangChain
Claude
LlamaIndex
Prompt Engineering
AI Agents
Gemini
LLM
RAG
DevOps
Rest API
GCP
Azure
Git
AWS
Docker
Linux
Apply
Backend Engineer 4 days ago
≈ $9k – $27k per year (Estimated) • In office • Bachelor's Degree • Pune
Python
Java
SQL
Databases
MySQL
PostgreSQL
AI/ML
LangChain
Claude
LlamaIndex
Prompt Engineering
AI Agents
Gemini
LLM
RAG
DevOps
Rest API
GCP
Azure
Git
AWS
Docker
Linux
Apply
≈ $20k – $41k per year (Estimated) • In office • 5+ years exp • Pune
SQL
Perl
Databases
Oracle
DevOps
Datadog
AWS
Unix
Apply
≈ $20k – $41k per year (Estimated) • In office • 5+ years exp • Pune
SQL
Perl
Databases
Oracle
DevOps
Datadog
AWS
Unix
Apply
Data Engineer 1 day ago
≈ $15k – $36k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Pune
DevOps
Azure
CI/CD
GitHub
Analytics
Azure Data Factory
Apply
≈ $19k – $39k per year (Estimated) • In office • 5+ years exp • Pune
SQL
Databases
MS SQL
DevOps
Datadog
Linux
Windows
Apply
Engineer 1 hour ago
≈ $13k – $32k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Pune
Apply
See all jobs
This is one of many
1,463,475 more open roles from verified company boards, updated every day.