368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$32k – $76k per year (Estimated)
Location
In office (Mumbai)
Seniority
Architect · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Quantiphi is an award-winning AI-first digital engineering company driven by the desire to solve transformational problems at the heart of business. Quantiphi solves the toughest and complex business problems by combining deep industry experience, disciplined cloud, and data-engineering practices, and cutting-edge artificial intelligence research to achieve quantifiable business impact at unprecedented speed.

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.

If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Data Architect

Exp Range : 8 - 13 Years

Job location : Mumbai , Bangalore, Trivandrum

Role Overview

The Data Architect is the senior technical owner of the platform's design. You will define and evolve the architectural blueprint - the canonical data model, the ingestion framework, the transformation patterns, the FHIR serialization layer, the bidirectional flow with FHIR-Repository, and the governance and observability frameworks that hold them together. You will set the standards every other role implements against.

This role works in close partnership with the customer's existing Health Data Engine technical owners, who carry deep operational knowledge of healthcare data at scale. Architectural decisions are made in dialogue with them, anchored on real volume requirements and known operational pain points. You will be expected to defend choices with technical depth and adapt them when warranted.

Key Responsibilities

  • Own the end-to-end architecture across ingestion (Flink + PySpark), storage (Iceberg on GCS with BigLake Metastore), transformation (dbt over Starburst), FHIR serialization (flat FHIR Iceberg → bundles → FHIR-Repository), and consumption (Starburst, FHIR API, data products). Maintain the architecture specification document and its companion design docs.

  • Lead the design and evolution of the Common Data Model (CDM): the dimensional, fact, bridge, and reference table classes; SCD2 semantics; hash key conventions; the write-authority matrix that governs cross-source field precedence. CDM is the project's central design artifact and the architect owns it.

  • Define the bidirectional FHIR flow with origin-tag-based loop prevention. Specify the FHIR repository egress interceptor contract, the loopback Flink consumer pipeline, and the boundary semantics between FHIR-repo-owned and CDM-owned fields.

  • Define the patient identity resolution architecture using Informatica MDM, including the synchronous-call pattern at ingestion, the asynchronous ECI change event pipeline, and the DLQ taxonomy for MDM failures.

  • Set the standards every other role implements against - naming conventions, hash algorithms (SHA-256 BINARY 32, 0x1F separator, NULL handling), Iceberg table properties (CoW vs MoR, partitioning, compaction cadence), DLQ taxonomy, observability (metric naming, structured logging, correlation ID propagation), and security (PHI handling, encryption, audit logging).

  • Own the spec-driven development framework's program-level and component-level specs. Approve significant changes to platform-wide rules. Architect specs are the constitution every implementing engineer references.

  • Drive the high-volume capacity design - sustained 50K msg/sec ingest target with 150K peak, billions of rows in CDM facts, thousand-concurrent Starburst workloads, thousands per second FHIR API. Lead capacity planning, validation, and the parallel-run cutover from the existing Health Data Engine.

  • Evaluate and recommend technology choices that are still open - Confluent Cloud tier, Starburst Galaxy vs Enterprise, FHIR-Repository sizing, multi-region failover topology. Build option analyses; defend recommendations with concrete trade-off matrices.

  • Provide architectural review for engineering work. Review specs at the component and unit level when they touch architectural concerns. Mentor data engineers and data modelers on design principles.

  • Lead architectural review meetings with customer technical stakeholders - including the Health Data Engine retirement team, governance, security, and clinical informatics. Translate complex trade-offs into language non-architect audiences can engage with.

  • Establish and evolve the disaster recovery, backup, and reprocessing strategy across all platform layers. RPO under 5 minutes for streaming, RTO under 4 hours for full platform recovery.

Required Skills and Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field

  • 8+ years of data engineering experience with at least 3-5 years in a Data Architect or Lead Data Engineer role

  • Deep expertise architecting and operating large-scale data platforms on Google Cloud Platform - Cloud Storage, Dataproc (specifically with Flink and Spark workloads), Cloud Composer, Secret Manager, Workload Identity, IAM, networking. BigQuery experience is useful background but is not used as the primary warehouse on this project

  • Hands-on experience with Apache Iceberg in production - table properties, partitioning strategies, snapshot semantics, schema evolution, CoW vs MoR write modes, compaction operations, and integration with metastore catalogs (BigLake, Glue, Polaris, or Nessie)

  • Production experience with Starburst (Galaxy or Enterprise) or Trino for analytical workloads against open table formats. Understanding of query planning, predicate pushdown, partition pruning, and workload separation patterns

  • Strong experience with streaming architectures using Apache Flink and Apache Kafka - stateful stream processing, exactly-once semantics, checkpointing, backpressure handling, and integration with Iceberg sinks

  • Expert SQL and strong programming skills in Python. Familiarity with PySpark and PyFlink for batch and streaming respectively

  • Deep experience with dbt - model materialization strategies, macros, tests, sources, exposures, project structure for large model graphs (hundreds of models). Specific experience with dbt-trino adapter is a strong plus

  • Comprehensive understanding of healthcare data standards: HL7v2 (ADT, ORU, ORM message types and segment-level parsing), CCDA, and FHIR R4 with US Core 6.1 profiles. Experience with FHIR Bundle assembly, profile validation, and resource versioning semantics

  • Hands-on experience with FHIR runtime platforms - FHIR-Repository or HAPI FHIR - including the interceptor framework, MDM module, and channel/subscription mechanisms

  • Experience with master data management for patient identity resolution. Informatica MDM specifically is preferred; equivalent experience with Verato, NextGate, or QuadraMed translates

  • Experience designing for HIPAA-regulated environments - PHI handling discipline, encryption at rest and in transit, audit logging conventions, BAA-relevant vendor decisions

  • Demonstrated ability to lead complex technical initiatives, make critical architectural decisions under uncertainty, and influence diverse stakeholders. Excellent written and verbal communication.

Nice-to-Have Skills

  • GCP Professional Data Engineer or Cloud Architect certification

  • Experience with Atlan or comparable governance platforms (Collibra, Alation, Unity Catalog)

  • Experience with multi-region active-active or active-passive deployments on GCP

  • Production experience operating Confluent Cloud at significant scale

  • Familiarity with spec-driven development workflows, particularly with AI-assisted code generation

  • Experience replacing or sunsetting legacy healthcare data platforms (Optum Health Data Engine, Innovaccer, Health Catalyst, Arcadia) is unique and highly relevant

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Mumbai
In office • Full-Time • 3+ years exp • Indonesia
DevOps
AWS
Azure
GCP
IAM
Apply
Java Developer 1 day ago
$31k – $54k per year (Estimated) • Remote • 5+ years exp • Tula
C#
JavaScript
TypeScript
Java
C#
.NET
Java
Spring Boot
Databases
Apache Kafka
ElasticSearch
PostgreSQL
RabbitMQ
Frontend
Angular
Bootstrap
React.js
DevOps
Docker
Git
Jenkins
Kubernetes
Rest API
GitLab
QA
Swagger
Apply
$27k – $58k per year (Estimated) • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Moscow
C#
JavaScript
SQL
TypeScript
C#
ASP.NET Core
Databases
Apache Kafka
MS SQL
Frontend
Angular
React.js
DevOps
Docker
Kubernetes
Apply
$27k – $66k per year (Estimated) • Remote/Hybrid • Full-Time • Moscow
Java
Databases
Apache Kafka
RabbitMQ
DevOps
Rest API
Apply
$19k – $53k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Pune
Bash
JavaScript
Python
TypeScript
Frontend
Angular
React.js
DevOps
AWS
Azure
Datadog
Docker
GCP
Grafana
Kubernetes
Prometheus
Splunk
IAM
Cybersecurity
Keycloak
Apply
$165k – $307k per year (Estimated) • Remote • Full-Time • 10+ years exp • Master's Degree • United States
Java
Python
SQL
Databases
Databricks
Snowflake
AI/ML
LangChain
LangSmith
LightGBM
LlamaIndex
LLM
LoRA
NLP
NumPy
Pandas
PEFT
PyTorch
QLoRA
RAG
Reinforcement Learning
RLHF
Scikit-learn
SciPy
Semantic Search
Spark
Transformers
Hugging Face
Semantic Search
AI Agents
DevOps
AWS
Azure
CI/CD
GCP
Apply
$22k – $52k per year (Estimated) • In office • Full-Time • 2+ years exp • Bengaluru • Mumbai
Apply
Architect - Chatbot 5 days ago
$40k – $87k per year (Estimated) • In office • Full-Time • 4+ years exp • Mumbai • Bengaluru
JavaScript
Node JS
Python
TypeScript
AI/ML
NLP
Frontend
Angular
JQuery
React.js
DevOps
AWS
CI/CD
Management
Jira
Marketing
Salesforce
Apply
$138k – $259k per year (Estimated) • Remote • Full-Time • 10+ years exp • United States
SQL
Databases
PostgreSQL
Redis
Snowflake
AI/ML
AI Agents
Edge AI
DevOps
Amazon EC2
Amazon EKS
Ansible
AWS
Chef
Docker
GCP
Kubernetes
Platform Engineering
Puppet
VMWare
Amazon CloudWatch
Amazon S3
Apply
$30k – $58k per year (Estimated) • In office • Full-Time • 4+ years exp • Bengaluru
Python
SQL
Databases
Snowflake
DevOps
AWS
Azure
Azure DevOps
Analytics
ETL/ELT
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$20k – $50k per year (Estimated) • Remote • Full-Time • 5+ years exp • Mumbai
QA
BrowserStack
Apply
$16k – $36k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 12+ years exp • Mumbai
Apply
$35k – $79k per year (Estimated) • In office • Full-Time • 12+ years exp • Mumbai
Management
ServiceNow
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.