368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$37k – $53k per year
Location
In office (Gurgaon)
Seniority
Senior · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Minfy Technologies provides cloud migration, data engineering and managed services. Its practice focuses on public cloud adoption for enterprises and public sector clients. The company also builds artificial intelligence solutions on cloud platforms.

Reports to: Head of Data & Analytics / Delivery Leadership

ABOUT THE ROLE

==============

We are seeking a Principal Data Engineer to serve as the most senior technical authority in our data and analytics practice. This is a hands-on leadership role: you will own the architecture and engineering standards for AWS-native data platforms, lead and mentor multi-pod engineering teams, and act as the trusted technical counterpart to customer stakeholders.

You will set the technical direction for large-scale batch and streaming platforms, make the build-versus-buy and pattern decisions that shape delivery for years, and remain close enough to the code to review it, tune it, and unblock the team when it matters. Success in this role is measured as much by the capability you build in others as by the systems you build yourself.

KEY RESPONSIBILITIES

====================

ARCHITECTURE & TECHNICAL DESIGN

- Own end-to-end architecture for AWS-native data platforms - lakehouse, warehouse, streaming, and data product layers - from discovery through production hardening.

- Produce and maintain architecture artefacts: solution design documents, HLD/LLD, data flow and lineage diagrams, ADRs (Architecture Decision Records), and reference implementations.

- Define the target-state roadmap and lead modernisation and migration programmes (on-premises Hadoop/Teradata/Informatica → AWS; legacy ETL → Glue/EMR/Iceberg).

- Evaluate and select services, table formats, and third-party tooling with clear trade-off analysis on cost, performance, operability, and lock-in.

- Design for non-functional requirements from day one - availability, recoverability (RTO/RPO), scalability, multi-tenancy, and disaster recovery.

PLATFORM & PIPELINE ENGINEERING

- Direct the design and build of scalable batch and streaming pipelines using AWS Glue, Amazon EMR, Amazon Kinesis, Amazon MSK, and AWS Lambda.

- Model and optimise analytical data stores in Amazon Redshift and S3-based data lakes - partitioning strategy, file formats, compaction, distribution and sort keys, workload management.

- Set the standards for CDC and ingestion patterns from operational databases and SaaS sources using AWS DMS, Glue connectors, and Kinesis.

- Own orchestration patterns across Amazon MWAA, AWS Step Functions, and Amazon EventBridge, including idempotency, retry, backfill, and SLA-breach handling.

- Own data cataloguing, lineage, and fine-grained access control via AWS Glue Data Catalog and AWS Lake Formation.

- Define the data quality and observability framework - contracts, expectations, freshness and volume checks, reconciliation, and alerting.

PERFORMANCE TUNING & COST OPTIMISATION

- Lead deep performance engineering across Redshift, Athena, Spark on EMR/Glue, and streaming workloads: query plan analysis, skew and spill remediation, partition pruning, caching, concurrency scaling, and shuffle optimisation.

- Establish benchmarking and profiling practice - define baselines, instrument workloads, and drive measurable improvements in latency, throughput, and job runtime.

- Own the FinOps posture for the data estate: right-sizing, Spot and Graviton adoption, storage tiering and lifecycle policies, Redshift RA3/serverless sizing, and per-workload cost attribution and chargeback.

- Set and enforce cost and performance SLOs, and run regular optimisation reviews with engineering and finance stakeholders.

ENGINEERING STANDARDS & CODE REVIEW

- Define and enforce engineering standards: coding conventions, repository structure, branching strategy, testing pyramid, documentation, and definition of done.

- Act as final reviewer and approver on critical pull requests; run structured code review sessions and raise the review bar across the team.

- Own CI/CD for data pipelines - automated testing (unit, integration, data quality), linting, security scanning, environment promotion, and release management.

- Champion infrastructure as code (Terraform, AWS CDK, or CloudFormation) and reusable, modular, well-tested platform components over bespoke one-off builds.

- Drive technical debt visibility and remediation planning alongside feature delivery.

TEAM LEADERSHIP & MENTORING

- Lead and technically line-manage a team of data engineers across multiple squads; allocate work, set technical goals, and own delivery quality.

- Mentor and coach senior and mid-level engineers; run design clinics, brown-bag sessions, and structured upskilling and certification paths.

- Contribute to hiring - technical screening, interview panel design, and calibration of the evaluation bar.

- Provide input to performance reviews, career development conversations, and succession planning for key technical roles.

- Build a culture of ownership, documentation, and blameless post-incident learning.

STAKEHOLDER MANAGEMENT & COMMUNICATION

- Act as the senior technical point of contact for customer architects, data leaders, and business sponsors; translate business objectives into technical roadmaps and vice versa.

- Present architecture, trade-offs, risks, and cost implications to both engineering audiences and CxO-level stakeholders with equal clarity.

- Manage expectations on scope, sequencing, and delivery risk; escalate early with options rather than problems.

- Support pre-sales and solutioning - effort estimation, technical proposals, solution walkthroughs, and proof-of-concept design.

- Partner with analysts, data scientists, and product owners to shape production-grade, consumable data products.

GOVERNANCE, SECURITY & COMPLIANCE

- Embed security controls into every pipeline - encryption at rest and in transit, KMS key management, IAM least privilege, VPC and network isolation, secrets management.

- Own PII discovery, classification, masking, tokenisation, and retention patterns; ensure designs meet applicable regulatory obligations (India DPDP Act 2023, GDPR where relevant) and audit requirements under ISO 27001 and SOC 2.

- Define data governance operating model in partnership with security and compliance functions - data ownership, access request workflows, and audit logging via AWS CloudTrail and Lake Formation.

REQUIRED SKILLS AND QUALIFICATIONS

==================================

- Bachelor's degree in Computer Science, Information Technology, Data Analytics, or a related field. Master's degree is an advantage.

- 10-15 years of overall IT experience, with 8+ years designing and delivering data platforms on AWS.

- Demonstrable track record as the lead architect or principal engineer on at least two large-scale AWS data platform builds or migrations.

- Deep expertise across AWS services and architectures, including:

- Compute: EC2, Amazon EKS, Amazon ECS, AWS Lambda, AWS Fargate, AWS Batch.

- Storage & Databases: Amazon S3, Amazon RDS, Amazon Aurora, Amazon Redshift, DynamoDB, Amazon Keyspaces, Amazon ElastiCache.

- Data & Analytics: AWS Glue (ETL, Data Catalog, DataBrew), Amazon EMR, Amazon Athena, Amazon Kinesis (Data Streams, Firehose, Managed Service for Apache Flink), Amazon MSK, AWS DMS, AWS Lake Formation, Amazon MWAA, AWS Step Functions, Amazon QuickSight.

- Operations & Monitoring: Amazon CloudWatch, AWS CloudTrail, AWS X-Ray, AWS Cost Explorer, AWS Well-Architected Tool.

- Expert-level SQL, including complex analytical patterns and performance tuning on very large datasets.

- Strong programming proficiency in Python and PySpark, with the ability to set code quality standards and review others' work critically.

- Proven experience with distributed data processing internals (Spark execution model, partitioning, shuffles, memory management).

- Hands-on experience with infrastructure as code and CI/CD for data workloads.

- Demonstrated experience leading, mentoring, and growing engineering teams.

- Excellent written and verbal communication; comfortable presenting to and negotiating with senior stakeholders.

- Strong analytical, problem-solving, and critical-thinking skills, with sound judgement under ambiguity.

- Experience working in Agile delivery environments, including sprint planning, estimation, and cross-team dependency management.

- Ability to operate independently, manage multiple concurrent engagements, and deliver to tight timelines.

PREFERRED SKILLS (NICE TO HAVE)

===============================

- AWS Certified Data Engineer - Associate, AWS Certified Solutions Architect - Professional, or AWS Certified Data Analytics - Specialty.

- Hands-on experience with modern lakehouse table formats - Apache Iceberg, Delta Lake, or Apache Hudi - including migration and optimisation at scale.

- Experience with another hyperscaler (Azure or GCP), demonstrating breadth in data engineering.

- Experience with dbt, Great Expectations, Soda, Monte Carlo, or equivalent transformation and data-quality tooling.

- Exposure to data mesh, data contracts, or domain-oriented data product operating models.

- Experience integrating ML and GenAI workloads - feature stores, Amazon SageMaker, Amazon Bedrock, vector stores, and RAG data pipelines.

- Real-time and event-driven architecture experience with Apache Flink or Kafka Streams.

- Prior experience in a consulting or professional services environment with direct client ownership.

- Open-source contributions, conference speaking, or published technical writing.

WHAT SUCCESS LOOKS LIKE IN THE FIRST 12 MONTHS

==============================================

- Architecture standards, reference patterns, and reusable platform components adopted across all data engineering pods.

- Measurable improvement in pipeline reliability (SLA adherence) and reduction in cost-per-workload across the data estate.

- A code review and quality gate process operating consistently, with defect escape rate trending down.

- At least two engineers visibly progressed to the next level through structured mentoring.

- Recognised by customer stakeholders as the go-to technical authority for the data platform.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Gurgaon
Applied - AI Engineer 10 hours ago
$25k – $69k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru • Pune
Java
Python
AI/ML
AI Agents
AWS Bedrock
Fine-tuning
Google ADK
LangGraph
LLM
LoRA
RAG
Semantic Search
LangChain
PEFT
A2A
Amazon SageMaker
AWS Strands Agents
NIST AI RMF
Semantic Search
Model Context Protocol
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Git
GitOps
gRPC
Kubernetes
OpenTelemetry
Rest API
Terraform
Vector
Amazon S3
IAM
GitLab
Apply
$34k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Java
Python
Python
Asyncio
FastAPI
Databases
Amazon Aurora
AI/ML
AWS Bedrock
Google ADK
LangChain
LangGraph
LLM
RAG
A2A
LLM Guardrails
AI Agents
Model Context Protocol
DevOps
Amazon EKS
AWS
Envoy
Kubernetes
API Gateway
IAM
Cybersecurity
Zero Trust
Apply
$75k – $198k per year (Estimated) • In office • Full-Time • Dublin
Java
Kotlin
Java
Spring Boot
Databases
Amazon Aurora
Apache Kafka
PostgreSQL
DevOps
AWS
CI/CD
Dynatrace
GitLab CI
Kubernetes
Trunk-Based Development
GitLab
Apply
$99k – $219k per year (Estimated) • In office • Full-Time • Dublin
Java
Kotlin
Databases
Amazon Aurora
PostgreSQL
DevOps
AWS
CI/CD
GitLab CI
Kubernetes
GitLab
Apply
$32k – $72k per year (Estimated) • In office • Full-Time • 6+ years exp • Bengaluru
C++
Java
Python
YARA
Databases
Amazon Aurora
DevOps
Azure
GCP
Kubernetes
Cybersecurity
MITRE ATT&CK
Suricata
YARA
Zeek
Apply
$120k – $150k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Wilmington
Management
Confluence
Jira
Trello
Apply
ML Engineer 1 day ago
$180k – $220k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Wilmington
Python
SQL
AI/ML
LightGBM
LLM
PyTorch
Scikit-learn
Spark
TensorFlow
XGBoost
Amazon SageMaker
Feature Store
LLM Guardrails
DevOps
AWS
Analytics
A/B Testing
Apply
AI Engineer 1 day ago
$21k – $32k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
AI/ML
AWS Bedrock
Embeddings
Function Calling
LLM
PyTorch
Spark
TensorFlow
Context Engineering
LLM Evaluation
LLM Guardrails
LLMOps
AI Agents
RAG
DevOps
AWS
Apply
up to $26k per year • In office • Full-Time • 12+ years exp • Hyderabad
DevOps
AWS
Azure
FinOps
GCP
Platform Engineering
SLI/SLO/SLA
IAM
Apply
up to $24k per year • In office • Full-Time • 12+ years exp • Mumbai
DevOps
AWS
Azure
FinOps
GCP
Platform Engineering
SLI/SLO/SLA
IAM
Apply
$16k – $36k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Gurgaon
Apply
$26k – $69k per year (Estimated) • In office • Full-Time • 5+ years exp • Gurgaon
Java
DevOps
Git
Apply
$32k – $83k per year (Estimated) • In office • Full-Time • 5+ years exp • Gurgaon
Python
Python
pySpark
Databases
Microsoft Fabric
AI/ML
Spark
DevOps
Azure
Apply
$24k – $65k per year (Estimated) • In office • Full-Time • 3+ years exp • Gurgaon
Databases
Snowflake
Apply
$22k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Noida • Gurgaon
C++
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.