Citi is looking for a visionary and highly technical Lead Data Engineer - AI & Distributed Analytics within the Metrics AI & Analytics Group to architect, scale, and optimize the enterprise data platforms that power our advanced analytics, machine learning, and Generative AI capabilities. In this high-impact role, you will lead a team of data engineers to design and deliver production-grade, high-performance data pipelines that bridge raw data and intelligent applications across the enterprise.
You will work at the forefront of data innovation, collaborating directly with Data Scientists, AI Researchers, Product Managers, Enterprise Architects, and business stakeholders to shape the technical direction of Citi's data ecosystem. The ideal candidate is a hands-on engineering leader with deep expertise in big data technologies and hands-on experience building and scaling high-volume distributed data lakes, including modern data orchestration, ingestion, processing, and distribution platforms, with a proven track record of delivering production-grade data solutions at scale.
Responsibilities
Define and drive the technical vision, architecture, and roadmap for Citi's enterprise Data Lake and Analytics platform, working closely with Citi Enterprise Architects to ensure alignment with current and future cloud-native and hybrid-cloud strategies.
Lead, mentor, and develop a team of data engineers, setting technical standards, conducting architecture and code reviews, fostering technical excellence, and promoting a culture of continuous learning and agile delivery.
Collaborate with business leaders, product owners, data scientists, and enterprise architects to translate complex business requirements into scalable technical solutions and platform capabilities.
Provide hands-on technical leadership in the design, development, deployment, and support of enterprise-scale data engineering solutions. Hands-on development is required.
Design, build, and maintain robust batch and real-time streaming data pipelines using Apache Iceberg, Starburst, Startree/Apache Pinot, Apache Kafka, Apache Flink, and related technologies. Demonstrate expertise designing large-scale data transmission, ingestion, transformation, and processing pipelines using a Java-based technology stack.
Architect and implement Feature Stores to standardize feature engineering and ensure consistent, reliable data delivery for model training and real-time ML inference.
Optimize data storage and query performance across Enterprise Data Lakes and Lakehouses, leveraging platforms such as Delta Lake, Apache Iceberg, Databricks, and other modern data warehousing technologies.
Implement end-to-end metadata management and data lineage tracking to provide full visibility into how data flows from source systems through transformations to analytics consumers and AI models.
Establish automated data quality frameworks, including validation rules, reconciliation checks, monitoring, and controls to ensure high-fidelity data and compliance with enterprise governance standards.
Ensure all data platforms and pipelines meet enterprise security requirements, encryption standards, and data privacy regulations such as GDPR and CCPA.
Evaluate and prototype emerging data technologies, frameworks, and tools, translating findings into actionable recommendations that keep Citi's data platform at the cutting edge.
Drive continuous improvement initiatives focused on platform scalability, reliability, operational excellence, and performance optimization across distributed data systems.
Required qualifications & skills
Bachelor's degree, university degree, or equivalent professional experience in Computer Science, Data Engineering, Information Systems, or a quantitative field.
6+ years of professional experience in data engineering, software engineering, or data platform development, including 3+ years in a technical lead or engineering leadership role.
Expert-level proficiency in Java (required) and SQL, with additional expertise in Python highly desirable.
Demonstrated expertise designing and building large-scale data transmission, ingestion, transformation, and processing pipelines using Java-based technologies.
Deep expertise in big data frameworks including Apache Iceberg, Startree/Apache Pinot, Hive/HDFS, and distributed computing principles.
Hands-on experience building and operating real-time streaming pipelines using Apache Kafka and Apache Flink at high volume and scale.
Demonstrated ability to design and build data infrastructure specifically supporting machine learning, advanced analytics, or AI applications in production environments.
Experience working with modern table formats and transactional storage layers such as Apache Iceberg, Delta Lake, or Apache Hudi, with strong understanding of lakehouse architectures.
Solid experience with workflow orchestration tools such as Apache Airflow or equivalent modern orchestrators to manage complex pipeline dependencies.
Strong understanding of data governance, metadata management, data lineage, data quality frameworks, and enterprise security practices.
Experience building and scaling high-volume distributed data lakes, including modern data orchestration, ingestion, processing, and distribution capabilities.
Exceptional communication skills with the ability to articulate complex technical concepts and data architectures to both technical and non-technical stakeholders.
Strong problem-solving capabilities with demonstrated success troubleshooting and optimizing complex distributed systems.
Beneficial skills & qualifications
Master's degree in Computer Science, Data Engineering, Information Systems, or a related quantitative discipline.
Additional programming proficiency in Python and modern data engineering frameworks.
Experience with cloud data platforms such as Databricks, including cluster tuning, platform administration, and lakehouse architecture optimization.
Familiarity with vector databases such as Milvus, Pinecone, Qdrant, or Chroma for GenAI and large language model use cases.
Deep understanding of dimensional modeling, Data Vault, and schema-on-read/write design patterns for enterprise analytics workloads.
Hands-on experience with containerization and infrastructure tooling including Docker, Kubernetes, and Terraform for Infrastructure as Code.
Experience working with modern enterprise data warehousing and analytics platforms.
Demonstrated experience mentoring engineers, conducting architecture reviews, and building high-performing engineering teams.
Professional Competencies
Strategic Thinking: Ability to align technical decisions with long-term business goals, enterprise data strategy, and AI initiatives.
Communication: Exceptional communication skills, with the ability to articulate complex data architectures to highly technical engineers, business stakeholders, and executive leadership.
Problem Solving: A methodical approach to debugging complex distributed systems, resolving performance bottlenecks, and driving operational excellence.
Mentorship: A passion for developing talent, conducting constructive code and architecture reviews, and elevating the team's overall technical capability.
What we offer
At Citi, you will work at genuine scale, building data infrastructure that directly shapes how one of the world's leading financial institutions leverages AI, advanced analytics, and Generative AI. This is a role where your technical decisions carry real weight, your leadership helps develop the next generation of data engineers, and your work delivers measurable impact across the enterprise.
Hybrid working model with 3 days in the office and 2 days working remotely, giving you flexibility alongside strong team collaboration.
Opportunity to lead high-impact, enterprise-scale data engineering initiatives at the intersection of big data, machine learning, advanced analytics, and Generative AI.
Access to continuous learning and professional development resources to keep your technical skills at the forefront of the industry.
A performance-driven environment where your contributions are recognized and directly tied to your career progression.
Competitive compensation and a comprehensive financial wellbeing package reflective of the seniority and strategic importance of the role.
Wellbeing support and work-life balance programmes designed to help you perform at your best inside and outside of work.
Collaboration with world-class Data Scientists, AI Researchers, Enterprise Architects, and business leaders on problems of genuine global scale.
The opportunity to influence Citi's long-term data and AI strategy through technical leadership, innovation, and platform transformation initiatives.
Build the data foundation that powers Citi's AI-driven future, lead the next generation of enterprise data platforms, and take ownership of one of the most strategically important engineering environments in global financial services.
------------------------------------------------------
Job Family Group:
Technology------------------------------------------------------
Job Family:
Applications Development------------------------------------------------------
Time Type:
Full time------------------------------------------------------
Primary Location:
Jersey City New Jersey United States------------------------------------------------------
Primary Location Full Time Salary Range:
$142,320.00 - $213,480.00In addition to salary, Citi’s offerings may also include, for eligible employees, discretionary and formulaic incentive and retention awards. Citi offers competitive employee benefits, including: medical, dental & vision coverage; 401(k); life, accident, and disability insurance; and wellness programs. Citi also offers paid time off packages, including planned time off (vacation), unplanned time off (sick leave), and paid holidays. For additional information regarding Citi employee benefits, please visit citibenefits.com. Available offerings may vary by jurisdiction, job level, and date of hire.
------------------------------------------------------
Most Relevant Skills
Please see the requirements listed above.------------------------------------------------------
Other Relevant Skills
For complementary skills, please see above and/or contact the recruiter.------------------------------------------------------
Anticipated Posting Close Date:
------------------------------------------------------
Automated Processing and AI
We use automated processing, including artificial intelligence, for our legitimate business interests (or our reasonable and appropriate business purposes) to identify and align the candidate's skills and abilities with a specific job opening. Additionally, if you so choose, or consent, we can match your skills and abilities to other suitable roles at Citi.
Importantly, all our hiring processes and decisions, including determining your suitability for a role, are conducted, checked, and decided by individuals. Our automated processing and AI do not involve relying on automatic or autonomous decision-making. Please refer to any Jurisdictional Considerations, with specific provisions for your country (where relevant) for further details.
Illinois residents - AI Notice and Right
------------------------------------------------------
Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.
View Citi’s EEO Policy Statement and the Know Your Rights poster.

