{"id":1412095,"url":"https://alion.io/job/citi-senior-data-engineer-vice-president-python-development","title":"Senior Data Engineer - Vice President - Python Development","company":{"id":7841,"name":"Citi","domain":"citi.com","url":"https://alion.io/company/citi","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":100,"open_postings":269,"ghost_share":0.007,"stale_share":0,"repost_share":0.104,"time_to_fill_p50_days":11,"computed_at":"2026-09-30T05:45:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"head","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":19500,"max_usd":49000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":61},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"CI/CD","optional":false},{"name":"Copilot","optional":false},{"name":"Dask","optional":false},{"name":"Devin","optional":false},{"name":"ETL/ELT","optional":false},{"name":"Kubernetes","optional":false},{"name":"NumPy","optional":false},{"name":"OpenShift","optional":false},{"name":"Polars","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"SQLAlchemy","optional":false},{"name":"AI Agents","optional":true},{"name":"Airflow","optional":true},{"name":"Dagster","optional":true},{"name":"Machine Learning","optional":true},{"name":"Prefect","optional":true}],"status":"live","first_seen_at":"2026-09-28T19:24:34Z","employer_posted_date":"2026-09-28","last_verified_at":"2026-09-29T08:57:17Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"Technology\nJoin a small, high-impact engineering team in Citi Markets Technology building the data foundation for greenfield Generative AI products across asset classes. We are looking for an experienced Data Engineer who combines deep Python expertise with strong engineering judgement and a practical understanding of how to build reliable, high-performance data platforms.\nThe role goes beyond constructing ETL pipelines. You will help define how billions of records from diverse Markets data sources are collected, validated, transformed, governed, and made available to production AI applications with consistently low retrieval latency.\nIf you want to work on ambitious data engineering problems, shape a platform from the ground up, and help define how data is engineered for production AI in Markets, this is the role for you.\nThe Team\nWe are a fast-moving team specialising in Generative AI within Markets Technology. We build greenfield products that span multiple asset classes and solve real business problems using modern AI and data engineering approaches.\nOur applications depend on a robust and well-designed data foundation. That means creating pipelines and serving layers that can process billions of records while preserving data quality, provenance, security, and operational reliability.\nThe team is still small enough for every engineer to have genuine influence. Data engineering is a core part of the product architecture, not a downstream support function.\nThe Role\nAs a Vice President in the team, you will be a hands-on senior engineer responsible for designing and building the data foundation for greenfield Generative AI products across Markets, including our conversational AI platform.\nYou will develop production-grade data pipelines and services, primarily using Python, to ingest, validate, transform, enrich, and serve data from a wide range of internal and external sources. You will help create a curated, high-performance data-serving layer that enables low-latency retrieval by our AI and application services.\nYou will contribute directly to code while also shaping architecture, engineering standards, data models, quality controls, and operational practices. We are looking for someone who can make pragmatic technology choices, challenge assumptions, and take long-term ownership of the platform.\nWhat You’ll Do\nDesign, build, and evolve the data architecture supporting Generative AI products across Markets\nDevelop scalable Python pipelines that ingest, validate, transform, enrich, and integrate billions of records from diverse sources\nBuild curated, queryable datasets and high-performance serving layers for low-latency application access\nDesign and optimise the application’s data storage strategies, initially centred on PostgreSQL and Parquet\nSelect suitable processing approaches for each workload, from efficient in-process and columnar processing to distributed frameworks where required\nOptimise ingestion, transformation, storage, indexing, and query performance\nEstablish robust controls for data quality, reconciliation, lineage, schema evolution, idempotency, and recovery\nBuild production observability into data pipelines, including metrics, logging, alerting, and operational diagnostics\nDesign solutions that respect data classification, entitlements, access controls, and security requirements\nContribute directly to code, architecture reviews, technical standards, and the wider engineering direction of the team\nBuild automated tests and CI/CD pipelines that enable reliable and repeatable delivery\nCollaborate closely with AI engineers, software engineers, architects, product partners, and Markets stakeholders\nUse AI-assisted engineering tools, including Devin and GitHub Copilot, to improve development quality and productivity\nWhat We’re Looking For\nExtensive hands-on experience in data engineering, software engineering, or a closely related discipline\nDeep practical expertise in Python and the ability to build maintainable, production-grade software\nA proven track record of designing and delivering large-scale data platforms or data-intensive applications\nStrong SQL skills and extensive experience with relational databases, particularly PostgreSQL\nExperience with data modelling, indexing, partitioning, query optimisation, and database performance tuning\nStrong knowledge of the Python data ecosystem, including libraries such as pandas, PyArrow, SQLAlchemy, and NumPy\nPractical experience working with columnar formats such as Apache Parquet and selecting efficient storage and serialisation strategies\nExperience processing large-scale datasets using technologies such as Apache Spark, Dask, Polars, or equivalent frameworks\nStrong understanding of ETL and ELT architecture, including incremental processing, idempotency, failure recovery, and schema evolution\nExperience implementing automated data quality controls, reconciliation, lineage, monitoring, and operational alerting\nStrong software engineering fundamentals, including design, testing, maintainability, code review, and CI/CD\nExperience deploying and operating services on container platforms such as Kubernetes or OpenShift\nUnderstanding of data governance and security practices, including encryption, masking, classification, entitlements, and fine-grained access control\nThe ability to operate in ambiguity, take ownership, and influence the technical direction of a product\nStrong communication skills and a collaborative approach to engineering\nWhat Makes This Role Different\nThis is not a role focused on maintaining legacy ETL jobs or moving data between systems without understanding how it will be used.\nYou will help build the data foundation of a new generation of AI products in Markets. The engineering challenges include integrating complex datasets, processing billions of records, delivering consistently low retrieval latency, and meeting the quality, security, and reliability standards expected of production financial systems.\nThe role offers the opportunity to influence the architecture from an early stage, work closely with AI and application engineers, and take genuine ownership of a platform that is central to the product.\nPreferred Experience\nExperience with workflow orchestration technologies such as Apache Airflow, Dagster, or Prefect\nFamiliarity with Generative AI applications, retrieval architectures, LLMs, agentic systems, or structured evaluation\nExperience building data foundations for search, retrieval, analytics, machine learning, or AI applications\nKnowledge of financial instruments, trading concepts, and data structures within FX, Equities, or other capital markets domains\nExperience working with temporal, reference, market, or transactional data\nA bachelor’s or master’s degree in Computer Science, Engineering, or another relevant quantitative discipline, or equivalent professional experience\n------------------------------------------------------\nJob Family Group:\nTechnology------------------------------------------------------\nJob Family:\nApplications Development------------------------------------------------------\nTime Type:\nFull time------------------------------------------------------\nMost Relevant Skills\nPlease see the requirements listed above.------------------------------------------------------\nOther Relevant Skills\nFor complementary skills, please see above and/or contact the recruiter.------------------------------------------------------\nCiti is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.\nIf you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.\nView Citi’s EEO Policy Statement and the Know Your Rights poster.","description_format":"text","description_chars":7983,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Financial Services","Wealth Management & Financial Advisors","Payment Processing & Gateways","Consumer Loans & Pawnshops"],"lifecycle":[{"event":"open","at":"2026-09-28T19:24:34Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":1,"expected_fill_days":11,"reasons":["conf:20","velocity","win:early","comp:brand"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/citi-senior-data-engineer-vice-president-python-development","json_url":"https://alion.io/job/citi-senior-data-engineer-vice-president-python-development.json","meta":{"generated_at":"2026-09-30T06:40:51Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4721,"day_limit":5000,"remaining_today":279,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}