{"id":1230745,"url":"https://alion.io/job/optum-senior-data-engineer","title":"Senior Data Engineer","company":{"id":20903,"name":"Optum","domain":"optum.com","url":"https://alion.io/company/optum","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India","Chennai, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":20000,"max_usd":41000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":51},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon ECS","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Amazon EventBridge","optional":false},{"name":"Amazon S3","optional":false},{"name":"Apache Iceberg","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"AWS Step Functions","optional":false},{"name":"CI/CD","optional":false},{"name":"CloudFormation","optional":false},{"name":"Databricks","optional":false},{"name":"Delta Lake","optional":false},{"name":"Docker","optional":false},{"name":"Git","optional":false},{"name":"IAM","optional":false},{"name":"Incident Management","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false},{"name":"HIPAA","optional":true},{"name":"Kubernetes","optional":true}],"status":"live","first_seen_at":"2026-09-17T11:24:24Z","employer_posted_date":null,"last_verified_at":"2026-09-17T11:24:24Z","board_verified":false,"closed_at":null,"days_open":13,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":13},"description":"Senior Data Engineer - Databricks & AWS\n\nJob Description :\n\nWe are looking for a Senior Data Engineer with 5+ years of experience in designing, developing, and operating scalable data platforms and production-grade data pipelines.\n\nThe ideal candidate will have strong hands-on expertise in Python, SQL, PySpark, Databricks, AWS, and modern data lakehouse technologies, along with exposure to healthcare data standards and privacy requirements.\n\nThe role involves building resilient data ingestion and processing pipelines, optimizing distributed workloads, implementing data quality and observability controls, and contributing to production-grade data engineering practices.\n\nKey Responsibilities :\n\n- Design, develop, and maintain scalable data pipelines and data processing solutions using Python, SQL, Apache Spark/PySpark, and Databricks.\n\n- Build and operate production-grade pipelines using Databricks Auto Loader, Delta Lake, Workflows, Unity Catalog, Notebooks/Jobs, and SQL Warehouses.\n\n- Optimize Spark and Databricks workloads, including cluster configuration, job performance, query optimization, and distributed processing.\n\n- Develop cloud-native data solutions using AWS services, including S3, IAM, VPC, ECS/EKS or Lambda, Step Functions, EventBridge, CloudWatch, Secrets Manager, and KMS.\n\n- Work with Apache Iceberg, including tables, catalogs, snapshots, schema and partition evolution, compaction, and interoperability with different query engines.\n\n- Design resilient ingestion frameworks for structured and semi-structured data, with appropriate validation, data quality checks, error handling, monitoring, and replay mechanisms.\n\n- Implement engineering best practices around Git, pull requests, CI/CD, Docker, Infrastructure as Code, automated testing, observability, and incident response.\n\n- Collaborate with cross-functional teams to understand data requirements and translate them into scalable technical solutions.\n\n- Ensure appropriate security, access controls, data governance, and privacy measures are incorporated into data engineering workflows.\n\n- Contribute to the responsible adoption of AI-assisted coding tools, critically reviewing, testing, securing, and productionizing AI-generated code.\n\nRequired Skills & Experience :\n\n- 5+ years of overall experience in Data Engineering.\n\n- Advanced proficiency in Python and SQL.\n\n- Strong hands-on experience with Apache Spark/PySpark, distributed processing, and performance tuning.\n\n- Production experience with Databricks, including :\n\n1. Auto Loader\n\n2. Delta Lake\n\n3. Databricks Workflows\n\n4. Unity Catalog\n\n5. Notebooks and Jobs\n\n6. SQL Warehouses\n\n7. Cluster and job optimization\n\n- Strong hands-on experience with AWS, particularly :\n\n1. Amazon S3\n\n2. IAM\n\n3. VPC\n\n4. ECS/EKS or Lambda\n\n5. Step Functions\n\n6. EventBridge\n\n7. CloudWatch\n\n8. Secrets Manager\n\n9. KMS\n\n- Practical knowledge of Apache Iceberg, including catalogs, snapshots, schema/partition evolution, compaction, and query-engine interoperability.\n\n- Experience with Git, CI/CD, Docker, Terraform/CloudFormation, automated testing, observability, and production incident management.\n\n- Experience working with semi-structured data and building reliable ingestion pipelines.\n\n- Strong understanding of data quality, validation, error handling, monitoring, and replay/reprocessing mechanisms.\n\nHealthcare Data Experience :\n\n- Experience with healthcare data and standards is highly desirable, including :\n\n1. FHIR R4 and/or HL7 v2\n\n2. OMOP Common Data Model (OMOP CDM)\n\n3. Clinical terminologies and healthcare data standards\n\n4. Handling of PHI (Protected Health Information)\n\n5. HIPAA-aligned engineering and security controls\n\n6. Data privacy and de-identification concepts\n\nAI & Engineering Practices :\n\n- Demonstrated responsible use of AI coding assistants/tools in software and data engineering workflows.\n\n- Ability to critically review AI-generated code for correctness, security, performance, maintainability, and production readiness.\n\n- Strong focus on testing, documentation, code quality, and engineering best practices.\n\nEducation :\n\n- UG : Any Graduate\n\nEmployment Type :\n\n- Full Time, Permanent\n\nDepartment :\n\n- Data Science & Analytics\nSkills\nData Engineering, AWS, Databricks, Data Pipeline, SQL, Python, PySpark, Apache Spark, Data Warehousing, AWS Lambda","description_format":"text","description_chars":4330,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Primary Care & Medical Centers","Revenue Cycle & Medical Billing","Health Data & Interoperability","Health Insurance & Benefits"],"lifecycle":[{"event":"open","at":"2026-09-25T14:00:00Z"}],"liveness":{"score":70,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.775,"p_room":0.9,"age_days":12,"expected_fill_days":23,"reasons":["seen:12","velocity","win:mid"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/optum-senior-data-engineer","json_url":"https://alion.io/job/optum-senior-data-engineer.json","meta":{"generated_at":"2026-09-30T19:29:20Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"search","counted_by":"address","units_charged":0,"used_today":0,"day_limit":null,"remaining_today":null,"minute_limit":null,"resets_at":"2026-10-01T00:00:00Z"}}}