{"id":1931054,"url":"https://alion.io/job/source-meridian-184-data-engineer","title":"184. Data Engineer","company":{"id":688542,"name":"Source Meridian","domain":"sourcemeridian.com","url":"https://alion.io/company/sourcemeridian","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"junior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Medellín, Colombia"],"countries":[],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":41000,"max_usd":120000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":434},"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Airflow","optional":false},{"name":"Amazon S3","optional":false},{"name":"AWS","optional":false},{"name":"AWS Glue","optional":false},{"name":"Databricks","optional":false},{"name":"dbt","optional":false},{"name":"ETL/ELT","optional":false},{"name":"pySpark","optional":false},{"name":"Scala","optional":false},{"name":"Snowflake","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"IAM","optional":true},{"name":"Python","optional":true}],"status":"live","first_seen_at":"2026-10-05T21:07:34Z","employer_posted_date":"2026-10-05","last_verified_at":"2026-10-10T00:29:53Z","board_verified":true,"closed_at":null,"days_open":4,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":4},"description":"We’re looking for a Data Engineer to join Source Meridian.\nAbout Source Meridian\nSource Meridian is a development software company that works to solve the industry’s most challenging problems in healthcare practices. We are laser focused on specific technologies in the healthcare and life science industries: Healthcare technology, artificial intelligence, and healthcare interoperability.\nAbout the Role\nWe're looking for a Data Engineer to help build and operate an AWS-native data platform processing healthcare claims data and tokenized identifiers. You'll design and implement Spark-based pipelines that transform, intersect, and enrich tokenized datasets stored primarily as Parquet on S3, queried via Athena and related AWS services. This environment intentionally avoids managed lakehouse platforms (e.g., no Databricks and no Snowflake)-you'll be doing \"real\" data engineering directly on AWS.\nWhat You’ll Do\nBuild and maintain Spark pipelines to process large-scale Parquet datasets on S3.\n\nImplement tokenization workflows, including transit token → real token conversion and dataset intersection/join logic.\n\nProcess and deliver healthcare claims datasets for matched individuals, ensuring accurate identity mapping and data integrity.\n\nOrchestrate data pipelines using Airflow and/or AWS-native orchestration tools when appropriate.\n\nDevelop reliable, testable, and observable ETL/ELT processes (retries, idempotency, monitoring, reprocessing).\n\nOptimize performance and cost across Spark jobs, S3 partitioning/layout, and Athena query patterns.\n\nContribute to dbt models when applicable (transformations, documentation, data quality checks).\n\nCollaborate with cross-functional stakeholders in a healthcare environment, with a strong focus on privacy and secure data handling.\n\nRequired Qualifications\n2+ years of professional experience in Data Engineering.\n\nStrong experience with Apache Spark (PySpark or Scala), including joins, intersections, partitioning, and performance tuning.\n\nStrong hands-on experience with the AWS data stack, including:\nAmazon S3 (Parquet datasets, partition strategies, data layout best practices)\n\nAmazon Athena (SQL, query optimization, managing large datasets)\n\nFamiliarity with AWS-native data lake patterns (Glue Catalog, Lake Formation concepts are a plus)\n\nExperience building and operating pipelines using Airflow (DAGs, scheduling, dependencies, backfills).\n\nExcellent SQL skills and solid data modeling fundamentals.\n\nAdvanced English level: able to lead technical discussions, write clear documentation, and work directly with US-based stakeholders.\n\nNice to Have\nExperience with dbt (core, tests, documentation, exposures).\n\nFamiliarity with healthcare data (claims data, eligibility, member-level datasets).\n\nExperience with tokenization, identity resolution, or privacy-preserving data workflows.\n\nKnowledge of AWS security concepts such as IAM, KMS, encryption, and secure data handling.\n\nExperience running Spark on AWS (e.g., EMR) or Spark-on-containers architectures.\n\nTech Stack\nAWS-native architecture\n\nAmazon S3 + Parquet (core storage layer)\n\nAmazon Athena (query engine)\n\nApache Spark (no Databricks)\n\nAirflow (orchestration)\n\ndbt (optional, as applicable)\n\nSoft Skills\nStrong and empathetic leadership.\n\nProven client-facing experience.\n\nExcellent communication skills.\n\nStrong expectation management abilities.\n\nStrategic mindset with a solution-oriented approach and strong decision-making skills.\n\nWhat We Offer\nPermanent contract\nLearning and continuous growth environment \nBenefits package focused on health and well-being \nCompetitive salary based on experience\n Apply only if you reside in Colombia or Ecuador \nAt Source Meridian, you’ll be part of a high-impact tech-health company, building products that truly make a difference.\nIf you meet the profile - or know someone who might be interested - apply now!\nWe’d love to meet you","description_format":"text","description_chars":3908,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"Advanced (C1)","optional":false}]},"benefits":[],"hiring_locations":[{"name":"Colombia","iso":null,"kind":"country"},{"name":"Ecuador","iso":"EC","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Biotechnology"],"lifecycle":[{"event":"open","at":"2026-10-05T21:57:01Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":3,"expected_fill_days":16,"reasons":["conf:1","win:early","comp:junior"],"computed_at":"2026-10-09T06:01:00Z"},"pay":null,"html_url":"https://alion.io/job/source-meridian-184-data-engineer","json_url":"https://alion.io/job/source-meridian-184-data-engineer.json","meta":{"generated_at":"2026-10-10T02:22:39Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4346,"day_limit":5000,"remaining_today":654,"minute_limit":60,"resets_at":"2026-10-11T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":688542},"rest":"https://alion.io/mcp/rest/get_company?id=688542"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fsource-meridian-184-data-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fsource-meridian-184-data-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fsource-meridian-184-data-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/source-meridian-184-data-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fsource-meridian-184-data-engineer"}]}