{"id":1577276,"url":"https://alion.io/job/geisinger-senior-platform-data-engineer","title":"Senior Platform Data Engineer","company":{"id":1774033,"name":"Geisinger","domain":"geisinger.org","url":"https://alion.io/company/geisinger","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"B","score":80,"open_postings":70,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":16,"computed_at":"2026-10-02T05:45:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":115000,"max_usd":226000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":769},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Apache Kafka","optional":false},{"name":"Databricks","optional":false},{"name":"Embeddings","optional":false},{"name":"Feature Store","optional":false},{"name":"Hybrid Search","optional":false},{"name":"LLM","optional":false},{"name":"Pinecone","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Qdrant","optional":false},{"name":"RAG","optional":false},{"name":"Reranking","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Weaviate","optional":false}],"status":"live","first_seen_at":"2026-04-16T00:00:00Z","employer_posted_date":"2026-04-16","last_verified_at":"2026-10-02T16:46:58Z","board_verified":true,"closed_at":null,"days_open":170,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":169},"description":"Location:\nWork from home (Pennsylvania)Shift:\nDays (United States of America)Scheduled Weekly Hours:\n40Worker Type:\nRegularExemption Status:\nYesJob Summary:\nThe Senior Platform Data Engineer owns roadmap, priorities, platform standards, and architecture reviews; provides formal input on performance reviews. This position makes clinical data ready for AI at scale: owning the shared data products, retrieval infrastructure, and platform administration that the entire AI portfolio depends on. Owns Real-time data feeds. Reusable clinical data models and feature pipelines. RAG retrieval infrastructure (ingestion, chunking, embeddings, vector DB, retrieval pipelines). Databricks platform administration.Job Duties:\nStreams data from Epic SDE, ADT feeds, lab results, and other clinical sources into Databricks for downstream model consumption.\n\nCurates shared clinical feature tables (patient demographics, labs, vitals, diagnoses, utilization history, imaging metadata) in Databricks/Unity Catalog that multiple AI programs consume for model training, validation, and monitoring.\n\nOwns RAG Infrastructure, the shared retrieval-augmented generation platform that agentic and generative AI programs use to ground LLM outputs in organizational knowledge.\n\nDesigns and operates document ingestion pipelines: normalizing clinical documents, policies, guidelines, and unstructured data sources into formats ready for embedding and retrieval.\n\nImplements and optimizes chunking strategies tailored to healthcare content (e.g., preserving clinical note structure, section-aware chunking for guidelines and protocols).\n\nManages the embedding pipeline: selecting, tuning, and versioning embedding models (domain-specific clinical models where they outperform general-purpose).\n\nAdministers the vector database: schema design, indexing, metadata management, access controls, and performance tuning.\n\nBuilds and maintains retrieval pipelines: hybrid search (vector + keyword/BM25), reranking, and relevance filtering to maximize retrieval precision for downstream agents and LLM applications.\n\nEstablishes data quality gates for RAG: automated profiling, completeness checks, and accuracy scoring before content enters the vector store.\n\nMonitors retrieval quality metrics (Precision@K, Recall@K, MRR) and continuously optimize retrieval performance.\n\nDatabricks workspace configuration and Unity Catalog governance.\n\nCluster policies, compute management, and cost monitoring.\n\nManges user/group management and access control.\n\nAdministrator for Feature Store.\n\nWork is typically performed in an office environment. Accountable for satisfying all job specific obligations and complying with all organization policies and procedures. The specific statements in this profile are not intended to be all-inclusive. They represent typical elements considered necessary to successfully perform the job.\n*Relevant experience may be a combination of related work experience and degree obtained (Master's Degree = 2 years).\nPosition Details:\nKey Technologies:\nDatabricks (Delta Live Tables, Feature Store, PySpark, Unity Catalog)\nEpic SDE / epic-ws for real-time clinical data extraction\nVector databases (Pinecone, Weaviate, Qdrant, or Databricks Vector Search)\nEmbedding models and pipelines (clinical domain-specific and general-purpose)\nSQL, pandas\nStreaming and batch ingestion patterns\nCDIS Data Warehouse (source system for batch clinical data)\nRequired Skills & Qualifications:\n5+ years in data engineering, with strong experience building both batch and streaming data pipelines\nExpert-level Databricks skills: Delta Live Tables, PySpark, Unity Catalog, Feature Store\nHands-on experience with real-time data ingestion (Kafka, Spark Structured Streaming, or comparable frameworks)\nStrong SQL and Python (pandas, PySpark) skills for data transformation and feature engineering\nExperience administering Databricks workspaces: cluster policies, compute management, access controls, cost monitoring\nFamiliarity with clinical data models and healthcare data sources (EHR extracts, ADT feeds, lab results, claims data) strongly preferred\nExperience with Epic data extraction methods (SDE, FHIR, epic-ws) a significant plus\nUnderstanding of data governance principles: lineage, quality monitoring, access controls\nEducation:\nBachelor's Degree-Related Field of Study (Required), Master's Degree-Related Field of Study (Preferred)Experience:\nMinimum of 5 years-Relevant experience* (Required)Certification(s) and License(s):\nSkills:\nOUR PURPOSE & VALUES: Everything we do is about caring for our patients, our members, our students, our Geisinger family and our communities.\nKINDNESS: We strive to treat everyone as we would hope to be treated ourselves. \nEXCELLENCE: We treasure colleagues who humbly strive for excellence. \nLEARNING: We share our knowledge with the best and brightest to better prepare the caregivers for tomorrow. \nINNOVATION: We constantly seek new and better ways to care for our patients, our members, our community, and the nation.\n SAFETY: We provide a safe environment for our patients and members and the Geisinger family. \nWe offer healthcare benefits for full time and part time positions from day one, including vision, dental and domestic partners. Perhaps just as important, we encourage an atmosphere of collaboration, cooperation and collegiality.\nWe know that a diverse workforce with unique experiences and backgrounds makes our team stronger. Our patients, members and community come from a wide variety of backgrounds, and it takes a diverse workforce to make better health easier for all. We are proud to be an affirmative action, equal opportunity employer and all qualified applicants will receive consideration for employment regardless to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or status as a protected veteran.","description_format":"text","description_chars":5889,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Health Care","Hospitals & Health Systems","Health Insurance & Benefits"],"lifecycle":[{"event":"open","at":"2026-10-01T10:42:18Z"}],"liveness":{"score":5,"band":"cold","label":"Long shot","p_open":1,"p_active":0.163,"p_room":0.28,"age_days":169,"expected_fill_days":16,"reasons":["conf:0","stale_co","velocity","win:tail","crowd:brand"],"computed_at":"2026-10-02T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/geisinger-senior-platform-data-engineer","json_url":"https://alion.io/job/geisinger-senior-platform-data-engineer.json","meta":{"generated_at":"2026-10-03T00:59:10Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":766,"day_limit":5000,"remaining_today":4234,"minute_limit":60,"resets_at":"2026-10-04T00:00:00Z"}}}