1,005,333open jobs
59,871companies
166,150added this week
Browse all
Salary
≈ $115k – $226k per year (Estimated)
Location
Hybrid (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Apr 16, 2026. Geisinger scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Geisinger is a nonprofit integrated health system headquartered in Danville, Pennsylvania, operating hospitals, clinics, a medical school and the Geisinger Health Plan for about 1.2 million people in central and northeastern Pennsylvania. Founded in 1915 as the George F. Geisinger Memorial Hospital, it runs facilities such as Geisinger Community Medical Center in Scranton and Geisinger Lewistown Hospital, and since 2024 it has been part of Risant Health, a nonprofit created by Kaiser Permanente. It hires registered nurses, patient care and telemetry technicians, phlebotomists, sterile processing and environmental services staff, nurse practitioners and health plan agents.

Location:

Work from home (Pennsylvania)

Shift:

Days (United States of America)

Scheduled Weekly Hours:

40

Worker Type:

Regular

Exemption Status:

Yes

Job Summary:

The Senior Platform Data Engineer owns roadmap, priorities, platform standards, and architecture reviews; provides formal input on performance reviews. This position makes clinical data ready for AI at scale: owning the shared data products, retrieval infrastructure, and platform administration that the entire AI portfolio depends on. Owns Real-time data feeds. Reusable clinical data models and feature pipelines. RAG retrieval infrastructure (ingestion, chunking, embeddings, vector DB, retrieval pipelines). Databricks platform administration.

Job Duties:

  • Streams data from Epic SDE, ADT feeds, lab results, and other clinical sources into Databricks for downstream model consumption.

  • Curates shared clinical feature tables (patient demographics, labs, vitals, diagnoses, utilization history, imaging metadata) in Databricks/Unity Catalog that multiple AI programs consume for model training, validation, and monitoring.

  • Owns RAG Infrastructure, the shared retrieval-augmented generation platform that agentic and generative AI programs use to ground LLM outputs in organizational knowledge.

  • Designs and operates document ingestion pipelines: normalizing clinical documents, policies, guidelines, and unstructured data sources into formats ready for embedding and retrieval.

  • Implements and optimizes chunking strategies tailored to healthcare content (e.g., preserving clinical note structure, section-aware chunking for guidelines and protocols).

  • Manages the embedding pipeline: selecting, tuning, and versioning embedding models (domain-specific clinical models where they outperform general-purpose).

  • Administers the vector database: schema design, indexing, metadata management, access controls, and performance tuning.

  • Builds and maintains retrieval pipelines: hybrid search (vector + keyword/BM25), reranking, and relevance filtering to maximize retrieval precision for downstream agents and LLM applications.

  • Establishes data quality gates for RAG: automated profiling, completeness checks, and accuracy scoring before content enters the vector store.

  • Monitors retrieval quality metrics (Precision@K, Recall@K, MRR) and continuously optimize retrieval performance.

  • Databricks workspace configuration and Unity Catalog governance.

  • Cluster policies, compute management, and cost monitoring.

  • Manges user/group management and access control.

  • Administrator for Feature Store.

Work is typically performed in an office environment. Accountable for satisfying all job specific obligations and complying with all organization policies and procedures. The specific statements in this profile are not intended to be all-inclusive. They represent typical elements considered necessary to successfully perform the job.

*Relevant experience may be a combination of related work experience and degree obtained (Master's Degree = 2 years).

Position Details:

Key Technologies:

  • Databricks (Delta Live Tables, Feature Store, PySpark, Unity Catalog)
  • Epic SDE / epic-ws for real-time clinical data extraction
  • Vector databases (Pinecone, Weaviate, Qdrant, or Databricks Vector Search)
  • Embedding models and pipelines (clinical domain-specific and general-purpose)
  • SQL, pandas
  • Streaming and batch ingestion patterns
  • CDIS Data Warehouse (source system for batch clinical data)

Required Skills & Qualifications:

  • 5+ years in data engineering, with strong experience building both batch and streaming data pipelines
  • Expert-level Databricks skills: Delta Live Tables, PySpark, Unity Catalog, Feature Store
  • Hands-on experience with real-time data ingestion (Kafka, Spark Structured Streaming, or comparable frameworks)
  • Strong SQL and Python (pandas, PySpark) skills for data transformation and feature engineering
  • Experience administering Databricks workspaces: cluster policies, compute management, access controls, cost monitoring
  • Familiarity with clinical data models and healthcare data sources (EHR extracts, ADT feeds, lab results, claims data) strongly preferred
  • Experience with Epic data extraction methods (SDE, FHIR, epic-ws) a significant plus
  • Understanding of data governance principles: lineage, quality monitoring, access controls

Education:

Bachelor's Degree-Related Field of Study (Required), Master's Degree-Related Field of Study (Preferred)

Experience:

Minimum of 5 years-Relevant experience* (Required)

Certification(s) and License(s):

Skills:

OUR PURPOSE & VALUES: Everything we do is about caring for our patients, our members, our students, our Geisinger family and our communities.

  • KINDNESS: We strive to treat everyone as we would hope to be treated ourselves.
  • EXCELLENCE: We treasure colleagues who humbly strive for excellence.
  • LEARNING: We share our knowledge with the best and brightest to better prepare the caregivers for tomorrow.
  • INNOVATION: We constantly seek new and better ways to care for our patients, our members, our community, and the nation.
  • SAFETY: We provide a safe environment for our patients and members and the Geisinger family.

We offer healthcare benefits for full time and part time positions from day one, including vision, dental and domestic partners. Perhaps just as important, we encourage an atmosphere of collaboration, cooperation and collegiality.

We know that a diverse workforce with unique experiences and backgrounds makes our team stronger. Our patients, members and community come from a wide variety of backgrounds, and it takes a diverse workforce to make better health easier for all. We are proud to be an affirmative action, equal opportunity employer and all qualified applicants will receive consideration for employment regardless to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or status as a protected veteran.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,005,333 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
United States
≈ $162k – $267k per year (Estimated) • Remote (United States) • 5+ years exp
SQL
Databases
MS SQL
AI/ML
AI Agents
Agentforce
DevOps
AWS
GitHub
Analytics
Pentaho
Management
Confluence
Jira
Agile
Apply
$154k – $186k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Bloomfield
Python
SQL
Groovy
Databases
Databricks
Apache Iceberg
Delta Lake
Apache Kafka
Apache Hudi
Teradata
AI/ML
Copilot
Spark
Model Context Protocol
Prompt Engineering
AWS Bedrock
Tokenization
DevOps
Rest API
Splunk
Terraform
Ansible
CI/CD
Jenkins
Git
AWS
Amazon EC2
Amazon S3
IAM
Amazon CloudWatch
Analytics
Tableau
ETL/ELT
Looker
Management
Agile
Scrum
Apply
$83k – $124k per year • Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Chicago • San Francisco • Scottsdale • New York
Python
Java
SQL
Scala
Python
pySpark
Java
Maven
AI/ML
Hadoop
Spark
Machine Learning
DevOps
AWS
Amazon S3
Linux
Analytics
Tableau
ETL/ELT
Apply
$110k – $150k per year • Hybrid • 8+ years exp • Bachelor's Degree • Minneapolis
Databases
Databricks
AI/ML
AI Agents
DevOps
Azure
AWS
Platform Engineering
Apply
≈ $117k – $230k per year (Estimated) • In office • Top Secret • 10+ years exp • Bachelor's Degree • United States
Python
Apply
$40k – $46k per year • In office • Full-Time • United States
Apply
≈ $40k – $95k per year (Estimated) • In office • Full-Time • United States
Apply
$40k – $46k per year • In office • Full-Time • United States
DevOps
Windows
Management
Outlook
SharePoint
Apply
≈ $46k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • United States
DevOps
Windows
Management
Outlook
SharePoint
Apply
≈ $67k – $180k per year (Estimated) • In office • Part-Time • PhD • United States
Apply
See all jobs
This is one of many
1,005,333 more open roles from verified company boards, updated every day.