Confirmed on the employer's own hiring board on Oct 2, 2026. First seen by Alion on Oct 1, 2026. Mutt Data scores B on the Alion truth index.
Join Our Remote Data Products & Machine Learning Startup!
At Muttdata, we build innovative Data Products and Machine Learning solutions that help companies solve complex business challenges. As a fast-growing, remote-first startup, we're passionate about technology, collaboration, and continuous learning.
We are looking for an experienced, ownership-driven Senior Data Engineer to join our team . You'll lead the architecture, design, and implementation of a next-generation, in-house clinical trial software platform built directly on Databricks, bridging the gap between software development and large-scale data engineering.
This role works closely with frontend developers, software architects, clinical research teams, and Clinical QA and Validation teams. It requires solid hands-on experience with Databricks and a strong understanding of clinical data standards and regulated environments. Ownership, clear communication, and the ability to build robust, compliant, and scalable solutions are essential to succeed in this fast-paced, collaborative environment.
What We Do
- Leveraging our expertise, we build modern Machine Learning systems for demand planning and budget forecasting.
- Developing scalable data infrastructures, we enhance high-level decision-making, tailored to each client.
- Offering comprehensive Data Engineering and custom AI solutions, we optimize cloud-based systems.
- Using Generative AI, we help e-commerce platforms and retailers create higher-quality ads, faster.
- Building deep learning models, we enhance visual recognition and automation for various industries, improving product categorization, quality control, and information retrieval.
- Developing recommendation models, we personalize user experiences in e-commerce, streaming, and digital platforms, driving engagement and conversions.
Our Partnerships
- Amazon Web Services
- Astronomer
- Databricks
Our Values
- We are Data Nerds
- We are Open Team Players
- We Take Ownership
- We Have a Positive Mindset
Curious about what we’re up to? Check out our case studies and dive into our blog post to learn more about our culture and the exciting projects we’re working on!
Responsibilities
- Design, build, and optimize enterprise data pipelines, lakehouse storage layers, and data models using Databricks (PySpark, Spark SQL, Delta Lake) to power custom clinical application backends.
- Collaborate with frontend developers, software architects, and clinical research teams to build API-driven endpoints, data ingestion engines, and query layers for proprietary clinical trial software.
- Build performant, standards-compliant data structures to store EDC outputs, audit trails, device telemetry, and patient-reported outcomes, enabling rapid querying and downstream analytics.
- Partner with Clinical QA and Validation teams to ensure database structures, data pipelines, and clinical data repositories comply with GxP, 21 CFR Part 11, HIPAA, and GDPR.
- Implement real-time and batch ingestion jobs connecting legacy clinical systems, central labs, EHRs, and wearable devices into a unified Databricks Lakehouse architecture.
- Monitor, troubleshoot, and optimize Spark jobs, Delta Lake tables, and query execution times to support high-throughput, low-latency clinical platform workflows.
Required Skills
- 4+ years of hands-on experience building production data pipelines and lakehouse architectures using Databricks, Delta Lake, and Apache Spark (PySpark or Scala).
- Demonstrated experience building, extending, or maintaining custom software applications for clinical trials (e.g., custom EDC, CTMS, Clinical Data Repositories, or eCOA/ePRO platforms).
- Deep understanding of clinical data standards and regulatory environments, including CDISC (SDTM, ADaM, CDASH), 21 CFR Part 11, GxP validation, and ICH-GCP guidelines.
- Strong experience with relational schema design, dimensional modeling, and unstructured data handling within Delta Lake environments.
- Proficiency in Python, SQL, RESTful API integrations, CI/CD pipelines, Git, and automated testing frameworks.
- Experience working in cloud environments (AWS preferred, Azure or GCP).
- Advanced English to discuss technical requirements and solutions with clients in the United States
Nice to have
- Bachelor's or Master's degree in Computer Science, Data Engineering, Bioinformatics, or a related quantitative field.
- Experience with Databricks Workflows, Delta Live Tables (DLT), and Unity Catalog governance.
- Background working in a validated system environment (Computer System Validation / CSV).
Perks
- Remote-first culture - work from anywhere!
- AWS, DBT, Google Cloud, Azure & Databricks certifications fully covered
- In-Company English Lessons.
- Birthday off + an extra vacation week (Mutt Week! )
- Referral bonuses - help us grow the team & get rewarded!
- Maslow: Monthly credits to spend in our benefits marketplace.
- Annual Mutters' Trip - an unforgettable getaway with the team!
- Monthly Childcare Reimbursement - Because supporting families matters too

