864,009open jobs
54,355companies
145,726added this week
Browse all
Salary
≈ $24k – $61k per year (Estimated)
Location
In office (Navi Mumbai, Mumbai, Thane, India)
Seniority
Junior · 2+ years exp

First seen by Alion on Sep 21, 2026.

Overview
Company
Impact
Profile match
MorningStar is a VC-funded Tokyo startup helping job seekers discover career opportunities they might otherwise miss. Its small, English-first team is preparing to launch its first product.

Job Description :

As a Data Scientist on the AI & L (Data Collection) team, you will own AI-powered solutions that extract structured information from PitchBook's reports, news, and other content. You will apply data analysis, machine learning, natural language processing (NLP), and generative AI to improve the quality, coverage, and timeliness of PitchBook data.

You will take end-to-end responsibility for data science initiatives, from problem definition, data exploration, and success metrics through model development, evaluation, production deployment, monitoring, and continuous improvement. Your work may include large language models (LLMs), retrieval-augmented generation (RAG), agentic workflows, and other information-extraction techniques.

You will collaborate with Product Managers, Machine Learning Engineers, Software Engineers, and domain experts to deliver scalable solutions. You will remain accountable for model quality and business impact after launching by evaluating performance, investigating regressions, and guiding improvements through data and experimentation. You will also contribute through peer reviews, reproducible work, documentation, and knowledge sharing.

You will join a multidisciplinary team of Data Scientists and Machine Learning Engineers developing AI and ML capabilities for PitchBook's data collection pipelines. Data Scientists own the analytical and modeling lifecycle and partner with engineering teams to operationalize, scale, and maintain successful solutions.

Primary Job Responsibilities :

- End-to-End Data Science Ownership : Own extraction and enrichment problems from discovery through production and continuous improvement. Define the problem, select data and methods, establish success criteria, evaluate results, and monitor outcomes.

- Problem Formulation & Data Strategy : Translate business requirements into measurable data science problems. Explore structured and unstructured data, identify quality and source-variability issues, and define training, validation, test, and labeling requirements with domain partners.

- Model Development & Experimentation : Design and optimize NLP, machine learning, and LLM solutions for document understanding and information extraction.

- Extraction Solution Development : Build extraction workflows using document parsing, preprocessing, chunking, feature engineering, embeddings, RAG, prompt engineering, fine-tuning, and agentic approaches.

- Evaluation & Error Analysis : Create representative evaluation datasets and metrics such as precision, recall, F1, field-level accuracy, coverage, confidence, and business impact. Use error analysis to guide model, prompt, data, and workflow improvements.

- Productionization & Model Ownership : Develop robust, testable model components and partner with ML Engineers to integrate solutions into production. Monitor quality, investigate regressions or drift, and prioritize improvements based on customer and business impact.

- Technical Trade-offs & Quality : Evaluate accuracy, coverage, latency, scalability, robustness, and cost. Recommend approaches using empirical evidence, write maintainable code, and document datasets, assumptions, experiments, limitations, and results.

- Collaboration & Innovation : Partner with Product, Data Collection, Engineering, Platform, and domain teams. Evaluate advances in NLP, generative AI, LLMs, and information extraction, and apply methods that deliver measurable value.

Skills and Qualifications :

- Bachelor's or Master's degree in Data Science, Computer Science, Statistics, Mathematics, Economics, Engineering, or a related quantitative field.

- 2+ years of experience in applied data science, machine learning, NLP, or information extraction.

- Demonstrated experience taking a data science or machine learning solution from problem definition and experimentation through production launch and ongoing improvement.

- Experience analyzing large, complex structured and unstructured datasets, including exploration, preprocessing, feature engineering, sampling, labeling, and dataset construction.

- Hands-on experience developing document intelligence or information-extraction solutions using techniques such as transformers, embeddings, RAG, LLMs, prompt engineering, fine-tuning, or agentic workflows.

- Strong understanding of experimental design, statistical reasoning, model evaluation, error analysis, and metrics such as precision, recall, F1, field-level accuracy, confidence, and coverage.

- Proficiency in Python and SQL, with experience using pandas, NumPy, scikit-learn, and PyTorch or TensorFlow.

- Experience with Hugging Face, LangChain, or comparable NLP and LLM frameworks; ability to write maintainable, testable model and data-processing code.

- Familiarity with cloud ML environments, version control, automated testing, model monitoring, containers, or data orchestration tools is beneficial.

- Strong communication and collaboration skills, including the ability to explain model behavior, limitations, trade-offs, and recommendations; experience with financial data, document intelligence, or large-scale data collection is a plus.

Working Conditions :

The job conditions for this position are in a standard office setting. Employees in this position use PC and phones on an ongoing basis throughout the day. Limited corporate travel may be required to remote offices or other business meetings and events.

Morningstar is an equal opportunity employer.

Skills

Data Science, Python, Machine Learning, NLP, SQL, PyTorch, Tensorflow, LangChain, Generative AI, RAG

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
864,009 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Navi Mumbai
≈ $92k – $168k per year (Estimated) • In office • London
Python
Databases
Apache Kafka
Trino
AI/ML
dbt
Mobile
Braze
DevOps
Platform Engineering
Analytics
ETL/ELT
Apply
≈ $115k – $185k per year (Estimated) • In office • London
Python
AI/ML
Machine Learning
Apply
AI Data Engineer 3 months ago
≈ $15k – $37k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • George Town
Python
SQL
Databases
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Hadoop
Spark
Reinforcement Learning
Anomaly Detection
Time Series Forecasting
DevOps
GCP
AWS
Analytics
ETL/ELT
Azure Data Factory
Apply
Data Engineer 2 months ago
$106k – $150k per year • Remote (United States) • Full-Time • 6+ years exp • Bachelor's Degree
SQL
C#
C#
.NET
Databases
MS SQL
Azure Cosmos DB
DevOps
Azure DevOps
Azure
CI/CD
Analytics
ETL/ELT
SSIS
Azure Data Factory
Management
SharePoint
Agile
Scrum
Apply
≈ $60k – $132k per year (Estimated) • In office • 3+ years exp • Barcelona
Python
SQL
Databases
Neo4j
Databricks
AI/ML
Spark
dbt
DevOps
Azure
CI/CD
Git
AWS
Analytics
ETL/ELT
Apply
≈ $12k – $30k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Manila
Python
JavaScript
SQL
Databases
Presto
AI/ML
Spark
Analytics
Microsoft Excel
Apply
In office • Full-Time • Bachelor's Degree • Singapore
Python
SQL
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Pune
JavaScript
TypeScript
SQL
Apex
Databases
SQLite
Frontend
RxJS
Angular
Bootstrap
Mobile
Ionic
Capacitor
Cordova
DevOps
Rest API
CI/CD
Git
Gitflow
Bitbucket
Windows
Apache HTTP Server
Management
Jira
QA
Playwright
Apply
Remote (United States) • Full-Time
SQL
Analytics
Power BI
Microsoft Excel
Management
Microsoft Office
Apply
≈ $85k – $176k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Calgary
Python
Apply
≈ $39k – $100k per year (Estimated) • In office • 5+ years exp • Navi Mumbai
Python
SQL
AI/ML
Embeddings
AI Agents
LLM
RAG
Hallucination
Agentic Workflows
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Apply
$64k – $94k per year • Hybrid • Full-Time • 3+ years exp • Toronto
JavaScript
TypeScript
Node JS
Frontend
Vue.js
Webpack
Next.js
Angular
React.js
Vite
Rollup
DevOps
AWS
Apply
DevOps Intern 1 month ago
In office • Internship • Tokyo
Python
SQL
DevOps
AWS
Kubernetes
Apply
≈ $26k – $54k per year (Estimated) • In office • Navi Mumbai
SQL
Databases
Oracle
Apply
≈ $29k – $68k per year (Estimated) • In office • 10+ years exp • Pune • Mumbai • Kolkata • Navi Mumbai
Apply
≈ $29k – $65k per year (Estimated) • Hybrid • 10+ years exp • Mumbai • Bengaluru • Navi Mumbai
Analytics
Informatica
Apply
≈ $24k – $54k per year (Estimated) • In office • 7+ years exp • Navi Mumbai
Management
Confluence
Jira
Agile
Scrum
Microsoft Office
Apply
≈ $30k – $81k per year (Estimated) • In office • 10+ years exp • Mumbai • Navi Mumbai
AI/ML
AI Agents
Marketing
Salesforce
HubSpot
Apply
See all jobs
This is one of many
864,009 more open roles from verified company boards, updated every day.