368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$20k – $49k per year (Estimated)
Location
Remote/Hybrid (India)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
ScienceLogic provides artificial intelligence driven monitoring for hybrid infrastructure. Its platform discovers assets, correlates events and automates remediation across clouds. Service providers and enterprises use it to run large heterogeneous estates.

About ScienceLogic…

ScienceLogic is redefining IT operations for the modern enterprise. Our AIOps platform empowers organizations to achieve Autonomic IT - where systems are self-healing, self-optimizing, and seamlessly aligned with business outcomes. We help enterprises and service providers gain unified visibility across hybrid and multi-cloud environments, automate workflows, and unlock performance at scale.

We’re accelerating digital transformation through the power of automation, AI, and analytics - giving IT and business leaders the tools to deliver superior customer experiences, drive efficiency, and innovate with confidence.

Data Scientist

About the Role

We run a suite of small, locally-hosted language models in production - not a single frontier API. That deliberate architecture defines this job: each model is more constrained than a giant hosted one, so product quality comes from how well we evaluate, route, prompt, ground, and orchestrate the models we have. Your job is to get the best possible outcomes out of that suite.

This is not classical predictive modeling. The object of measurement is the LLM system itself - its answers, retrieval, multi-step agent behavior, and reliability under adversarial and edge-case conditions. You'll define what "good" means for a non-deterministic system running on bounded local models, build the evaluation infrastructure that catches regressions, and turn interaction data into the analysis that tells engineering and product where to invest. You'll also build production prediction and trend capability - forecasting, anomaly detection, and early-warning signals over operational telemetry - that feeds directly into that system. You'll work across data scientists, ML/inference engineers, frontend, and product in an enterprise environment with real security and compliance constraints. If you think in eval suites, failure modes, and groundedness - and you're energized by squeezing reliable, high-quality behavior out of small models under real resource budgets - this is the role.

Key Responsibilities

Evaluation & Response Quality

· Design and own evaluation harnesses for LLM and agentic outputs - golden sets, regression suites, and rubric-based scoring.

· Build and calibrate LLM-as-judge pipelines; validate judges against human labels and control for their bias and variance.

· Define and track response-quality metrics: faithfulness/groundedness, hallucination rate, answer relevance and completeness, instruction-following, and persona adherence.

· Curate, version, and grow evaluation datasets as the product and its surfaces evolve.

· Benchmark the models in the suite against each other to decide which model handles which task, and quantify the quality cost of running smaller, local models versus larger alternatives.

Adversarial & Robustness Testing

· Red-team the system: prompt injection, jailbreaks, tool-misuse, and edge-case discovery.

· Design chaos and stress tests that probe model and agent reliability under degraded or hostile conditions.

· Characterize failure modes and feed them back into guardrails and regression coverage.

Retrieval & Agentic Trajectory Analysis

· Evaluate retrieval quality over the document corpus - recall@k, MRR/nDCG, context precision and recall - and run experiments on chunking, indexing, and hybrid retrieval strategies.

· Analyze multi-step agent trajectories: tool-call correctness, trajectory efficiency, replayable-state inspection, and guardrail-breach behavior.

· Assess intent classification and routing quality as measurable components, not black boxes.

Behavioral Regression & Drift

· Build standing evaluation that catches quality and behavioral regressions when a model in the suite is swapped, upgraded, or re-quantized, or when prompts and pipelines change.

· Monitor output-distribution and quality drift in production; distinguish genuine regressions from noise on stochastic outputs.

· Recommend and validate fixes through the levers available with local models - prompt changes, retrieval and grounding adjustments, routing changes, or model selection.

Predictive & Trend Modeling

· Build, ship, and own production models that forecast and surface trends from operational telemetry - capacity and resource forecasting, anomaly prediction, and early-warning signals on metrics and logs.

· Take these from prototype to production and keep them healthy: deployment, monitoring, recalibration, and retraining as data and behavior shift.

· Define accuracy and lead-time metrics that matter operationally - precision/recall on predicted incidents, forecast error, how far ahead a signal fires - not just offline scores.

· Wire predictive signals into the LLM and agentic layer so forecasts and trends feed reasoning, advisories, and operator-facing recommendations.

Domain & Value Analytics

· Apply AIOps/NOC analysis where it's the product: log anomaly detection, event correlation, and root-cause and problem analysis.

· Quantify the economics of the system - cost and token consumption per interaction, interaction-type taxonomies - and connect them to customer-facing value metrics like MTTR and operator-hours.

· Communicate findings to engineering and product stakeholders through clear, in-context analysis.

Method & Innovation

· Use LLM-assisted workflows to scale the work itself - drafting analyses, generating synthetic evaluation cases, and bootstrapping labeled data for human refinement.

· Track and adopt state-of-the-art evaluation, retrieval, and agentic-analysis techniques; bring the useful ones into the team's workflow.

Required Experience

  • Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or a related field - or equivalent experience.

  • 3+ years in data science, ML, or applied quantitative analysis.

  • Strong applied statistics, with the judgment to design sound experiments and significance tests on noisy, non-deterministic outputs (not just clean A/B conversion).

  • Experience building, deploying, and monitoring predictive or time-series models in production - forecasting, anomaly detection, or trend analysis - including recalibration as data shifts.

  • Demonstrated work evaluating, analyzing, or improving LLM or NLP systems - eval design, quality measurement, retrieval evaluation, or agent analysis.

  • Proficiency in Python.

  • Strong SQL and comfort querying large analytical datasets.

  • Fluency with foundation models and hands-on experience with the modern LLM evaluation and tooling layer - eval/harness frameworks, judge pipelines, and the libraries used to serve, prompt, and test models.

  • Ability to build analysis and visualization in code.

Preferred Qualifications

  • Experience getting strong results out of small or self-hosted/local models under compute, memory, or latency constraints - quantization-aware evaluation, prompt and context optimization, or model routing.

  • Experience with retrieval-augmented systems and retrieval evaluation at scale.

  • Experience with agentic frameworks and tool-use/orchestration analysis, including human-in-the-loop and replayable-state patterns.

  • Familiarity with red-teaming or adversarial robustness for LLMs.

  • Domain background in IT operations - AIOps, NOC, ITSM, observability, or anomaly detection on logs and telemetry.

  • Experience with large-scale analytical and big-data stores.

  • Cloud experience for data science and ML workloads.

  • Exposure to enterprise security and compliance constraints in a delivery context.

    Don’t meet every single requirement? Studies have shown that women and people of color are less likely to apply to jobs unless they meet every single qualification. At ScienceLogic, we are dedicated to building a diverse, inclusive and authentic workplace, so if you’re excited about this role but your past experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyways. You may be just the right candidate for this or other roles.

    All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other applicable legally protected characteristics in the location in which you are applying.

    About ScienceLogic

    ScienceLogic is a leader in IT Operations Management, providing modern IT operations with actionable insights to resolve and predict problems faster in a digital, ephemeral world. Its solution sees everything across cloud and distributed architectures, contextualizes data through relationship mapping, and acts on this insight through integration and automation.

    www.sciencelogic.com

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$122k – $200k per year • In office • Full-Time • 10+ years exp • Charlotte • New York
AI/ML
AI Agents
Context Engineering
Copilot
Hallucination
Human-in-the-Loop
Knowledge Graph
LLM
LLM Guardrails
LLMOps
Prompt Engineering
RAG
DevOps
Azure
CI/CD
GitHub
Kubernetes
Vector
Analytics
A/B Testing
Apply
$145k – $193k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Chicago • Washington • Denver • Jersey City • Boston
Python
AI/ML
AI Agents
Anomaly Detection
Embeddings
Fine-tuning
Function Calling
Hugging Face
LangChain
LLM
LLM Evaluation
LLM Guardrails
PyTorch
RAG
Red Teaming
Scikit-learn
Vertex AI
DevOps
Azure
GCP
Cybersecurity
Threat Modeling
Apply
Sr. UX Designer 3 hours ago
$76k – $161k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Toronto
Python
AI/ML
AI Agents
Hallucination
Human-in-the-Loop
LLM
LLM Guardrails
Design
Figma
Apply
$87k – $130k per year • In office • Full-Time • 6+ years exp • Murray
Python
TypeScript
JavaScript
Python
Alembic
FastAPI
Pydantic
SQLAlchemy
Databases
PostgreSQL
Redis
AI/ML
Embeddings
LLM
LLM Guardrails
Ollama
RAG
Frontend
React Query
React Router
React.js
Vite
DevOps
AWS
CI/CD
Docker
Docker Compose
IAM
Terraform
Cybersecurity
FedRAMP
NIST 800-53
QA
Pytest
Apply
$187k – $378k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp
Python
AI/ML
AI Agents
Edge AI
Embeddings
LLM
NumPy
Pandas
Spark
Analytics
A/B Testing
Apply
$16k – $47k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree
Bash
Python
DevOps
AIOps
Amazon EC2
AWS
CI/CD
CloudFormation
Docker
Jenkins
Kubernetes
Self-Healing
Terraform
VMWare
Apply
$27k – $65k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree
Bash
Go
Python
DevOps
AIOps
AWS
CI/CD
Jenkins
Platform Engineering
Self-Healing
Terraform
VMWare
Apply
$32k – $63k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree
JavaScript
Python
Frontend
React.js
DevOps
AIOps
CI/CD
Rest API
Self-Healing
Apply
$20k – $49k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree
JavaScript
Node JS
PHP
Python
AI/ML
Cursor
Tabnine
OpenAI Codex
DevOps
AIOps
Amazon EC2
AWS
CI/CD
Docker
Jenkins
Kubernetes
Self-Healing
VMWare
Management
Jira
QA
Cucumber
Apply
$135k – $155k per year • Remote • Full-Time • 5+ years exp
SQL
Databases
ClickHouse
AI/ML
Anomaly Detection
dbt
LLM
AI Agents
DevOps
AIOps
OpenTelemetry
Self-Healing
Analytics
Power BI
Tableau
Management
Jira
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.