368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$29k – $73k per year (Estimated)
Location
In office (Bengaluru, Tlaquepaque)
Seniority
Senior · 12+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
HP Inc. is a American multinational technology company specializing in personal computing, printing systems, and 3D printing solutions. Formed in 2015 following the corporate split of the original Hewlett-Packard Company, it manufactures a broad portfolio of consumer and enterprise products, including laptops, desktop PCs, workstation displays, and commercial printers.
Senior Machine Learning Engineer

Description -

We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale.

In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services.

You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers.

The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability.

Key Responsibilities

MLOps Platform and Architecture

  • Design and implement a scalable MLOps platform using AWS and Databricks.

  • Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models.

  • Build standardized workflows that move models from development and validation into staging and production.

  • Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure.

  • Establish clear separation between development, testing, staging, and production environments.

  • Design multi-region or multi-availability-zonearchitectures where required by business continuity and availability objectives.

CI/CD and Model Deployment

  • Build automated CI/CD pipelines for model code, inference services, infrastructure, configuration, and model artifacts.

  • Implement automated testing across the deployment lifecycle, including:

    • Unit testing

    • Integration testing

    • Model validation

    • Data contract validation

    • API and endpoint testing

    • Security testing

    • Performance and load testing

    • Regression testing

  • Automate model packaging, containerization, versioning, approval, promotion, and deployment.

  • Support deployment strategies such as blue-green deployments, canary releases, shadow deployments, and controlled traffic shifting.

  • Implement reliable rollback and roll-forward mechanisms for application code, infrastructure, model versions, prompts, and configuration.

  • Ensure deployments are reproducible, auditable, and recoverable.

Model and LLM Serving

  • Design and operate secure, scalable, low-latency inference endpoints.

  • Deploy models using appropriate services and patterns across AWS and Databricks, such as:

    • Databricks Model Serving

    • MLflow Model Registry

    • Amazon SageMaker

    • Amazon ECS or EKS

    • AWS Lambda, where appropriate

    • API Gateway

    • Application Load Balancers

  • Build synchronous, asynchronous, batch, and streaming inference capabilities.

  • Design serving architectures for LLM-powered applications, including:

    • Hosted foundation models

    • Open-source models

    • Fine-tuned models

    • Retrieval-augmented generation

    • Embedding services

    • Vector search

    • Prompt and response orchestration

    • Tool-calling and agentic workflows

  • Optimize inference performance, scalability, GPU utilization, concurrency, throughput, latency, and cost.

  • Implement autoscaling, request throttling, queuing, caching, timeout handling, and graceful degradation.

Reliability, Recovery, and Business Continuity

  • Build recoverable model-serving endpoints with clearly defined recovery time and recovery point objectives.

  • Implement automated health checks, failover mechanisms, retry policies, circuit breakers, and service recovery procedures.

  • Design backup and recovery processes for:

    • Model artifacts

    • Model registry metadata

    • Feature definitions

    • Deployment configurations

    • Infrastructure state

    • Prompts and application configuration

    • Vector indexes and knowledge-base assets

  • Create disaster recovery procedures and regularly test restoration and failover capabilities.

  • Ensure production services can recover from failed deployments, infrastructure outages, model errors, and upstream dependency failures.

  • Develop operational runbooks and incident response procedures.

Monitoring and Observability

  • Implement end-to-end observability for infrastructure, applications, models, data, and user interactions.

  • Monitor:

    • Availability

    • Request volume

    • Latency

    • Error rates

    • Resource utilization

    • Model performance

    • Data quality

    • Data drift

    • Concept drift

    • Prediction distributions

    • LLM response quality

    • Hallucination and safety indicators

    • Token consumption

    • Cost per request

  • Establish dashboards, alerts, service-level indicators, and service-level objectives.

  • Integrate monitoring with incident management and on-call processes.

  • Enable traceability from user requests through model inference, retrieval, orchestration, and downstream services.

  • Support root-cause analysis by maintaining structured logs, metrics, traces, model lineage, and deployment history.

Security and Governance

  • Implement security controls for model-serving environments, APIs, data access, and deployment pipelines.

  • Apply least-privilege access using AWS IAM, Databricks permissions, service principals, and role-based access control.

  • Secure secrets, credentials, API keys, certificates, and tokens using approved secrets-management solutions.

  • Implement encryption in transit and at rest.

  • Design private networking, endpoint controls, firewall rules, and secure connectivity patterns.

  • Support authentication, authorization, rate limiting, and tenant isolation for AI services.

  • Ensure models and LLM applications comply with organizational requirements for privacy, security, auditability, and responsible AI.

  • Maintain model lineage, approval records, version history, and deployment audit trails.

  • Implement controls for sensitive data, personally identifiable information, prompt injection, unsafe outputs, and unauthorized model access.

Infrastructure as Code and Automation

  • Build and maintain cloud infrastructure using Infrastructure as Code tools such as Terraform or AWS CloudFormation.

  • Automate environment provisioning, policy enforcement, deployment configuration, and platform upgrades.

  • Create reusable modules, templates, libraries, and deployment frameworks.

  • Implement configuration management and environment-specific parameterization.

  • Ensure infrastructure changes are peer-reviewed, tested, version-controlled, and traceable.

Collaboration and Engineering Standards

  • Partner with data scientists and ML engineers to productionize models and define deployment requirements.

  • Work with software engineering teams to integrate model endpoints into user-facing products and internal applications.

  • Collaborate with cybersecurity, architecture, legal, privacy, and governance teams.

  • Define MLOps engineering standards, design principles, coding practices, and operational requirements.

  • Conduct architecture reviews, code reviews, and production-readiness assessments.

Machine Learning Experience :

  • Significant experience in MLOps, platform engineering, DevOps, site reliability engineering, cloud engineering, or production machine learning.

  • Proven experience deploying and operating machine learning models in production.

  • Strong hands-on experience with AWS services and cloud architecture.

  • Strong hands-on experience with Databricks, including MLflow, model registries, jobs, clusters, permissions, and model-serving capabilities.

  • Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or AWS CodePipeline, Docker and containerized application deployment.

  • Experience with Kubernetes and managed container platforms such as Amazon EKS or ECS.

  • Proficiency in Python and experience building production-quality APIs and inference services.

  • Experience with REST APIs, asynchronous processing, event-driven architectures, and distributed systems.

  • Strong knowledge of Infrastructure as Code, preferably Terraform.

  • Experience with model versioning, artifact management, experiment tracking, and deployment promotion workflows.

  • Experience implementing monitoring, logging, tracing, alerting, and production support processes.

  • Strong understanding of high availability, disaster recovery, fault tolerance, and rollback strategies.

  • Knowledge of cloud networking, IAM, secrets management, encryption, and secure software delivery.

  • Training and inference workflows

  • Online and batch inference

  • Model serialization and packaging

  • Feature engineering and feature consistency

  • Model validation and evaluation,l drift and data drift

  • Model explainability and reproducibility

  • GPU-based model serving

  • Embeddings and vector databases

  • Retrieval-augmented generation

  • Prompt management and versioning

  • LLM evaluation and guardrails

  • Token limits, context management, and inference cost optimization

  • Responsible AI, content safety, and human-in-the-loop controls

This role does not necessarily require the candidate to develop new machine learning algorithms. However, the candidate must be able to understand model behavior, deployment constraints, performance characteristics, and operational risks.

Preferred Qualifications:

  • Bachelor’s or Master’s degree in Computer Science, Software Engineering, Data Engineering, Machine Learning, or a related discipline, or equivalent professional experience.

  • 12+ years of total experience

  • Experience deploying generative AI or LLM-based applications in production.

  • Experience with Amazon Bedrock, SageMaker, Databricks Mosaic AI, or similar AI platforms.

  • Experience with vector-search technologies such as Databricks Vector Search, OpenSearch, Pinecone, Weaviate, Milvus, or pgvector.

  • Experience with LLM application frameworks such as LangChain, LlamaIndex, Semantic Kernel, or equivalent orchestration tools.

  • Experience with observability platforms such as Datadog, Grafana, Prometheus, CloudWatch, OpenTelemetry, or Splunk.

  • Experience applying SRE practices to machine learning and AI systems.

  • Experience operating platforms in regulated, enterprise, or data-sensitive environments.

  • Relevant AWS, Databricks, Kubernetes, or cloud architecture certifications.

What Success Looks Like

Within the first 6 to 12 months, the successful candidate will:

  • Establish a standardized and automated path from model development to production.

  • Reduce the time required to deploy a new model or LLM service.

  • Enable repeatable deployments across development, staging, and production.

  • Provide reliable and secure endpoints through which users and applications can interact with AI.

  • Implement automated rollback and recovery for failed deployments.

  • Establish monitoring for service health, model performance, quality, security, and cost.

  • Improve availability, deployment frequency, change-failure rate, and recovery time.

  • Create reusable platform components that increase engineering productivity.

  • Establish production-readiness standards for machine learning and generative AI services.

Example Performance Indicators

  • Model deployment lead time

  • Deployment frequency

  • Percentage of automated deployments

  • Change-failure rate

  • Mean time to detect incidents

  • Mean time to recover

  • Endpoint availability

  • P95 and P99 inference latency

  • Model rollback success rate

  • Recovery-test success rate

  • Infrastructure provisioning time

  • Cost per inference request

  • Percentage of production models with complete lineage and monitoring

  • Number of security or compliance exceptions

  • Internal developer and data scientist satisfaction

Ideal Candidate Profile

You are a pragmatic platform engineer who understands that deploying a model is only the beginning. You think about security, reliability, monitoring, recovery, cost, governance, and the end-user experience from the start.

You are comfortable moving between cloud architecture, infrastructure automation, CI/CD pipelines, Python services, Databricks workflows, model registries, Kubernetes, and production incident management. You can translate experimental AI solutions into dependable services that users can trust.

Job -

Data & Information Technology

Schedule -

Full time

Shift -

No shift premium (India)

Travel -

Relocation -

Equal Opportunity Employer (EEO)-

HP, Inc. provides equal employment opportunity to all employees and prospective employees, without regard to race, color, religion, sex, national origin, ancestry, citizenship, sexual orientation, age, disability, or status as a protected veteran, marital status, familial status, physical or mental disability, medical condition, pregnancy, genetic predisposition or carrier status, uniformed service status, political affiliation or any other characteristic protected by applicable national, federal, state, and local law(s).

Please be assured that you will not be subject to any adverse treatment if you choose to disclose the information requested. This information is provided voluntarily. The information obtained will be kept in strict confidence.

For more information, review HP’sEEO Policy or read about your rights as an applicant under the law here: “ Know Your Rights: Workplace Discrimination is Illegal "

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
Head of Cyber Security 11 hours ago
$32k – $78k per year (Estimated) • Remote/Hybrid • Contractor • 5+ years exp • Saint Petersburg
DevOps
AWS
Azure
CI/CD
Docker
GCP
IAM
Kubernetes
Cybersecurity
ISO 27001
Least Privilege
NIST CSF
SOC 2
Management
Google Workspace
Apply
Full Stack Engineer 12 hours ago
$47k – $103k per year (Estimated) • In office • Full-Time • 12+ years exp • Hyderabad • Bengaluru
Java
SQL
TypeScript
JavaScript
Frontend
Angular
DevOps
AWS
Azure
Apply
$18k – $39k per year (Estimated) • In office • Full-Time • 3+ years exp • Kolkata
DevOps
Azure
Azure DevOps
Incident Management
Apply
Senior Data Scientist 11 hours ago
$28k – $55k per year (Estimated) • In office • Full-Time • Pune
Databases
Databricks
DevOps
AWS
Apply
$22k – $59k per year (Estimated) • In office • Full-Time • 3+ years exp • Navi Mumbai
DevOps
Azure
Apply
In office • Part-Time • 2+ years exp • Bachelor's Degree • Ness Ziona
C#
Java
AI/ML
Copilot
Cursor
DevOps
CI/CD
GitHub
Apply
In office • Full-Time • 4+ years exp • High School Diploma • Dalian
DevOps
Azure
Management
ServiceNow
Apply
$49k – $125k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Porto Alegre
Python
Databases
ElasticSearch
FAISS
AI/ML
Embeddings
Fine-tuning
LangChain
Multimodal AI
Prompt Engineering
PyTorch
RAG
Scikit-learn
TensorFlow
Transformers
Hugging Face
OpenAI
DevOps
Azure
CI/CD
Vector
Apply
$105k – $162k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Washington
Apply
$131k – $205k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Washington
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.