Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs

Mercari

Mercari is a leading online marketplace platform that enables individuals to easily buy and sell pre-owned and new items from their mobile devices or web browsers. The platform offers an intuitive, user-friendly experience across a wide range of product categories, including fashion, electronics, collectibles, and home goods.

We're looking for a Senior Platform Engineer with deep expertise in observability systems at scale. This role is part of our Platform Engineering division, which builds and operates the foundational infrastructure powering all of Mercari's services. You will drive the technical direction of the observability platform, identify areas for improvement, and build solutions that measurably reduce incident detection and mitigation times. You'll lead by example, mentor others, and help build a strong engineering culture within the team.

The ideal candidate is passionate about giving engineers the visibility they need to ship with confidence. You're someone who sees noisy alerts and blind spots as problems worth solving, and who gets energy from making complex systems understandable. You'll take ownership of our observability stack end-to-end - from data collection and pipeline efficiency to intelligent alerting and AI-assisted incident response. We value engineers who lead by example, mentor others, and build a strong engineering culture within the team.

Responsibilities:

  • Design, build, and operate Mercari's observability platform - covering metrics, logs, traces, and alerting at scale.
  • Drive measurable improvements in Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM) across all services.
  • Build AI-powered solutions for automated anomaly detection, alert correlation, and incident response assistance.
  • Develop self-service observability tooling that enables product engineers to instrument, monitor, and alert on their services independently.
  • Define and champion observability standards, best practices, and SLO frameworks across the engineering organisation.
  • Collaborate with other platform teams, SRE, Security, and product engineering teams to ensure comprehensive system visibility and reliability.
  • Automate operational workflows to reduce toil and improve the team's efficiency.
  • Lead technical decisions, mentor team members, and actively shape the engineering culture within the observability team.

Requirements:

  • 6+ years of experience building, operating, and maintaining scalable production systems.
  • Strong expertise in observability and monitoring platforms (Datadog, Prometheus, Grafana, or similar) in production environments.
  • Hands-on experience with Kubernetes and container orchestration in production.
  • Proficiency in Go or Python for building infrastructure tooling and services.
  • Experience with cloud platforms (GCP and/or AWS) and Infrastructure as Code (Terraform).
  • Deep understanding of metrics, logging, and distributed tracing, including instrumentation patterns and data pipeline design.
  • Experience designing and tuning alerting systems to reduce noise and improve incident detection.
  • Strong understanding of SLIs, SLOs, and error budgets as reliability frameworks.
  • Proven ability to develop internal tools and platforms that improve developer productivity.
  • Strong documentation and communication skills; able to write design docs and drive technical discussions.
  • Shared commitment to our company's mission and values.

Preferred Requirements:

  • Experience leveraging AI technologies for observability use cases (anomaly detection, alert correlation, root cause analysis).
  • Track record of measurably improving MTTD and MTTM in a microservices environment.
  • Experience with observability for large-scale distributed systems (500+ microservices).
  • Hands-on experience with OpenTelemetry for instrumentation and data collection.
  • Cost optimisation of observability data at scale (sampling strategies, data tiering, pipeline efficiency).
  • Demonstrated ability to lead technical direction, mentor engineers, and build engineering culture.
  • Passionate about improving developer experience through better platform tooling and self-service capabilities.
  • Contributions to or active participation in open-source observability communities.
  • Experience with incident management processes and tooling.

Recommended for you based on this role

Similar stack
Same company
In your city
9+ year exp • Bengaluru
Go
Java
PHP
Databases
MySQL
DevOps
AWS
GCP
SLI/SLO/SLA
Apply
Full-Time • Minato
Go
SQL
Databases
Google BigQuery
Google Cloud Spanner
AI/ML
Claude
Claude Code
Cursor
Gemini
DevOps
ArgoCD
CI/CD
Datadog
GCP
GitHub Actions
gRPC
Kubernetes
PagerDuty
Rest API
Spinnaker
Terraform
Management
Confluence
Slack
Apply
Full-Time • Minato
Go
SQL
Databases
Google BigQuery
Google Cloud Spanner
AI/ML
Claude
Claude Code
Cursor
Gemini
DevOps
ArgoCD
CI/CD
Datadog
GCP
GitHub Actions
gRPC
Kubernetes
PagerDuty
Rest API
Spinnaker
Terraform
Management
Confluence
Slack
Apply
Full-Time • Minato
Swift
Kotlin
Mobile
Jetpack Compose
SwiftUI
DevOps
CI/CD
Apply
Full-Time • Minato
Go
Databases
ElasticSearch
AI/ML
AI Agents
DevOps
Amazon EKS
Azure AKS
cert-manager
CI/CD
FinOps
GitHub Actions
Google GKE
Jenkins
Kubernetes
PagerDuty
SLI/SLO/SLA
Terraform
AWS
Azure
GCP
Apply
Full-Time • Minato
Go
SQL
Databases
Google BigQuery
Google Cloud Spanner
AI/ML
Claude
Claude Code
Cursor
Gemini
DevOps
ArgoCD
CI/CD
Datadog
GCP
GitHub Actions
gRPC
Kubernetes
PagerDuty
Rest API
Spinnaker
Terraform
Management
Confluence
Slack
Apply
Full-Time • Minato
Go
SQL
Databases
Google Cloud Spanner
DevOps
Cloudflare
GCP
gRPC
Kubernetes
Apply
Full-Time • Minato
SQL
Databases
Google BigQuery
AI/ML
AI Agents
LLM
DevOps
FinOps
Apply
Full-Time • Minato
Python
SQL
TypeScript
Databases
Google BigQuery
PostgreSQL
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
GitHub Actions
Kubernetes
Platform Engineering
Terraform
Cybersecurity
Least Privilege
PCI DSS
SLSA
Management
Google Workspace
Apply
Full-Time • Minato
AI/ML
AI Agents
DevOps
CI/CD
GCP
Incident Management
Kubernetes
Platform Engineering
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
Bengaluru