Mercari
We're looking for a Senior Platform Engineer with deep expertise in observability systems at scale. This role is part of our Platform Engineering division, which builds and operates the foundational infrastructure powering all of Mercari's services. You will drive the technical direction of the observability platform, identify areas for improvement, and build solutions that measurably reduce incident detection and mitigation times. You'll lead by example, mentor others, and help build a strong engineering culture within the team.
The ideal candidate is passionate about giving engineers the visibility they need to ship with confidence. You're someone who sees noisy alerts and blind spots as problems worth solving, and who gets energy from making complex systems understandable. You'll take ownership of our observability stack end-to-end - from data collection and pipeline efficiency to intelligent alerting and AI-assisted incident response. We value engineers who lead by example, mentor others, and build a strong engineering culture within the team.
Responsibilities:
- Design, build, and operate Mercari's observability platform - covering metrics, logs, traces, and alerting at scale.
- Drive measurable improvements in Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM) across all services.
- Build AI-powered solutions for automated anomaly detection, alert correlation, and incident response assistance.
- Develop self-service observability tooling that enables product engineers to instrument, monitor, and alert on their services independently.
- Define and champion observability standards, best practices, and SLO frameworks across the engineering organisation.
- Collaborate with other platform teams, SRE, Security, and product engineering teams to ensure comprehensive system visibility and reliability.
- Automate operational workflows to reduce toil and improve the team's efficiency.
- Lead technical decisions, mentor team members, and actively shape the engineering culture within the observability team.
Requirements:
- 6+ years of experience building, operating, and maintaining scalable production systems.
- Strong expertise in observability and monitoring platforms (Datadog, Prometheus, Grafana, or similar) in production environments.
- Hands-on experience with Kubernetes and container orchestration in production.
- Proficiency in Go or Python for building infrastructure tooling and services.
- Experience with cloud platforms (GCP and/or AWS) and Infrastructure as Code (Terraform).
- Deep understanding of metrics, logging, and distributed tracing, including instrumentation patterns and data pipeline design.
- Experience designing and tuning alerting systems to reduce noise and improve incident detection.
- Strong understanding of SLIs, SLOs, and error budgets as reliability frameworks.
- Proven ability to develop internal tools and platforms that improve developer productivity.
- Strong documentation and communication skills; able to write design docs and drive technical discussions.
- Shared commitment to our company's mission and values.
Preferred Requirements:
- Experience leveraging AI technologies for observability use cases (anomaly detection, alert correlation, root cause analysis).
- Track record of measurably improving MTTD and MTTM in a microservices environment.
- Experience with observability for large-scale distributed systems (500+ microservices).
- Hands-on experience with OpenTelemetry for instrumentation and data collection.
- Cost optimisation of observability data at scale (sampling strategies, data tiering, pipeline efficiency).
- Demonstrated ability to lead technical direction, mentor engineers, and build engineering culture.
- Passionate about improving developer experience through better platform tooling and self-service capabilities.
- Contributions to or active participation in open-source observability communities.
- Experience with incident management processes and tooling.
