We are seeking a Senior Site Reliability Engineer (SRE) to join our engineering team. This is a hands-on individual contributor role for an experienced SRE/DevOps/Platform Engineer who can take ownership of high-severity production incidents and drive continuous improvements in platform reliability and resilience. The successful candidate will work closely with global engineering teams, participate in the on-call rotation, troubleshoot production issues, improve observability, and implement automation to reduce operational toil.
The core responsibilities for the job include the following:
Incident Management and Production Support:
- Participate in the on-call rotation and lead high-severity production incidents.
- Drive incident triage, coordinate cross-functional response, and communicate status to stakeholders.
- Troubleshoot production alerts across AWS infrastructure and application layers.
- Distinguish transient issues from systemic failures and take appropriate corrective action.
Monitoring and Observability:
- Own and enhance Datadog monitoring standards.
- Design and maintain monitors, dashboards, SLOs, and SLIs.
- Improve signal-to-noise ratios and reduce unnecessary alert fatigue.
- Monitor system health and identify reliability risks proactively.
Reliability Engineering:
- Identify systemic weaknesses and implement reliability improvements.
- Lead initiatives focused on reducing operational toil and inefficiencies.
- Design and implement automation and process improvements.
- Drive reliability initiatives through implementation and adoption.
RCA, Runbooks, and Documentation:
- Conduct detailed Root Cause Analysis (RCA) for infrastructure and application failures.
- Create structured post-mortems with actionable follow-up items.
- Develop and maintain operational runbooks and SOPs.
- Document operational learnings and continuously improve troubleshooting processes.
Deployment and CI/CD:
- Monitor CI/CD pipelines during deployments.
- Identify potential reliability risks before and during releases.
- Initiate rollbacks following established procedures when system stability is at risk.
Collaboration and Mentorship:
- Mentor junior and mid-level engineers during complex investigations.
- Review runbooks, incident reports, and operational documentation.
- Promote strong troubleshooting and operational best practices across the team.
- Collaborate effectively with global engineering teams across time zones.
Requirements:
- 6-8 years of hands-on experience in SRE, DevOps, or platform engineering.
- Proven experience leading responses to high-severity production incidents.
- Strong hands-on experience with AWS infrastructure.
- Working knowledge of: Amazon ECS, IAM, VPC, ALB/NLB, RDS, S3 MSK, ElastiCache, Lambda, and CloudWatch.
- Strong experience with Terraform / Infrastructure as Code.
- Hands-on experience building and managing Datadog monitors and dashboards.
- Strong understanding of SLOs, SLIs, and error budgets.
- Strong Linux command-line skills.
- Experience writing Python automation scripts.
- Working knowledge of GitHub Actions or GitLab CI for ECS-based deployments.
- Understanding of architectural patterns such as microservices, pub/sub, and load balancing.
- Strong structured troubleshooting and problem-solving skills.
- Excellent written and verbal English communication skills.
Preferred Skills:
- Familiarity with Grafana.
- Experience with RDS or Cassandra performance metrics.
- Basic understanding of financial markets and market data, including equities, options, and market data feeds.
- Experience mentoring engineers or establishing operational best practices.
- Ability to independently learn unfamiliar systems using documentation and runbooks.
- Strong documentation and knowledge-sharing practices.
The ideal candidate demonstrates:
- Strong ownership and accountability.
- A systematic, hypothesis-driven approach to troubleshooting.
- Ability to remain effective under production pressure.
- Strong cross-functional communication.
- Proactive identification and resolution of reliability issues.
- Ability to drive issues from discovery through complete resolution.
- A continuous improvement mindset focused on automation, resilience, and operational excellence.

