Job Description
Thomson Reuters is strengthening its Site Reliability Engineering capability to help engineering and operations teams build, operate, and improve reliable production services.
The Site Reliability Engineer will support the tools, processes, and operational practices that help teams detect, investigate, respond to, and prevent production reliability issues. You will work with observability platforms, operational documentation, automation, deployment information, and AI-enabled tools to improve the quality and availability of context used during day-to-day operations and incidents.
This is a hands-on engineering role for someone who enjoys learning how complex systems work, improving operational readiness, and contributing practical solutions to production challenges. You will work closely with experienced SREs, Product Engineering teams, platform teams, and operations partners to maintain reliable services and reduce operational toil.
Rather than expecting you to know every architecture on day one, this role will help you build familiarity across products and platforms through maintained documentation, dashboards, runbooks, telemetry, deployment data, and operational tooling. When you identify gaps in that context, you will help improve the systems and processes that keep it current.
You will also use AI-enabled engineering and investigation tools responsibly to accelerate analysis, documentation, and operational workflows. You will apply technical judgment, validate outputs, and escalate when additional expertise or review is needed.
Key Responsibilities
Support and maintain SRE operational tooling, including dashboards, alerts, runbooks, service documentation, telemetry baselines, deployment visibility, and dependency information.
Use observability tools-including logs, metrics, traces, dashboards, and alerts-to investigate service-health issues, identify trends, and support incident response.
Participate in incident response by gathering relevant context, reviewing recent changes, following established runbooks, documenting findings, and helping coordinate technical follow-up actions.
Execute approved runbooks and mitigation procedures within established escalation, change-management, and decision-making processes.
Clearly document facts, observations, hypotheses, actions, and open questions during incidents, handoffs, and operational reviews.
Help improve the accuracy, completeness, and freshness of operational context used by engineering and operations teams during incidents and routine production support.
Contribute to automation and integration work that keeps operational information current, such as CI/CD notifications, deployment telemetry, change-correlation data, service ownership records, and monitoring configuration.
Review and validate operational artifacts, including runbooks, diagrams, dashboards, alerts, and AI-generated documentation, with guidance from senior engineers and service owners.
Assist with root-cause analysis, post-incident reviews, and follow-up work by identifying gaps in monitoring, documentation, automation, instrumentation, or operational processes.
Treat missing runbooks, outdated documentation, incomplete telemetry, and unclear service ownership as improvement opportunities; partner with the appropriate teams to help resolve those gaps.
Contribute to service-health and error-reduction initiatives using available SLO, error budget, incident, alerting, and operational data.
Partner with Product Engineering and platform teams to identify reliability and observability improvements, including monitoring gaps, alert quality, deployment visibility, capacity concerns, and failure-mode coverage.
Contribute directly to code, scripts, infrastructure configuration, dashboards, alerts, automation, and documentation that improve service reliability and reduce manual operational work.
Use AI-enabled coding and investigation tools to accelerate log review, documentation updates, runbook drafting, incident summarization, and hypothesis generation, while validating results before relying on them.
Provide actionable feedback when AI-enabled operational tools produce incomplete, inaccurate, or insufficiently supported outputs.
Participate in design reviews, sprint planning, and operational-readiness discussions, helping ensure reliability and observability considerations are addressed before production deployment.
Support blameless post-incident reviews focused on learning, systemic improvement, and preventing recurring issues.
Required Qualifications
3+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, platform engineering, systems engineering, production operations, software engineering, or a related technical field.
Working knowledge of at least two of the following areas: cloud infrastructure, distributed systems, observability and monitoring, networking, databases, CI/CD, containers, or infrastructure automation.
Experience working with production telemetry, including logs, metrics, traces, dashboards, monitoring platforms, or alerting systems.
Experience troubleshooting production issues, participating in incident response, or supporting business-critical applications and services.
Experience writing, maintaining, or improving operational documentation, runbooks, knowledge articles, or support procedures.
Experience with scripting or programming in one or more languages, such as Python, Bash, PowerShell, JavaScript, Java, Go, or a comparable language.
Willingness and ability to contribute to code, infrastructure configuration, dashboards, monitoring rules, alerts, automation, or documentation.
Ability to communicate clearly about technical findings, operational risks, and next steps with engineers and operational stakeholders.
Ability to work effectively in a collaborative environment, ask for help when needed, and learn from more experienced engineers.
Familiarity with AI-enabled coding, documentation, investigation, or operational-analysis tools, along with an understanding that outputs must be reviewed and validated.
Preferred Qualifications
Experience with cloud platforms, Kubernetes, containers, CI/CD tooling, infrastructure-as-code, or configuration-management tools.
Experience with observability platforms such as Datadog, Dynatrace, New Relic, Splunk, Grafana, Prometheus, Elastic, or similar technologies.
Experience defining or working with Service Level Objectives, Service Level Indicators, error budgets, service-health metrics, or incident-management processes.
Experience supporting 24/7 production environments or participating in an on-call rotation.
Experience improving dashboards, alerts, runbooks, deployment visibility, change correlation, service documentation, or operational workflows.
Familiarity with AI agent workflows, AI-assisted root-cause analysis, or AI-enabled incident-management tools.
Experience contributing to automation that reduces repetitive operational work and improves response consistency.
Experience participating in blameless post-incident reviews and helping drive corrective actions to completion.
Relevant certifications in cloud infrastructure, Kubernetes, DevOps, SRE, observability, or incident management are beneficial but not required.
What’s in it For You?
- Hybrid Work Model: We’ve adopted a flexible hybrid working environment for our office-based roles while delivering a seamless experience that is digitally and physically connected.
- Flexibility & Work-Life Balance: Flex My Way is a set of supportive workplace policies designed to help manage personal and professional responsibilities, whether caring for family, giving back to the community, or finding time to refresh and reset. This builds upon our flexible work arrangements, including work from anywhere for up to 8 weeks per year, empowering employees to achieve a better work-life balance.
- Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrow’s challenges and deliver real-world solutions. Our Grow My Way programming and skills-first approach ensures you have the tools and knowledge to grow, lead, and thrive in an AI-enabled future.
- Industry Competitive Benefits: We offer comprehensive benefit plans to include flexible vacation, two company-wide Mental Health Days off, access to the Headspace app, retirement savings, tuition reimbursement, employee incentive programs, and resources for mental, physical, and financial wellbeing.
- Culture: Globally recognized, award-winning reputation for inclusion and belonging, flexibility, work-life balance, and more. We live by our values: Obsess over our Customers, Compete to Win, Challenge (Y)our Thinking, Act Fast / Learn Fast, and Stronger Together.
- Social Impact: Make an impact in your community with our Social Impact Institute. We offer employees two paid volunteer days off annually and opportunities to get involved with pro-bono consulting projects and Environmental, Social, and Governance (ESG) initiatives.
- Making a Real-World Impact: We are one of the few companies globally that helps its customers pursue justice, truth, and transparency. Together, with the professionals and institutions we serve, we help uphold the rule of law, turn the wheels of commerce, catch bad actors, report the facts, and provide trusted, unbiased information to people all over the world.
In the United States, Thomson Reuters offers a comprehensive benefits package to our employees. Our benefit package includes market competitive health, dental, vision, disability, and life insurance programs, as well as a competitive 401k plan with company match. In addition, Thomson Reuters offers market leading work life benefits with competitive vacation, sick and safe paid time off, paid holidays (including two company mental health days off), parental leave, sabbatical leave. These benefits meet or exceeds the requirements of paid time off in accordance with any applicable state or municipal laws. Finally, Thomson Reuters offers the following additional benefits: optional hospital, accident and sickness insurance paid 100% by the employee; optional life and AD&D insurance paid 100% by the employee; Flexible Spending and Health Savings Accounts; fitness reimbursement; access to Employee Assistance Program; Group Legal Identity Theft Protection benefit paid 100% by employee; access to 529 Plan; commuter benefits; Adoption & Surrogacy Assistance; Tuition Reimbursement; and access to Employee Stock Purchase Plan.Thomson Reuters complies with local laws that require upfront disclosure of the expected pay range for a position. The base compensation range varies across locations. For any eligible US locations, unless otherwise noted, the base compensation range for this role is $70,800 USD - $131,400 USD. Base pay is positioned within the range based on several factors including an individual’s knowledge, skills and experience with consideration given to internal equity. Base pay is one part of a comprehensive Total Reward program which also includes flexible and supportive benefits and other wellbeing programs. This role may also be eligible for an Annual Bonus based on a combination of enterprise and individual performance.
About Us
Thomson Reuters informs the way forward by bringing together the trusted content and technology that people and organizations need to make the right decisions. We serve professionals across legal, tax, accounting, compliance, government, and media. Our products combine highly specialized software and insights to empower professionals with the data, intelligence, and solutions needed to make informed decisions, and to help institutions in their pursuit of justice, truth, and transparency. Reuters, part of Thomson Reuters, is a world leading provider of trusted journalism and news.
We are powered by the talents of 26,000 employees across more than 70 countries, where everyone has a chance to contribute and grow professionally in flexible work environments. At a time when objectivity, accuracy, fairness, and transparency are under attack, we consider it our duty to pursue them. Sound exciting? Join us and help shape the industries that move society forward.
As a global business, we rely on the unique backgrounds, perspectives, and experiences of all employees to deliver on our business goals. To ensure we can do that, we seek talented, qualified employees in all our operations around the world regardless of race, color, sex/gender, including pregnancy, gender identity and expression, national origin, religion, sexual orientation, disability, age, marital status, citizen status, veteran status, or any other protected classification under applicable law. Thomson Reuters is proud to be an Equal Employment Opportunity Employer providing a drug-free workplace.
Thomson Reuters makes reasonable accommodations for applicants with disabilities, including veterans with disabilities, and for sincerely held religious beliefs in accordance with applicable law. If you reside in the United States and require an accommodation in the recruiting process, you may contact our Human Resources Department at [email protected]. Disability accommodations in the recruiting process may include things like a sign language interpreter, making interview rooms accessible, providing assistive technology, or other relevant accommodations. Please note this email is not intended for general recruitment questions and we will promptly respond to inquiries regarding accommodations. More information on requesting an accommodation here.
Learn more on how to protect yourself from fraudulent job postings here.
More information about Thomson Reuters can be found on thomsonreuters.com

