707,373open jobs
41,962companies
100,977added this week
Browse all
Salary
$97k – $162k per year
Location
Remote/Hybrid (Buffalo, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
M&T Bank is an American regional bank founded in Buffalo in 1856 to serve manufacturers and traders, and it has grown into one of the larger commercial banks in the northeastern United States. It concentrates on commercial and business banking, commercial real estate lending and retail branches across New York, the mid-Atlantic and New England, a footprint substantially enlarged by its 2022 acquisition of People's United Financial. Headquartered in Buffalo and listed on the New York Stock Exchange, it has a long-standing reputation for conservative credit underwriting through economic cycles.

Overview

The Technical Engineer serves as a senior production support and reliability engineering professional responsible for ensuring the availability, stability, performance, and operational excellence of critical business applications and platforms.

This role combines strong troubleshooting expertise with modern Site Reliability Engineering (SRE), observability, automation, cloud operations, and incident management practices. The Technical Engineer partners with Engineering, Architecture, Infrastructure, Security, Product, and Vendor teams to proactively identify operational risks, improve system resilience, accelerate incident resolution, and continuously enhance customer and employee experiences.

The ideal candidate possesses deep technical knowledge of application support, distributed systems, cloud technologies, monitoring platforms, automation tools, and modern operational practices. They are passionate about eliminating repetitive work through automation and leveraging AI-powered tools to improve operational efficiency and support outcomes.

Primary Responsibilities

Production Support & Incident Management

  • Serve as a technical escalation point for critical production incidents, outages, and service degradation events.
  • Lead troubleshooting and root cause analysis efforts across applications, integrations, infrastructure, cloud services, APIs, and supporting technologies.
  • Coordinate incident response activities involving application teams, infrastructure teams, vendors, and business stakeholders.
  • Restore service quickly while ensuring long-term corrective actions are identified and implemented.
  • Participate in major incident management processes and post-incident reviews.

Site Reliability Engineering (SRE)

  • Apply Site Reliability Engineering principles to improve platform reliability, scalability, resilience, and operational efficiency.
  • Define and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational performance metrics.
  • Drive reduction of operational toil through automation and process improvement.
  • Support production readiness reviews and operational acceptance processes.
  • Participate in disaster recovery, resiliency, failover, and business continuity testing.

Troubleshooting & Problem Management

  • Analyze complex system behavior using logs, metrics, traces, performance data, and monitoring tools.
  • Perform deep technical investigations across application, infrastructure, data, network, and cloud environments.
  • Identify recurring issues, trends, and systemic problems to reduce future incidents.
  • Lead root cause analysis (RCA) activities and implement preventive solutions.
  • Develop technical recommendations that improve system stability, performance, and reliability.

Observability & Monitoring

  • Design, implement, and optimize monitoring, alerting, logging, and observability solutions.
  • Develop dashboards and health indicators providing visibility into application and platform performance.
  • Partner with engineering teams to improve observability through instrumentation, distributed tracing, synthetic monitoring, and telemetry collection.
  • Continuously refine alerting strategies to reduce false positives and alert fatigue.
  • Establish operational health metrics and reliability reporting.

Automation & Scripting

  • Develop and maintain automation solutions that improve operational efficiency and service reliability.
  • Create scripts, tools, and workflows to automate diagnostics, health checks, remediation activities, and routine support tasks.
  • Leverage Infrastructure as Code (IaC) and automation frameworks where appropriate.
  • Drive continuous improvement through operational automation and self-healing capabilities.
  • Partner with engineering teams to integrate automation into deployment and operational workflows.

Desired Scripting Technologies:

  • PowerShell
  • Python
  • Bash/Shell
  • SQL
  • REST APIs
  • Workflow automation platforms

AI-Assisted Operations & Innovation

  • Leverage AI and Generative AI tools to improve incident analysis, troubleshooting, knowledge management, and operational efficiency.
  • Utilize AI-powered operational insights to identify patterns, anomalies, and emerging risks.
  • Contribute to development of intelligent support capabilities including chatbots, operational copilots, automated RCA generation, and knowledge recommendations.
  • Evaluate opportunities to improve production support through AI-enabled automation and predictive analytics.
  • Promote responsible AI practices aligned with enterprise governance and security requirements.

Support Playbooks & Knowledge Management

  • Develop, maintain, and continuously improve support runbooks, operational procedures, troubleshooting guides, and recovery playbooks.
  • Ensure support documentation remains accurate, actionable, and aligned with production environments.
  • Establish standardized operational processes supporting incident response and service recovery.
  • Capture lessons learned from incidents and incorporate improvements into support practices.
  • Build and maintain operational knowledge repositories to improve support consistency and reduce resolution times.

Cloud Operations

  • Support cloud-hosted and hybrid application environments, including Azure-based platforms and services.
  • Assist engineering teams in implementing resilient and observable cloud architectures.
  • Monitor cloud resource health, performance, utilization, and operational readiness.
  • Support cloud deployments, platform upgrades, and operational change activities.
  • Partner with cloud engineering teams on modernization and reliability initiatives.

Collaboration & Technical Leadership

  • Partner closely with Engineering, Product, Architecture, Infrastructure, Security, QA, and Vendor teams.
  • Review operational readiness of new systems and platform enhancements.
  • Mentor junior support engineers and provide technical guidance during incident response activities.
  • Promote operational excellence, reliability engineering principles, and continuous improvement practices.
  • Serve as a subject matter expert within assigned technology domains.

Governance, Risk & Compliance

  • Ensure production support activities comply with enterprise risk, security, regulatory, and audit requirements.
  • Identify operational risks and escalate issues appropriately.
  • Support implementation of internal controls and operational governance standards.
  • Participate in audit, compliance, and regulatory review activities as required.

Education & Experience Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • Minimum 5+ years of experience supporting enterprise applications, platforms, or infrastructure in production environments.
  • Experience troubleshooting complex application, infrastructure, integration, or cloud-related issues.
  • Strong knowledge of incident management, production support processes, and operational best practices.
  • Experience with monitoring, observability, logging, and alerting platforms.
  • Experience with scripting and automation technologies.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent communication and collaboration skills.

Preferred Qualifications

Reliability Engineering & Operations

  • Experience with Site Reliability Engineering (SRE) practices.
  • Experience defining and managing SLIs, SLOs, and operational metrics.
  • Experience supporting distributed systems, APIs, microservices, and cloud-native applications.
  • Experience performing production readiness reviews and operational assessments.

Observability & Monitoring

Experience with platforms such as:

  • Dynatrace
  • Splunk
  • Datadog
  • Azure Monitor
  • Grafana
  • Prometheus
  • OpenTelemetry
  • AppDynamics

Cloud Technologies

  • Experience supporting Azure cloud environments.
  • Familiarity with Azure App Services, AKS, Functions, Storage, Event Hub, Service Bus, and monitoring services.
  • Understanding of cloud security and operational best practices.

Automation & Scripting

  • Advanced PowerShell scripting.
  • Python development and automation.
  • REST API integration.
  • Infrastructure as Code concepts.
  • CI/CD tools and deployment pipelines.

AI & Modern Operations

  • Experience using Microsoft Copilot, Azure AI services, Microsoft Foundry, or similar AI-enabled platforms.
  • Familiarity with AI-assisted troubleshooting and operational analytics.
  • Experience implementing AI-enabled support workflows or knowledge management solutions.
  • Understanding of intelligent automation and operational copilots.

What Great Looks Like

A successful Production Reliability Engineer:

  • Resolves complex incidents quickly and effectively.
  • Proactively identifies and eliminates recurring production issues.
  • Uses automation and scripting to eliminate manual support work.
  • Builds comprehensive support playbooks and operational runbooks.
  • Leverages AI to accelerate troubleshooting and operational insights.
  • Establishes strong observability and monitoring capabilities.
  • Partners effectively with engineering teams to improve reliability.
  • Drives a culture of operational excellence through SRE principles.
  • Continuously improves customer experience through system stability, resilience, and performance.
M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $97,100.00 - $161,800.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

Location

Buffalo, New York, United States of America
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
707,373 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Buffalo
$116k – $194k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Buffalo
Python
SQL
PowerShell
Bash
AI/ML
Copilot
DevOps
Rest API
Splunk
OpenTelemetry
Datadog
Dynatrace
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Self-Healing
AppDynamics
Azure AKS
Incident Management
Apply
$87k – $198k per year • In office • TS/SCI • Full-Time • 8+ years exp • Bachelor's Degree • Chantilly
Python
JavaScript
PHP
TypeScript
SQL
C#
C++
Node JS
Apex
Bash
Lua
Visual Basic
C#
.NET
Apex
MuleSoft
Databases
MS SQL
Frontend
Angular
React.js
DevOps
Rest API
GCP
Azure
CI/CD
Windows Server
AWS
Analytics
Tableau
Informatica
Management
SharePoint
Agile
Apply
$87k – $198k per year • In office • TS/SCI • Full-Time • 5+ years exp • Bachelor's Degree • Aberdeen
Python
Go
JavaScript
PHP
PowerShell
C#
C++
Node JS
Bash
PHP
Drupal
C#
.NET
Frontend
React.js
DevOps
Azure DevOps
Azure
CI/CD
Windows Server
Jenkins
AWS
GitLab
Linux
Cybersecurity
Active Directory
Management
Agile
Apply
$84k – $208k per year (Estimated) • In office • San Francisco
Python
Bash
DevOps
Git
Bitbucket
Linux
Management
Confluence
Jira
Apply
SIM Triage Operator 2 hours ago
$45k – $108k per year (Estimated) • In office • Palo Alto
Python
SQL
Management
Jira
Apply
$140k – $233k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Buffalo
AI/ML
Agentic Workflows
Machine Learning
DevOps
Azure DevOps
Azure
CI/CD
Kubernetes
Platform Engineering
Azure AKS
Incident Management
Management
Agile
Apply
$140k – $233k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Buffalo
Python
PowerShell
AI/ML
Anomaly Detection
DevOps
Splunk
OpenTelemetry
Datadog
Dynatrace
Azure
CI/CD
Grafana
Platform Engineering
Incident Management
Apply
$44k – $74k per year • In office • Full-Time • 2+ years exp • High School Diploma • Farmington
Apply
$116k – $194k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Buffalo
Python
SQL
PowerShell
Bash
AI/ML
Copilot
DevOps
Rest API
Splunk
OpenTelemetry
Datadog
Dynatrace
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Self-Healing
AppDynamics
Azure AKS
Incident Management
Apply
$42k – $72k per year • In office • Full-Time • 1+ year exp • High School Diploma • Washington
Apply
$196k – $265k per year • Remote/Hybrid • Full-Time • 9+ years exp • Bachelor's Degree • Deerfield • Schaumburg • Evanston • Buffalo • Northbrook
Apply
$103k – $172k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Buffalo
Management
Agile
Apply
Design Drafter 1 day ago
$70k – $90k per year • Equity • In office • Full-Time • 3+ years exp • Associate's Degree • Buffalo
Apply
$70k – $95k per year • Equity • In office • Full-Time • Bachelor's Degree • Buffalo
VHDL
MATLAB
MATLAB
Simulink
Apply
$125k – $165k per year • Equity • Remote • Full-Time • 7+ years exp • Bachelor's Degree • Buffalo
Analytics
Microsoft Excel
Apply
See all jobs
This is one of many
707,373 more open roles from verified company boards, updated every day.