368,910open jobs
9,449companies
47,822added this week
Browse all
Location
In office
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
Softtek is a Mexican technology company founded in 1982 and headquartered in Monterrey. It provides software development, digital transformation and managed services worldwide. The company pioneered the nearshore delivery model from Latin America.

Proficient Mission-Critical Application Monitoring & Support Engineer

Responsibilities

Support and maintain mission-critical applications in a production environment. The role focuses on proactively monitoring application health, managing alerts and incidents, performing advanced troubleshooting, analyzing logs, conducting root cause analysis (RCA), and ensuring compliance with strict SLAs. The candidate will work closely with application teams, infrastructure teams, vendors, and business stakeholders to maintain service availability, performance, and reliability.

Responsibilities include monitoring dashboards and business-critical transactions, investigating and resolving incidents, performing log and trend analysis, tuning alerts to reduce noise and improve monitoring effectiveness, documenting activities in ServiceNow, leading or contributing to RCAs, maintaining operational procedures and SOPs, and driving continuous improvement initiatives. The ideal candidate should demonstrate strong analytical and troubleshooting skills, a high sense of urgency, operational ownership, and the ability to communicate effectively during critical situations.

Preferred Qualifications

  • 3 to 5+ years of experience in Application Support, Production Support, Monitoring Operations, SRE (Site Reliability Engineering) or IT Operations.
  • Experience supporting mission-critical applications in production environments.
  • Hands-on experience with monitoring and observability platforms such as New Relic, Dynatrace, Datadog, Splunk, Grafana, AppDynamics, or similar tools.
  • Strong experience in application and infrastructure log analysis and troubleshooting.
  • Experience managing incidents in SLA-driven environments.
  • Knowledge of SQL and database troubleshooting.
  • Understanding of APIs, integrations, and distributed application architectures.
  • Experience with cloud environments such as Azure, AWS, or Google Cloud Platform.
  • Familiarity with ServiceNow or similar ITSM platforms.
  • Knowledge of ITIL Incident Management, Problem Management, and Change Management processes.
  • Experience participating in or leading Root Cause Analysis (RCA) activities.
  • Experience supporting high-availability, 24x7 operational environments.
  • Excellent verbal and written communication skills in English.

Required Technical Skills

  • Monitoring and observability tools (New Relic, Dynatrace, Datadog, Splunk, Grafana, AppDynamics, or equivalent).
  • Log analysis and troubleshooting.
  • SQL and database querying.
  • Incident management and escalation processes.
  • Application performance monitoring (APM).
  • Basic cloud troubleshooting (Azure, AWS, GCP).
  • ServiceNow or ITSM platforms.
  • Application infrastructure fundamentals (web servers, APIs, integrations, middleware).
  • Dashboard creation and operational reporting.
  • Alert tuning and threshold optimization.

Preferred Certifications

  • ITIL Foundation Certification.
  • New Relic Certified Professional.
  • Splunk Core Certified User/Power User.
  • Dynatrace Associate Certification.
  • Microsoft Azure Fundamentals (AZ-900).
  • AWS Cloud Practitioner.
  • Google Associate Cloud Engineer.
  • Site Reliability Engineering (SRE) training or certification.
  • ServiceNow Fundamentals Certification.

Key Competencies

  • Strong analytical and problem-solving skills.
  • Incident ownership and accountability.
  • Sense of urgency and ability to work under pressure.
  • Root cause analysis and critical thinking.
  • Effective stakeholder communication.
  • Customer-focused mindset.
  • Operational leadership during critical incidents.
  • Risk identification and proactive issue prevention.
  • Continuous improvement mindset.
  • Collaboration across technical and business teams.
  • Attention to detail and execution discipline.

Success Profile

A successful candidate is someone who proactively identifies risks before they impact the business, effectively manages incidents through resolution, performs detailed log and trend analysis, communicates clearly with stakeholders, drives RCA and corrective actions, and continuously improves monitoring and operational processes. They take full ownership of production issues, demonstrate strong technical judgment, and contribute to maintaining the stability, availability, and performance of mission-critical applications.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,910 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$143k – $274k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Antonio • Charlotte • Colorado Springs • Plano • Phoenix
DevOps
IAM
Cybersecurity
CyberArk
Microsoft Entra ID
PCI DSS
Robotics
Path Planning
Management
ServiceNow
Apply
$126k – $195k per year • Equity • In office • Full-Time • 2+ years exp • Bachelor's Degree • San Diego
Java
SQL
Databases
MariaDB
MySQL
Oracle
PostgreSQL
Trino
Mobile
Clean Architecture
JUnit
DevOps
CI/CD
Management
ServiceNow
QA
Selenium
TestNG
Apply
IT Support Specialist 3 hours ago
$55k – $127k per year (Estimated) • In office • Full-Time • Associate's Degree • Los Angeles
DevOps
Azure
Cybersecurity
HIPAA
Least Privilege
Microsoft Entra ID
Management
Jira
ServiceNow
Marketing
Zendesk
Apply
$232k – $405k per year • Equity • In office • Full-Time • 10+ years exp • PhD • Santa Clara
Python
AI/ML
AI Agents
Computer Vision
DPO
Fine-tuning
Function Calling
GRPO
Knowledge Distillation
LLM
Model Context Protocol
Multimodal AI
Post-training
PyTorch
Reinforcement Learning
SFT
Synthetic Data
Management
ServiceNow
Apply
$143k – $243k per year • Equity • In office • Full-Time • 5+ years exp • San Francisco
Go
JavaScript
Python
AI/ML
AI Agents
Context Engineering
LLM
Prompt Engineering
DevOps
AWS
Azure
Rest API
Cybersecurity
FedRAMP
Okta
Management
Jira
ServiceNow
Marketing
Zendesk
Apply
POS TP.NET Leader 1 day ago
$28k – $73k per year (Estimated) • Remote/Hybrid • 5+ years exp • Bengaluru
SQL
C#
C#
.NET
Databases
MS SQL
DevOps
SLI/SLO/SLA
Management
ServiceNow
Apply
Remote/Hybrid • 7+ years exp • Bachelor's Degree
C#
JavaScript
Python
SQL
TypeScript
C#
ASP.NET Core
Blazor
Dapper
Databases
Oracle
Frontend
Angular
Bootstrap
JQuery
Vue.js
DevOps
AWS
Azure
Azure DevOps
CI/CD
Git
Apply
$33k – $107k per year (Estimated) • In office • Monterrey
TypeScript
JavaScript
Frontend
Ant Design
GraphQL
Lighthouse
Material UI
Next.js
React.js
Redux
Tailwind CSS
Mobile
State Management
DevOps
Azure
Azure DevOps
CI/CD
QA
Cypress
Jest
Apply
$30k – $101k per year (Estimated) • Remote/Hybrid • 4+ years exp • Madrid
Java
Java
Spring Boot
Databases
Apache Kafka
DevOps
Dynatrace
Apply
$53k – $120k per year (Estimated) • Remote/Hybrid • 5+ years exp • Madrid
Java
Java
Spring Boot
Apply
See all jobs
This is one of many
368,910 more open roles from verified company boards, updated every day.