661,377open jobs
38,546companies
98,203added this week
Browse all
Salary
$22k – $53k per year (Estimated)
Location
In office (Taguig)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
EviSmart is Autopilot for dental lab operations. Scans in, cases out. A universal inbox and zero-touch workflows cut admin chaos so you scale without headcount.

• On-site • Full-Time • Night Shift

Why EviSmart

  • 300 people. Two hubs: Vancouver HQ and Manila operations.
  • 145% year-over-year SaaS growth - the market is responding.
  • 28 countries. One platform. The dental industry's Autopilot.
  • An in-house AI model research and development team building proprietary intelligence.

How We Work

  • We ship before we're 100% certain. We write things down because we have two offices and memory is lossy. We debate loudly and move without resentment. We treat the customer's real problem as more important than an elegant internal process. If you've spent time waiting for permission to try something obvious - you'll notice the difference here immediately.

When something goes wrong in production overnight, the Production Support & Incident Response Lead is the person leading the response.

EviSmart is looking for a Production Support & Incident Response Lead  to own incident response and platform continuity during our night operations.

This is not a role where you simply monitor dashboards, create a ticket, and wait for Engineering.

You will be the lead incident response person on shift. You are expected to investigate first, understand what is happening, determine the safest way to restore operations, bring in the right technical people when necessary, and remain accountable for the incident until the platform is stable.

Our platform supports 2,000+ dental labs, so an issue in production can quickly become a real business problem for our customers. The goal is simple: keep cases moving and minimize disruption.

What you'll own

  • You will be responsible for the health and continuity of the platform during your coverage.

That means:

  • Lead the response to production incidents during the night shift

  • Continuously monitor platform health and act on warning signs before they become customer-impacting problems

  • Personally perform first-line investigation using logs, dashboards, monitoring tools, diagnostic commands and available system access

  • Determine the impact and likely source of an issue before escalating

  • Look for safe workarounds or restoration options when a permanent fix is not immediately available

  • Decide when an issue can be handled at your level and when Engineering, DevOps or another specialist needs to be brought in

  • Command the incident even after technical teams become involved: keep people aligned, decisions moving and communication clear

  • Keep Application Support and other stakeholders informed during active incidents

  • Document incidents, root causes, workarounds and follow-up actions

  • Make sure recurring issues don't simply become accepted problems

  • Provide a complete handoff to the daytime team with nothing dropped overnight

  • Develop one Application Support teammate into a reliable backup who can eventually handle routine night triage independently

The role's ownership of monitoring, incident command, proactive customer communication, post-mortems and backup development is explicit in the operating playbook.

What this role is NOT • This is not  a traditional Service Delivery Manager or ITIL governance position.

It is also not a pure DevOps or Software Engineering role. You don't need to be the person who writes the permanent code fix for every problem. But you do need enough technical depth to investigate intelligently before asking someone else to solve it.

If your normal incident process is: Alert → Create ticket → Escalate → Wait ...  then this probably isn't the right role.

We're looking for someone whose instinct is closer to:

Detect → Investigate → Isolate → Restore or Work Around → Escalate Intelligently → Command Through Resolution → Prevent Recurrence

The kind of person we're looking for

You may currently be a: Senior Application Support Engineer, L2/L3 Application Support Engineer, Production Support Engineer, Application Operations Engineer, Technical Operations Engineer, Platform Support Engineer, or similar.

More important than your current title is how you operate.

You are someone who:

  • Has personally supported live production applications

  • Can investigate an unfamiliar production problem without immediately needing someone to tell you what to check

  • Is comfortable working with logs, dashboards, APIs, databases and monitoring/observability tools

  • Understands enough infrastructure and application behavior to distinguish between likely application, database, API/connectivity and infrastructure problems

  • Thinks about business continuity, not only technical resolution

  • Can make sensible decisions with incomplete information

  • Knows when a workaround is safer and faster than waiting for the perfect fix

  • Stays calm when customers are affected and several teams are involved

  • Communicates clearly during incidents without creating noise

  • Can challenge or direct technical teams when an incident needs movement

  • Notices patterns and asks why the same problem keeps happening

  • Doesn't need constant hand-holding

  • Is comfortable being accountable when they are the most senior incident-response person available

Here's a good way to know whether you'll enjoy this role.

It's 2 AM. A production issue is preventing a customer from processing cases. The permanent fix requires an engineer who isn't immediately available.

What do you do?

We're looking for someone who doesn't stop at "I'll escalate it."

We want someone who starts asking:

What's actually broken?

What's the business impact?

What changed?

What can I verify myself?

Can I safely restore the previous working state?

Is there another way to keep the customer's operation moving?

Who genuinely needs to be involved?

What can we do now instead of waiting until morning?

That's the mindset we're hiring for.

What success looks like

Your first 90 days are designed to progressively prove that we can trust you with the night.

First 30 days:  Learn the platform and prove you can troubleshoot real issues and identify the correct workaround without being walked through every step.

By 60 days:  Independently monitor the platform, recognize warning signals and own selected production tickets through resolution.

By 90 days:  Independently command night incidents from detection through restoration, handle more complex issues, proactively identify problems and effectively delegate to your trained backup.

Ultimately, success means >99% platform health during night operations, incidents declared quickly, complete post-mortems, no dropped night-to-day handoffs, and a backup capable of independently handling night triage.

The ideal person for this role will have:

  • Strong hands-on experience in Application Support, Production Support, SaaS Operations or a similar production environment

  • Experience supporting systems in a 24/7 or on-call environment

  • Real production incident troubleshooting experience

  • Experience with monitoring and observability tools such as Grafana, Datadog, CloudWatch, Splunk, Kibana, Azure Monitor or similar

  • Working knowledge of logs, APIs, SQL/databases, cloud environments and basic diagnostic tools

  • Experience with incident response, root-cause analysis and production releases

  • Strong judgment around escalation, risk and business continuity

  • The ability to communicate confidently with both technical teams and business stakeholders

Deep DevOps expertise is not required. What matters is that you can investigate intelligently, understand what you're seeing, take the safest action available at your level, and know when specialist intervention is genuinely necessary.

Important items to take note of before you apply:

This is a permanent night-shift role.

This is also an Individual Contributor role, not a traditional people-management position. You will act as the functional point person during night coverage and will help develop a designated backup, but you will not be joining to manage a large team.

The responsibility is significant because you are the person we need to trust when the daytime team isn't around.

If you're already strong in Application or Production Support and you're looking for an opportunity where you're given more ownership, more decision-making authority and the chance to become the person trusted to lead production incidents, we'd like to hear from you.

Apply and tell us about the toughest production problem you've personally solved.

Apply at https://www.evismart.com/careers

EviSmart 

  •   Philippines 
  •   evismart.com
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
661,377 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Taguig
$180k – $210k per year • Remote • Top Secret • Full-Time • 10+ years exp • Bachelor's Degree • United States
SQL
Databases
PostgreSQL
Apache Iceberg
Apache Kafka
AI/ML
Flink
DevOps
GCP
GitHub Actions
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Docker
Kubernetes
OpenStack
GitLab
Management
Confluence
Jira
Agile
Apply
$29k – $72k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Master's Degree • Pune • Chennai
Python
SQL
Scala
Databases
PostgreSQL
Oracle
Cassandra
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Copilot
Hadoop
Spark
Airflow
Dagster
Prefect
Flink
Devin
DevOps
GCP
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
AWS
Amazon S3
Amazon Kinesis
Apply
$25k – $51k per year (Estimated) • In office • Full-Time • Bengaluru • Pune • Mumbai • Chennai • Gurgaon
Python
Python
Flask
FastAPI
Django
AI/ML
LangChain
LlamaIndex
Vertex AI
Fine-tuning
RLHF
Quantization
Prompt Engineering
Computer Vision
NLP
PEFT
AWS Bedrock
Transformers
TensorFlow
PyTorch
LLM
OpenAI
Hugging Face
DevOps
GCP
Azure
Git
AWS
Docker
Apply
$70k – $182k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Burnaby
Python
SQL
Apply
$63k – $111k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Warsaw
DevOps
Terraform
Azure
CI/CD
Kubernetes
Bicep
Cybersecurity
Microsoft Sentinel
Microsoft Defender
MITRE ATT&CK
Microsoft Entra ID
Apply
$14k – $32k per year (Estimated) • In office • Taguig
Apply
$18k – $44k per year (Estimated) • In office • Taguig
JavaScript
SQL
C#
C#
.NET
Databases
MS SQL
AI/ML
Cursor
Claude
LLM
Lovable
Frontend
React.js
DevOps
Git
Management
Agile
Scrum
Apply
Automation QA 5 days ago
$12k – $33k per year (Estimated) • In office • Taguig
AI/ML
Cursor
Claude
LLM
DevOps
CI/CD
Management
Jira
Agile
Scrum
Apply
$22k – $50k per year (Estimated) • In office • 4+ years exp • Manila
Python
SQL
Databases
Databricks
Azure SQL Database
AI/ML
dbt
DevOps
Azure
Analytics
ETL/ELT
Fivetran
Azure Data Factory
Management
Jira
QuickBooks
Apply
$58k – $150k per year (Estimated) • In office • Vancouver
Apply
Sales Manager 8 hours ago
In office • Full-Time • Taguig
Apply
$13k – $29k per year (Estimated) • In office • Full-Time • Master's Degree • Taguig
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Taguig
Chips/EDA
OpenLane
Marketing
Salesforce
Apply
$23k – $51k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Taguig
Apply
$15k – $35k per year (Estimated) • In office • Taguig
JavaScript
Node JS
Node JS
Commander.js
Management
Jira
Apply
See all jobs
This is one of many
661,377 more open roles from verified company boards, updated every day.