Resunova

Production Support & Incident Commander

Evismart
Taguig
Posted October 6, 2026
Apply on company site → See how your résumé matches →

Incident Commander | SaaS Production Operations

Permanent night shift • Individual Contributor • Full onsite – BGC, Taguig

This is not an escalation coordinator role. We need someone who can make real-time production recovery decisions when systems fail and customers are already affected.

When an incident happens, your job is not:

Follow procedure → identify resolver → escalate → wait → send updates.

Your job is:

Understand what is actually affected → determine the fastest safe recovery path → make the call → direct the teams involved → keep the customer operational → stay on it until production is restored.

What this actually looks like

EviSmart is a live SaaS platform with multiple applications, services, dependencies and operational workflows.

When something breaks, the answer will not always be written in a runbook.

You need to be able to determine:

• What is actually failing?

• What is the customer impact?

• Which systems or dependencies are involved?

• Is the normal recovery path working?

• How long should recovery reasonably take?

• Can Production/Application Support restore it?

• Do we actually need Engineering?

• Is there a workaround or manual process that keeps the customer operational?

• At what point do we stop waiting and change the recovery strategy?

Different failures have different recovery paths. During the incident, you own the clock and the decision.

If the normal recovery route isn't working, we expect you to recognize that quickly and change course rather than continue following a process that is no longer protecting the customer.

That can mean restoring the application, rolling something back, implementing a safe workaround, directing Support or Engineering, or switching operations to a manual process while the underlying issue is being resolved.

You stay accountable until production is operational again.

What we mean by Incident Commander

You command both the people and the recovery strategy.

Engineering may perform the permanent technical fix. Production Support may execute part of the recovery. Customer Support may need to communicate a workaround.

But during the incident, you are responsible for determining what needs to happen next and driving everyone toward recovery.

You need enough technical depth to investigate production issues yourself using things such as logs, monitoring, application behavior, databases, APIs, recent releases and system dependencies.

You also need enough judgment to know when to continue investigating, when to restore, when to escalate technically, and when the safest decision is to keep the customer moving another way.

Incident Management experience alone is not enough

You may have years of Incident Management experience and still find this role difficult.

Our production environment is a large, interconnected ecosystem. An issue in one part of the system can affect another application, workflow, team or customer process.

If most of your experience has been following established procedures, running bridges, coordinating resolver groups and escalating technical decisions to someone else, this is probably not the right role for you.

We need someone who can reason through an unfamiliar production problem even when:

the documentation is incomplete, the normal process isn't working, information is still coming in, teams disagree on the cause, and the person you would normally escalate to isn't available.

You won't know our ecosystem on Day 1. We can teach you that.

What is much harder to teach is the ability to analyze a complex system, make a defensible decision with incomplete information, take responsibility for that decision, and change course when the evidence tells you you're wrong.

If your background isn't a perfect match but you genuinely believe you can demonstrate that level of technical reasoning and judgment to us during the interview, we're willing to give you the opportunity to prove it.

We're likely to be interested if you come from

Production Support, Application Support, SaaS Operations, Site Reliability, Technical Operations, Incident Response, or another environment where you have personally handled business-critical production failures.

We're particularly interested in people who have worked with multiple interconnected applications or services and have personally made recovery decisions during live incidents.

The simplest test:

At 2 AM, production is down, the runbook isn't solving it, Engineering isn't available, and customers need to keep operating.

Can we trust you to figure out what needs to happen next, make the call, and own the outcome?

If yes, we want to talk to you.

Discover jobs matched to your résumé at Resunova.