Hong Kong, Hong Kong SAR, Hong Kong

Reliability Engineer (Trading Platforms)

 Job Description:

Job Description:

  • Automate repeatable triage workflows to help first-line teams respond faster and more consistently (e.g., alert enrichment, routing, correlation, and operational runbooks).
  • Identify monitoring/alerting gaps and drive improvements in visibility and alert quality.
  • Track reliability and availability across critical trading applications and their dependencies. Partner with users, development teams, and IT to pinpoint where service levels are degrading.
  • Triage incoming alerts, issues, and escalations—assessing impact, urgency, and ownership.
  • Determine when incident criteria are met, declare incidents, and act as Incident Commander.
  • Coordinate responders and stakeholders; keep incident calls focused on facts, mitigation, and recovery.
  • Maintain clear timelines, actions, and status updates throughout the incident lifecycle.
  • Recover and stabilize systems using approved runbooks. Escalate cleanly through the defined support/development path when the issue exceeds documented recovery steps.
  • Support post-incident review (PIR) follow-ups and recurring issue reviews.
  • Ensure smooth handovers across EMEA, AMER, and APAC using a single global model: one incident standard and one handover process.

Requirements:

  • Experience in production operations, SRE, NOC/command center, trading operations, or a comparable first-line technical role—ideally in a trading, financial services, or other latency-sensitive environment.
  • Strong triage and prioritization skills: you can separate facts from assumptions under pressure and keep the response moving.
  • Clear communication (verbal and written): status updates are understandable to both traders and engineers.
  • Broad technical understanding (not just deep specialist knowledge): enough to collaborate effectively across domains and interfaces.
  • Solid Linux and networking fundamentals, plus the ability to quickly interpret alerts, logs, dashboards, and symptoms.
  • Working knowledge of common operational tasks across adjacent teams (application support, infrastructure, connectivity, data).
  • Familiarity with incident and observability tooling (e.g., PagerDuty or equivalent, Jira Service Management or equivalent, Grafana, Prometheus, log search).
  • Scripting/automation skills (Python preferred; Bash and Go are a plus), applied to triage, enrichment, routing, and correlation (not product code).
  • Exposure to containerized/cloud-hosted production environments (Kubernetes, Docker, GCP) is a plus.
  Required Skills:

Environment Handover Go APAC Data Routing Connectivity Prometheus Mitigation Steps Support Search Development Interfaces Grafana Financial Services Gcp Operations Trading Ownership Bash Timelines Reviews Reliability Availability Infrastructure Networking Automation Kubernetes Pressure Linux Docker JIRA Python Communication Management