About the role
Reliability & Run is the team that gets paged when a customer's regulated workload is unhappy at four in the morning, and the team that spends the following month making sure it cannot happen the same way twice. We run from the Kwun Tong office, which is deliberate — it is twenty minutes from three of the data centres we work in and most of our customers' operations floors.
The staff role is new. We have strong senior SREs and a run book culture, but nobody whose job is the two-year view: what the platform should look like when we are running eighty environments instead of forty-two, and what we have to stop doing to get there.
You would work with the platform engineering leads, the security engineering team and, regularly, with customers' own SRE and change-management functions. On-call here is one week in six, paid, and the escalation path has never had a lone hero at the end of it.
What you'll do
- Set the reliability architecture across environments: SLOs, error budgets, capacity models and the failure modes we accept.
- Lead the response to major incidents and write the postmortems that leadership reads and acts on.
- Drive the observability platform — Prometheus, Grafana and OpenTelemetry — toward signals that predict rather than describe.
- Reduce toil measurably: identify the top five recurring manual interventions each quarter and remove them with code.
- Own the disaster-recovery and failover programme, including the live exercises customers' regulators ask to observe.
- Mentor five SREs and set the technical bar for the on-call rotation across both offices.
- Represent CloudPeak in customers' change advisory boards when a migration needs an engineer, not an account manager.
What we're looking for
- Eight or more years in SRE, platform or infrastructure engineering, including staff- or principal-level scope.
- Deep Kubernetes operational experience — you have broken and recovered a production cluster.
- Strong coding ability in Go or Python; automation here is software, not shell scripts pasted into a wiki.
- Infrastructure as code at scale with Terraform, including module design others depend on.
- A track record of leading incident response and writing postmortems that changed engineering practice.
- Business-level English; the reliability team operates entirely in English.
Nice to have
- Multi-cloud experience across AWS and Azure.
- Service mesh operations at scale (Istio or Linkerd).
- Experience with regulated change windows and HKMA cloud guidance.
What you get
- Paid on-call allowance and time back in lieu after major incidents
- 13th month salary plus discretionary bonus
- All certifications funded, including CKA and cloud architect tracks
- 25 days annual leave at this level
- Medical and dental for the family
- Hybrid — two days a week at the Kwun Tong office
Skills & keywords
Aisha Rahman
VP Engineering · reviews applications personally
Listing ID JOB-TECH-005 · Closes 18 Sept 2026 · HKjobs never asks candidates to pay a fee. Report this listing