- Work on enterprise-scale platforms and Kubernetes environments
- Drive automation and SRE initiatives with real business impact.
Our client is a well-established organisation in Hong Kong operating mission-critical enterprise systems. They are investing in modern platform technologies, automation and reliability engineering to support business-critical digital services.
Job Description
- Own the day-to-day operation, maintenance and support of enterprise-scale systems, including standalone and Kubernetes-based environments.
- Administer and support Kubernetes platforms, including deployment, configuration, monitoring, troubleshooting and operational maintenance.
- Review vendor deliverables, including technical documentation, source code, testing evidence and deployment materials.
- Coordinate with business users, vendors, infrastructure and platform teams to support releases, enhancements and incident resolution.
- Implement and enhance Site Reliability Engineering (SRE) practices, including monitoring, alerting, logging and incident management.
- Investigate production incidents, conduct root-cause analysis and drive preventive actions.
- Automate operational processes through scripting, standardisation and workflow improvements.
- Maintain operational documentation, runbooks, support procedures and knowledge-base materials.
- Provide occasional after-hours support for production releases, maintenance activities and incident handling. Compensation leave will be provided.
- Degree in Computer Science, Information Technology or related disciplines.
- Minimum 4 years of IT experience, including at least 2 years in system development, application implementation and/or enterprise system support.
- Hands-on experience in Kubernetes operations, including deployment, configuration, monitoring, troubleshooting and incident resolution.
- Kubernetes administration experience is advantageous.
- Familiarity with Docker and container technologies.
- Strong stakeholder management and coordination skills across multiple teams.
- Experience working with Agile/Scrum methodologies and tools such as Jira and Confluence.
- Good understanding of RESTful APIs, Single Sign-On (SSO) and microservices architecture.
- Proficiency in Bash and Python scripting.
- Strong automation mindset with experience improving operational efficiency.
- Experience using Ansible Playbooks for deployment, configuration or operational automation is highly preferred.
- Experience with monitoring, alerting and observability platforms is an advantage.
- Strong analytical, problem-solving and incident-management capabilities.
- Strong communication skills in English, Cantonese and Mandarin.
- Competitive monthly salary.
- Comprehensive all-inclusive benefits package.
Job ID JN-082026-7089872
We help organisations find and develop the talent that matters most. From experienced professionals and executive search to scaled recruitment and out...
More jobs From Michael Page
Michael Page
Hong Kong
Michael Page
Hong Kong
Michael Page
Hong Kong
Boost your career
Find thousands of job opportunities by signing up to eFinancialCareers today.More Jobs Like This