- Penang George Town Pulau Pinang Malaysia
Lokasi Kerja
Penerangan Kerja
Tanggungjawab
We are seeking a proactive and reliability-focused engineer to strengthen the operational excellence of Remote Care platforms. This role is responsible for ensuring system stability, performance, and availability by delivering operational support, incident management, and continuous service improvements.
Key Responsibilities
1. Operational Support
Support production operations and issue investigations.
Participate in troubleshooting and cross-functional escalations.
Ensure operational stability and timely issue resolution.
2. Incident Management
Lead and coordinate incident response activities for production systems and services.
Perform incident triage, impact assessment, and root cause analysis (RCA).
Drive timely resolution of critical issues and ensure effective stakeholder communication throughout the incident lifecycle.
Track recurring incidents, identify systemic issues, and implement preventive measures to reduce operational risk.
Maintain incident documentation, post-mortem reports, and continuous improvement actions.
3. System Performance Monitoring and Tuning
Monitor system health, performance metrics, and service availability across Remote Care platforms.
Analyze trends and proactively identify potential performance bottlenecks or reliability risks.
Collaborate with engineering teams to optimize application, database, and infrastructure performance.
Develop and enhance monitoring dashboards, alerting mechanisms, and operational KPIs.
Recommend and implement performance tuning initiatives to improve system stability, scalability, and customer experience.
4. Quality
Assist with monitoring and addressing system performance and quality issues.
Identify trends of system health, engage in data review and provide recommendations for improvement.
Provide updates to stakeholders with regards to incident reports and performance metrics.
We Are Looking For
Required
Bachelor’s degree in Software Engineering, Systems Engineering, IT, Computer Science, or related field.
Minimum 3 years of experience working in site reliability, software engineering, computer engineering or a related technical discipline.
Strong analytical, problem-solving, and data interpretation skills.
Excellent communication skills with the demonstrated ability to work effectively in cross-functional teams and translate technical complexity for non-technical stakeholders.
Experience in software quality, system operations, or reliability engineering.
Ability to work with complex systems and large datasets.
Strong communication and stakeholder management skills.
Proficiency in at least one systems or automation programming language (e.g., Python, PowerShell) for building tooling, automation, and operational system.
Preferred
Experience with metrics visualization and reporting tools.
Familiarity with automation frameworks and CI/CD pipelines.
Knowledge of cloud environments and distributed systems.
Experience in regulated industries such as medical devices or healthcare.
Peringatan Penting
Jangan pernah kongsikan maklumat bank atau kad kredit anda semasa memohon pekerjaan. Elakkan membuat sebarang pembayaran atau mengisi survey yang tidak berkaitan. Jika ada yang mencurigakan, sila laporkan iklan pekerjaan ini segera.