jobs in NEPTUNEZ SINGAPORE PTE. LTD.

Full Time Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible ) Jobs, Salary up to SGD 7,200 in NEPTUNEZ SINGAPORE PTE. LTD. Islandwide (Singapore) - Maukerja

Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )

NEPTUNEZ SINGAPORE PTE. LTD.

SGD7,200 - SGD7,200 Per Month

Islandwide (Singapore)

Share
Save

Working Location

  • Islandwide (Singapore) Singapore

Job Description

Responsibilities

Responsibilities

  • Manage and maintain highly available, scalable, and reliable production environments.
  • Lead incident management, root cause analysis (RCA), and problem management activities to ensure service stability.
  • Develop and maintain automation solutions using Python and Shell scripting to improve operational efficiency.
  • Implement and enhance monitoring, logging, and observability using Splunk and other enterprise monitoring tools.
  • Administer, troubleshoot, and optimize Linux servers and production infrastructure.
  • Provide L2/L3 production support and resolve critical application and infrastructure issues within SLA.
  • Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools.
  • Collaborate with development and infrastructure teams to improve system reliability, performance, and deployment processes.
  • Identify opportunities for process improvement and implement automation to reduce manual effort.
  • Create and maintain operational documentation, runbooks, and best practices for production support.

Requirements

  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • 8+ years of experience in Site Reliability Engineering (SRE), Production Support, or Infrastructure Operations.
  • Strong hands-on experience with Python automation and Shell scripting.
  • Strong experience in Dynatrace, AppDynamics
  • Hands on experience in Splunk for monitoring, log analysis, troubleshooting, and observability.
  • Strong experience in administering and troubleshooting Linux environments.
  • Experience in incident management, RCA, problem management, and production support for mission-critical systems.
  • Experience with cloud platforms such as AWS and/or Oracle Cloud Infrastructure (OCI).
  • Hands-on experience with Infrastructure-as-Code tools such as Terraform and configuration management tools like Ansible.
  • Experience in SQL, and application performance tuning is preferred.
  • Excellent analytical, troubleshooting, communication, and collaboration skills.

Important Information

Never provide your bank or credit card details when applying for jobs. Do not transfer any money or complete unrelated online surveys. If you see something suspicious, Report this Job ad.

Learn More