Site Reliability Engineer
2 days ago
As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of continuous learning and accountability.
Key Responsibilities:
- Design and implement resilient system architectures that support high availability and scalability.
- Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
- Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
- Conduct thorough post-mortem analyses following incidents, driving continuous improvement through root cause identification and solution implementation.
- Collaborate with development and operations teams to establish best practices in system reliability and incident management.
- Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines).
- Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs), maintaining high standards of service delivery.
- Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements.
- Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.
Qualifications:
- Proficiency in programming languages such as Python, Golang, Java, or similar, focusing on operational efficiency.
- Minimum experience of 3 years and above in related field.
- Demonstrated experience in system architecture and design, prioritizing reliability, and scalability.
- Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems.
- Experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management.
- Strong expertise in Linux system administration.
- Proven experience in troubleshooting application support issues with a focus on performance and connectivity.
- Familiarity with networking concepts and effective troubleshooting techniques.
- Excellent problem-solving abilities and a proactive approach to operational challenges.
- Ability to work independently while effectively collaborating within a team environment.
Preferred Skills:
- Familiarity with monitoring tools and performance optimization techniques.
- Experience in scripting or automation for system administration tasks.
- Knowledge of networking concepts and troubleshooting methodologies.
- Hands-on knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud) and their services.
- Familiarity with DevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization.
-
Site Reliability Engineer
2 weeks ago
Kuala Lumpur, Kuala Lumpur, Malaysia Kneat Full time 80,000 - 120,000 per yearSite Reliability Engineer – Kuala Lumpur, MalaysiaKneat enables regulated organizations to move from paper-based validation to intelligent, digitized, paperless solutions. And we do it through the ongoing development of a powerful, purpose-built software platform. In 2014, after eight years of intensive software development, we launched Kneat Gx—the...
-
Site Reliability Engineer
2 weeks ago
Kuala Lumpur, Kuala Lumpur, Malaysia Kneat Full time 80,000 - 120,000 per yearSite Reliability Engineer – Kuala Lumpur, MalaysiaKneat enables regulated organizations to move from paper-based validation to intelligent, digitized, paperless solutions. And we do it through the ongoing development of a powerful, purpose-built software platform. In 2014, after eight years of intensive software development, we launched Kneat Gx—the...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia VCB Malaysia Berhad Full time 144,000 - 156,000 per yearOverview:As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia PeopleScope Full time 60,000 - 120,000 per yearSite Reliability EngineerJob Description:Ability to debug scripts and automate routine tasks in OS, network, database or application servers. Coding experience beyond simple scripts; Experience in Devops process, programming knowledge in at least one of the following languages: Java, Python, or Go; Scripting skills in at least of the following:...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia Abhidi Solution Private Limited Full time 120,000 - 180,000 per yearJob Title: Site Reliability Engineer (SRE)Job Type: Permanent positionWork Location Kuala LumpurResponsibilities:Strong hands-on experience with VMware solutionsStrong experience with patch management for OS & middlewareExperience in VMware server templating/blueprints (RedHat & Windows)Experience with Infrastructure-as-Code, orchestration, configuration...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia Hunters International Full time 19,000 per yearOverview:As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia Unison Group Full time 120,000 - 240,000 per yearAs a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia Aisling Group Full time 90,000 - 120,000 per yearCOMPANY PROFILE: Our client is a Tech Ecommerce Scale-Up that provides a single platform for customers to shop for the best price online. Not only that, they also provide data and insights to customers on latest trends and e-commerce sector. They are looking for a Site Reliability Engineers (SREs) who are responsible for keeping all services and...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia Swift Transportation Full time 80,000 - 120,000 per yearABOUT USWe're the world's leading provider of secure financial messaging services, headquartered in Belgium. We are the way the world moves value – across borders, through cities and overseas. No other organisation can address the scale, precision, pace and trust that this demands, and we're proud to support the global economy. We're unique too. We were...
-
Site Reliability Engineer
2 days ago
Kuala Lumpur, Kuala Lumpur, Malaysia Guidewire Software Full time 120,000 - 240,000 per yearSummaryAt Guidewire, we deliver the software that Property and Casualty (P&C) insurance companies rely on to protect their customers during crises, natural disasters, accidents, and cyber risks. Our core applications enable insurers to sell and underwrite policies, settle claims, and bill their customers. We also offer a suite of innovative products for data...