Talent.com
CareCone Group
Sr Site Reliability Engineer (SRE)CareCone Group • Sydney, NSW, AU
Search for other jobs
Sr Site Reliability Engineer (SRE)

Sr Site Reliability Engineer (SRE)

CareCone Group • Sydney, NSW, AU
1 day ago
Job description

Job Description

Role: Sr Site Reliability Engineer (SRE) \n

Experience: 8-10 years

Location: Sydney Key Technologies \n

AWS | EKS | ECS | EC2 | Lambda | GitLab | Terraform | Ansible | Python | Bash | Linux | Kubernetes | Docker | PostgreSQL | Oracle | Vault | CloudWatch | OpenTelemetry | Grafana | Splunk | ELK | SLI | SLO | SLA | DevOps | Site Reliability Engineering

Role Summary \n

We are seeking an experienced AWS DevOps & Site Reliability Engineer (SRE) to build, automate, operate, and continuously improve cloud platforms and mission-critical applications. The role focuses on AWS cloud engineering, infrastructure automation, CI/CD, platform reliability, observability, incident management, resiliency, and operational excellence within a highly regulated enterprise environment.

\n

The successful candidate will drive automation, reliability, performance, scalability, and availability of cloud-native platforms while partnering with engineering teams to improve release quality, operational stability, and customer experience.

Key Responsibilities \n

    \n
  • Design, build, and maintain AWS cloud infrastructure and platform services.
  • \n
  • Develop and manage CI/CD pipelines using GitLab.
  • \n
  • Support cloud migration and platform modernization initiatives.
  • \n
  • Implement Infrastructure as Code (Terraform) and configuration management (Ansible).
  • \n
  • Automate deployments, environment provisioning, database refreshes, and operational processes.
  • \n
  • Manage secrets, certificates, access controls, and cloud security controls.
  • \n
  • Develop reusable infrastructure modules, deployment standards, and platform engineering patterns.
  • \n

Site Reliability Engineering \n

    \n
  • Define and implement reliability engineering practices, operational standards, and platform blueprints.
  • \n
  • Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
  • \n
  • Improve platform reliability, scalability, resilience, and fault tolerance for mission-critical applications.
  • \n
  • Drive proactive reliability improvements through automation and elimination of operational toil.
  • \n
  • Conduct capacity planning, performance tuning, and resource optimization activities.
  • \n
  • Support disaster recovery planning, backup validation, resilience testing, and business continuity initiatives.
  • \n
  • Participate in production readiness reviews and ensure operational requirements are embedded into solution designs.
  • \n

Observability & Monitoring \n

    \n
  • Design and implement monitoring, logging, alerting, and observability frameworks.
  • \n
  • Build and maintain dashboards, health checks, metrics, and operational reporting.
  • \n
  • Enhance end-to-end system visibility using CloudWatch, Grafana, Splunk, ELK, OpenTelemetry, or similar technologies.
  • \n
  • Drive alert tuning and noise reduction to improve operational effectiveness.
  • \n
  • Establish monitoring standards across applications, infrastructure, databases, and integration components.
  • \n

Incident & Operational Management \n

    \n
  • Lead incident triage, troubleshooting, root cause analysis, and post-incident reviews.
  • \n
  • Support production systems and provide timely resolution of infrastructure and application issues.
  • \n
  • Drive reduction of Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
  • \n
  • Develop operational runbooks, knowledge articles, and recovery procedures.
  • \n
  • Collaborate with engineering and support teams to implement preventative actions and continuous improvement initiatives.
  • \n
  • Participate in on-call and major incident management activities where required.
  • \n
  • Promote DevOps, SRE, and Platform Engineering best practices.
  • \n
  • Collaborate with development teams to improve deployment reliability and automation maturity.
  • \n
  • Integrate security, observability, and operational controls into CI/CD pipelines.
  • \n
  • Support release management, change governance, and deployment strategies.
  • \n
  • Champion Infrastructure as Code, automation, and self-service platform capabilities.
  • \n

Mandatory Skills AWS Cloud \n

    \n
  • AWS EC2, ECS/EKS, Lambda, VPC, IAM, S3, CloudWatch
  • \n
  • RDS/Aurora
  • \n
  • Route53
  • \n
  • Secrets Manager
  • \n
  • AWS Networking and Security Services
  • \n

DevOps & Automation \n

    \n
  • Terraform
  • \n
  • Ansible
  • \n
  • Bash and Python scripting
  • \n
  • Infrastructure Automation
  • \n

Site Reliability Engineering \n

    \n
  • Production Support and Incident Management
  • \n
  • SLI / SLO / SLA implementation and management
  • \n
  • Reliability Engineering practices
  • \n
  • High Availability and Fault-Tolerant Architecture Design
  • \n
  • Capacity Planning and Performance Optimization
  • \n
  • Disaster Recovery and Business Continuity
  • \n
  • Root Cause Analysis and Problem Management
  • \n

Observability \n

    \n
  • Monitoring, Logging and Alerting
  • \n
  • CloudWatch
  • \n
  • Grafana
  • \n
  • Splunk / ELK
  • \n
  • OpenTelemetry
  • \n
  • Operational Metrics and Dashboarding
  • \n

Platform & Containers \n

    \n
  • Containerisation (Docker)
  • \n
  • Networking and Infrastructure Troubleshooting
  • \n

Security \n

    \n
  • Secrets Management Solutions (Vault preferred)
  • \n
  • IAM and Access Controls
  • \n
  • Cloud Security Best Practices
  • \n

Preferred Skills \n

    \n
  • AWS Certified Solutions Architect / DevOps Engineer Certification
  • \n
  • Experience with Aurora PostgreSQL and Oracle migration programs
  • \n
  • Chaos Engineering and Resilience Testing
  • \n
  • Experience implementing SRE operating models
  • \n
  • Service Mesh technologies
  • \n
  • Financial Services, Banking, Capital Markets, or other regulated industries
  • \n
  • ITIL processes including Incident, Problem, Change and Release Management
  • \n

#J-18808-Ljbffr

Create a job alert for this search

Sr Site Reliability Engineer (SRE) • Sydney, NSW, AU