Job Description
Role: Sr Site Reliability Engineer (SRE) \n
Experience: 8-10 years
Location: Sydney Key Technologies \n
AWS | EKS | ECS | EC2 | Lambda | GitLab | Terraform | Ansible | Python | Bash | Linux | Kubernetes | Docker | PostgreSQL | Oracle | Vault | CloudWatch | OpenTelemetry | Grafana | Splunk | ELK | SLI | SLO | SLA | DevOps | Site Reliability Engineering
Role Summary \n
We are seeking an experienced AWS DevOps & Site Reliability Engineer (SRE) to build, automate, operate, and continuously improve cloud platforms and mission-critical applications. The role focuses on AWS cloud engineering, infrastructure automation, CI/CD, platform reliability, observability, incident management, resiliency, and operational excellence within a highly regulated enterprise environment.
\n
The successful candidate will drive automation, reliability, performance, scalability, and availability of cloud-native platforms while partnering with engineering teams to improve release quality, operational stability, and customer experience.
Key Responsibilities \n
- \n
- Design, build, and maintain AWS cloud infrastructure and platform services. \n
- Develop and manage CI/CD pipelines using GitLab. \n
- Support cloud migration and platform modernization initiatives. \n
- Implement Infrastructure as Code (Terraform) and configuration management (Ansible). \n
- Automate deployments, environment provisioning, database refreshes, and operational processes. \n
- Manage secrets, certificates, access controls, and cloud security controls. \n
- Develop reusable infrastructure modules, deployment standards, and platform engineering patterns. \n
Site Reliability Engineering \n
- \n
- Define and implement reliability engineering practices, operational standards, and platform blueprints. \n
- Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs). \n
- Improve platform reliability, scalability, resilience, and fault tolerance for mission-critical applications. \n
- Drive proactive reliability improvements through automation and elimination of operational toil. \n
- Conduct capacity planning, performance tuning, and resource optimization activities. \n
- Support disaster recovery planning, backup validation, resilience testing, and business continuity initiatives. \n
- Participate in production readiness reviews and ensure operational requirements are embedded into solution designs. \n
Observability & Monitoring \n
- \n
- Design and implement monitoring, logging, alerting, and observability frameworks. \n
- Build and maintain dashboards, health checks, metrics, and operational reporting. \n
- Enhance end-to-end system visibility using CloudWatch, Grafana, Splunk, ELK, OpenTelemetry, or similar technologies. \n
- Drive alert tuning and noise reduction to improve operational effectiveness. \n
- Establish monitoring standards across applications, infrastructure, databases, and integration components. \n
Incident & Operational Management \n
- \n
- Lead incident triage, troubleshooting, root cause analysis, and post-incident reviews. \n
- Support production systems and provide timely resolution of infrastructure and application issues. \n
- Drive reduction of Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR). \n
- Develop operational runbooks, knowledge articles, and recovery procedures. \n
- Collaborate with engineering and support teams to implement preventative actions and continuous improvement initiatives. \n
- Participate in on-call and major incident management activities where required. \n
- Promote DevOps, SRE, and Platform Engineering best practices. \n
- Collaborate with development teams to improve deployment reliability and automation maturity. \n
- Integrate security, observability, and operational controls into CI/CD pipelines. \n
- Support release management, change governance, and deployment strategies. \n
- Champion Infrastructure as Code, automation, and self-service platform capabilities. \n
Mandatory Skills AWS Cloud \n
- \n
- AWS EC2, ECS/EKS, Lambda, VPC, IAM, S3, CloudWatch \n
- RDS/Aurora \n
- Route53 \n
- Secrets Manager \n
- AWS Networking and Security Services \n
DevOps & Automation \n
- \n
- Terraform \n
- Ansible \n
- Bash and Python scripting \n
- Infrastructure Automation \n
Site Reliability Engineering \n
- \n
- Production Support and Incident Management \n
- SLI / SLO / SLA implementation and management \n
- Reliability Engineering practices \n
- High Availability and Fault-Tolerant Architecture Design \n
- Capacity Planning and Performance Optimization \n
- Disaster Recovery and Business Continuity \n
- Root Cause Analysis and Problem Management \n
Observability \n
- \n
- Monitoring, Logging and Alerting \n
- CloudWatch \n
- Grafana \n
- Splunk / ELK \n
- OpenTelemetry \n
- Operational Metrics and Dashboarding \n
Platform & Containers \n
- \n
- Containerisation (Docker) \n
- Networking and Infrastructure Troubleshooting \n
Security \n
- \n
- Secrets Management Solutions (Vault preferred) \n
- IAM and Access Controls \n
- Cloud Security Best Practices \n
Preferred Skills \n
- \n
- AWS Certified Solutions Architect / DevOps Engineer Certification \n
- Experience with Aurora PostgreSQL and Oracle migration programs \n
- Chaos Engineering and Resilience Testing \n
- Experience implementing SRE operating models \n
- Service Mesh technologies \n
- Financial Services, Banking, Capital Markets, or other regulated industries \n
- ITIL processes including Incident, Problem, Change and Release Management \n
#J-18808-Ljbffr