$2,581.00 Fixed
Cascade Apps
Contract · Remote · Flexible hours
About the role
Cascade Apps is modernizing its customer‑facing platform to achieve sub‑second latency and 99.99% uptime. As a Site Reliability Engineer you will lead the performance and reliability overhaul, focusing on the new "PulseStream" service layer that powers real‑time analytics for clients.
Key responsibilities
- Instrument PulseStream with OpenTelemetry, create dashboards in Grafana, and set SLOs for latency and error rates.
- Automate failover testing in AWS using Terraform and Chaos Monkey scripts to validate resilience.
- Refactor CI/CD pipelines in GitHub Actions to include canary deployments and automated rollback triggers.
- Optimize Kubernetes pod autoscaling policies (HPA/VPA) based on custom metrics from Prometheus.
- Collaborate with backend teams to rewrite critical endpoints in Go, reducing average response time by 30%.
Must-have skills
- Extensive experience with Kubernetes and container orchestration.
- Proficiency in AWS services (EC2, RDS, S3, CloudWatch).
- Strong scripting skills in Bash/Python for automation.
- Deep understanding of SRE principles, monitoring, and incident response.
- Hands‑on with Terraform or CloudFormation for IaC.
Nice to have
- Experience with Service Mesh technologies (Istio, Linkerd).
- Familiarity with chaos engineering tools.
- Proposal: 0
- Less than 2 month
Carl Atkins
,
Member since
Oct 27, 2025
Total Job