$2,939.00 Fixed
DataPro Services
Contract · Remote · Flexible hours
About the role
DataPro Services is scaling its B2B platform with a new multi‑tenant capability called TenantPulse. As a Site Reliability Engineer you will own the reliability, performance, and observability of this feature from design through production launch.
Key responsibilities
- Design and implement Terraform modules for isolated tenant environments on AWS.
- Set up Prometheus‑Grafana dashboards to monitor latency and error rates for TenantPulse.
- Automate canary deployments using Argo CD and Helm for zero‑downtime releases.
- Develop chaos‑engineering tests with Gremlin to validate fault tolerance across tenants.
- Collaborate with product owners to define SLOs and create alerting policies in PagerDuty.
- Document runbooks and hand‑off procedures for on‑call engineers.
Must-have skills
- 5+ years in site reliability or DevOps engineering.
- Deep experience with AWS services (EKS, RDS, S3).
- Proficiency in Docker, Kubernetes, and Helm.
- Infrastructure‑as‑code expertise using Terraform.
- Strong scripting skills in Bash/Python for automation.
Nice to have
- Experience with multi‑tenant SaaS architectures.
- Familiarity with service mesh technologies like Istio.
- Proposal: 0
- Less than 3 month
Sam Vasquez
,
Member since
Oct 27, 2025
Total Job