Site Reliability Engineer
About the Role
Apply software engineering principles to operations to ensure our platforms are highly scalable, reliable, and fault-tolerant.
Responsibilities
- Establish and monitor SLIs, SLAs, and SLOs.
- Automate infrastructure provisioning and configuration management.
- Lead blameless post-mortems and implement incident prevention.
- Optimize system architecture to handle massive traffic spikes.
Requirements
- 5+ years of combined development and operations experience.
- Deep expertise in Kubernetes orchestration.
- Proficiency in programming (Go, Python, or Ruby).
- Strong experience with observability tools (Grafana, New Relic).
Nice to Have
- Experience with chaos engineering practices.
- Background in distributed systems architecture.
How to Apply
Send your resume to [email protected] with the subject line "Site Reliability Engineer – CA".
Apply Now