Site Reliability Engineer (SRE)
Bright Vision Technologies
Jersey City, New Jersey, United States · Posted Jul 18
Job description
Responsibilities: Ensure the availability and performance of large-scale distributed systems by applying software engineering principles to infrastructure. Responsibilities include defining SLOs/SLIs, leading incident response, and automating operational toil to improve reliability.
Requirements: Requires a Bachelor's degree and 5+ years of experience in SRE, DevOps, or production engineering. Candidates must be proficient in Python, Go, or Java and have deep hands-on experience with Kubernetes and Linux at scale.
Key skills: Kubernetes, Python, Go, Prometheus, Grafana, CI/CD, Linux, Distributed Systems, Observability, Incident Response, Chaos Engineering, Terraform, Java, Bash, OpenTelemetry, ELK/EFK
Keywords: Site Reliability Engineering, SRE, DevOps, Kubernetes, Python, Go, Java, Linux, Prometheus, Grafana, OpenTelemetry, ELK, EFK, Datadog, CI/CD, Distributed Systems, SLO, SLI, Error Budgets, Chaos Engineering, Service Mesh, Istio, Linkerd, Consul, AWS, Azure, GCP, Capacity Planning, Performance Engineering, Incident Management