Site Reliability Engineer (SRE)

Bright Vision Technologies

Jersey City, New Jersey, United States · Posted Jul 18


Job description

Responsibilities: Ensure the availability and performance of large-scale distributed systems by applying software engineering principles to infrastructure. Responsibilities include defining SLOs/SLIs, leading incident response, and automating operational toil to improve reliability.

Requirements: Requires a Bachelor's degree and 5+ years of experience in SRE, DevOps, or production engineering. Candidates must be proficient in Python, Go, or Java and have deep hands-on experience with Kubernetes and Linux at scale.

Key skills: Kubernetes, Python, Go, Prometheus, Grafana, CI/CD, Linux, Distributed Systems, Observability, Incident Response, Chaos Engineering, Terraform, Java, Bash, OpenTelemetry, ELK/EFK

Keywords: Site Reliability Engineering, SRE, DevOps, Kubernetes, Python, Go, Java, Linux, Prometheus, Grafana, OpenTelemetry, ELK, EFK, Datadog, CI/CD, Distributed Systems, SLO, SLI, Error Budgets, Chaos Engineering, Service Mesh, Istio, Linkerd, Consul, AWS, Azure, GCP, Capacity Planning, Performance Engineering, Incident Management

Land this job faster with Remote Job Match

Free account: browse thousands of remote roles, no degree needed. Upgrade to tailor your resume to each job with AI and prep for the interview.

Create free account

← Browse more remote jobs