Site Reliability Engineer
Microsoft
United States · Posted Jul 15
Job description
Responsibilities: The role involves designing, deploying, and monitoring Azure infrastructure while acting as a Designated Responsible Individual (DRI) for live site operations. Responsibilities include reducing incident volume, implementing mitigations for complex issues, and ensuring high availability and security of cloud services.
Requirements: Requires a Bachelor's or Master's degree in Computer Science or a related field with 1-5+ years of experience in software, network, or systems engineering. Candidates must have experience managing physical infrastructure and be able to pass a Microsoft Cloud Background Check.
Key skills: Site Reliability Engineering, Azure, Cloud Services, Distributed Systems, Network Engineering, Systems Administration, Software Engineering, Monitoring, Security, Automation, Physical Infrastructure Management, GPU Support, InfiniBand, Incident Management, Service Architecture, Project Management
Keywords: Azure, SRE, Cloud Computing, AI Infrastructure, Distributed Systems, Control Plane, Data Plane, DRI, SLA, Postmortem, GPU, InfiniBand, Observability, Scalability, Physical Infrastructure, Microsoft Cloud, Network Engineering, Systems Administration, Automation, Security Screening