Overview
Site Reliability Engineer Kubernetes Expertise Jobs in Invermere, Canada at Net2Source (N2S)
Position: Site Reliability Engineer with Kubernetes Expertise
Location: Invermere
Become a crucial part of L1 Site Reliability Engineering focused on monitoring and automating operational tasks across enterprise applications. Leverage your skills with Kubernetes, APIs, and multi-cloud environments to ensure seamless performance.
This L1 Site Reliability Engineer role demands up to five years in IT operations, NOC, or SRE roles. You will be involved in monitoring systems using Grafana, Splunk, and Prometheus, while also triaging incidents and following standard runbooks for rapid resolution. Your expertise in automation with Python or Bash will streamline processes, enhancing operational workflow.
Key Responsibilities:
• Monitor systems with Grafana, Datadog, and AIOps tools
• Execute predefined runbooks for quick incident resolution
• Validate Kubernetes performance using dashboard metrics
• Collect and analyze logs for proactive issue detection
• Communicate effectively with stakeholders throughout incidents
Requirements:
• 2–5 years in IT operations or SRE roles
• Strong knowledge of Linux and networking principles
• Familiarity with AWS, Azure, or GCP
• Experience with observability tools and ITSM solutions
• Proven troubleshooting skills using structured methods
Utilize your technical skills in Kubernetes and automation to optimize enterprise applications in this dynamic site reliability role.
#J-18808-Ljbffr
Title: Site Reliability Engineer Kubernetes Expertise
Company: Net2Source (N2S)
Location: Invermere, Canada
Category: