Senior Site Reliability Engineer

$127,200 - $245,900/Yr

Cribl - Lansing, MI

posted 5 months ago

Full-time - Senior
Remote - Lansing, MI

About the position

Cribl is on a mission to unlock the value of all observability data, and we are looking for a Senior Site Reliability Engineer to join our team. In this role, you will be part of a remote-first company that values collaboration, curiosity, and a customer-first approach. As a Senior Site Reliability Engineer, you will work closely with a team of technical engineers dedicated to delivering high-quality software while enjoying a fun and engaging work environment. Your contributions will be vital in envisioning, creating, deploying, testing, and shipping Cribl products that fundamentally change technology and empower our customers with control over their observability data. In this position, you will engage with various teams to enhance service delivery and reliability throughout the entire lifecycle of our systems. You will measure and monitor production systems, focusing on availability, latency, and overall health. Your role will involve investigating the causes of errors and instability in our production cloud services, driving teams towards operational excellence. You will collaborate with product and platform teams to advocate for changes that improve reliability, resilience, and observability, while also identifying and reducing toil through innovative automation solutions. Additionally, you will have on-call responsibilities to ensure the smooth operation of our services. This is an exciting opportunity for those who are passionate about reliability and have strong opinions on improving systems. If you enjoy being involved from conception to production and want to make a real impact in a rapidly growing company, we encourage you to apply and join our herd at Cribl.

Responsibilities

  • Engage with teams and improve service delivery and reliability across their entire lifecycle
  • Measure and monitor all production systems with an eye towards availability, latency and overall system health
  • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
  • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
  • Help identify and drive down toil with creative innovation and automation
  • Participate in on-call responsibilities

Requirements

  • Extensive experience with enterprise scale continuous delivery environments
  • 5+ years of experience as a DevOps or SRE
  • Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
  • Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible
  • Experience with sustainable incident response in a blameless environment
  • Knowledge of cloud platforms (prefer Azure) and container + orchestration technologies
  • Experience with APM and Observability and related tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.
  • Background in Linux Systems Engineering
  • Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc.
  • Comfortable with a high level of autonomy and working with a distributed team

Nice-to-haves

  • Knowledge of Azure
  • Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
  • A love for high quality and a knack for testing
  • Opinions about dashboards, metrics, and SLO's

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Short-term disability insurance
  • Life insurance
  • Paid holidays
  • Paid time off
  • Fertility treatment benefit
  • 401(k) plan
  • Equity
  • Eligibility for a discretionary company-wide bonus
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service