Senior Site Reliability Engineer

$127,200 - $245,900/Yr

Cribl - Raleigh, NC

posted 5 months ago

Full-time - Senior

Remote - Raleigh, NC

About the position

Cribl is on a mission to unlock the value of all observability data, and we are looking for a Senior Site Reliability Engineer to join our team. As a remote-first company, we empower our employees to do their best work from anywhere. In this role, you will be part of a collaborative and motivated engineering organization that is dedicated to delivering high-quality software while enjoying a fun and engaging work environment. You will have the opportunity to contribute to the development, deployment, testing, and shipping of Cribl products, which are designed to give users unprecedented control over their real-time data. As a Senior Site Reliability Engineer, you will engage with various teams to enhance service delivery and reliability throughout the entire lifecycle of our systems. Your responsibilities will include measuring and monitoring production systems to ensure availability, latency, and overall health. You will investigate the root causes of errors and instability in our cloud services, driving teams towards operational excellence. Additionally, you will collaborate with product and platform teams to advocate for changes that improve reliability, resilience, and observability. Your role will also involve identifying and reducing toil through innovative solutions and automation, as well as participating in on-call responsibilities. This position is ideal for individuals who are passionate about reliability and have strong opinions on how to improve systems. You will be involved from conception to design to development and beyond, providing your creative input on cloud scaling, reliability, and high availability. If you are ready to make a real impact and be part of a team that is fundamentally changing technology, we want to hear from you!

Responsibilities

Engage with teams and improve service delivery and reliability across their entire lifecycle
Measure and monitor all production systems with an eye towards availability, latency and overall system health
Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
Help identify and drive down toil with creative innovation and automation
Participate in on-call responsibilities

Requirements

Extensive experience with enterprise scale continuous delivery environments
5+ years of experience as a DevOps or SRE
Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible
Experience with sustainable incident response in a blameless environment
Knowledge of cloud platforms (prefer Azure) and container + orchestration technologies
Experience with APM and Observability and related tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.
Background in Linux Systems Engineering
Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc.
Comfortable with a high level of autonomy and working with a distributed team

Nice-to-haves

Knowledge of Azure
Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
A love for high quality and a knack for testing
Opinions about dashboards, metrics, and SLO's

Benefits

Health insurance
Dental insurance
Vision insurance
Short-term disability insurance
Life insurance
Paid holidays
Paid time off
Fertility treatment benefit
401(k) plan
Equity
Eligibility for a discretionary company-wide bonus

Senior Site Reliability Engineer

About the position

Responsibilities

Requirements

Nice-to-haves

Benefits

Tools

Career Hubs

Guides

Company