Staff Site Reliability Engineer (SRE)

Cribl - Columbus, OH

posted 3 months ago

Full-time - Senior

Remote - Columbus, OH

About the position

Cribl Inc is seeking a Staff Site Reliability Engineer to join our mission to unlock the value of all observability data. As the data engine for IT and Security, Cribl provides users with a new level of observability, intelligence, and control over their real-time data. In this role, you will be part of a remote-first engineering organization, contributing to the envisioning, creation, deployment, testing, and shipping of Cribl products. You will join a team of technical engineers committed to delivering high-quality software while enjoying a fun and collaborative work environment. This position offers a unique opportunity to be part of a company that is fundamentally changing technology and putting customers in full control of their observability data. If you are passionate about reliability and have strong opinions on improving systems, this role may be the perfect fit for you.

Responsibilities

Engage with teams and improve service delivery and reliability across their entire lifecycle.
Measure and monitor all production systems with an eye towards availability, latency, and overall system health.
Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence.
Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability.
Help identify and drive down toil with creative innovation and automation.
Participate in on-call responsibilities.

Requirements

Extensive experience with enterprise scale continuous delivery environments.
8+ years of experience with a DevOps or SRE job title.
Development experience with JavaScript/Node.js/TypeScript in a Linux/Mac environment.
Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible.
Experience with sustainable incident response in a blameless environment.
Knowledge of cloud platforms (preferably AWS) and container + orchestration technologies.
Experience with APM and Observability tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry, etc.
Background in Linux Systems Engineering.
Experience with incident response related tools such as PagerDuty, FireHydrant.

Staff Site Reliability Engineer (SRE)

About the position

Responsibilities

Requirements

Tools

Career Hubs

Guides

Company