Staff Site Reliability Engineer (SRE) - 5168362004_11-6854

$144,000 - $144,000/Yr

Cribl - Springfield, IL

posted 3 months ago

Full-time - Senior

Remote - Springfield, IL

About the position

Cribl Inc is seeking a Staff Site Reliability Engineer (SRE) to join our mission to unlock the value of all observability data. As the data engine for IT and Security, many of the biggest names in the most demanding industries trust Cribl to solve their most pressing data needs. This role is remote, allowing you to work from anywhere while being part of a collaborative and motivated team that is passionate about putting customers first. You will be part of the engineering organization, contributing to envisioning, creating, deploying, testing, and shipping Cribl products. In this position, you will engage with teams to improve service delivery and reliability across their entire lifecycle. You will measure and monitor all production systems with a focus on availability, latency, and overall system health. Your role will involve seeking out the causes of errors and instability in our production cloud services and driving teams towards better operational excellence. You will also engage with product and platform teams to improve and evolve systems by advocating for changes that enhance reliability, resilience, and observability. Additionally, you will help identify and reduce toil through creative innovation and automation, and you will have on-call responsibilities as part of the role. This is an exciting opportunity to be part of something that is fundamentally changing technology. If you are passionate about reliability and have strong opinions on how to improve systems, we want to hear from you!

Responsibilities

Engage with teams and improve service delivery and reliability across their entire lifecycle
Measure and monitor all production systems with an eye towards availability, latency and overall system health
Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
Help identify and drive down toil with creative innovation and automation
On-call responsibilities

Requirements

Extensive experience with enterprise scale continuous delivery environments
8+ years of experience with a DevOps or SRE job title
Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible
Experience with sustainable incident response in a blameless environment
Knowledge of cloud platforms (prefer AWS) and container + orchestration technologies
Experience with APM and Observability and related tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.
Background in Linux Systems Engineering
Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc.
Comfortable with a high level of autonomy and working with a distributed team

Nice-to-haves

Knowledge of Cloud and application security
Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
A love for high quality and a knack for testing
Opinions about dashboards, metrics, and SLO's

Benefits

Health insurance
Dental insurance
Vision insurance
Short-term disability insurance
Life insurance
Paid holidays
Paid time off
Fertility treatment benefit
401(k) plan
Equity
Eligibility for a discretionary company-wide bonus

Staff Site Reliability Engineer (SRE) - 5168362004_11-6854

About the position

Responsibilities

Requirements

Nice-to-haves

Benefits

Tools

Career Hubs

Guides

Company