Staff Site Reliability Engineer (SRE)

$144,000 - $278,000/Yr

Cribl - Tallahassee, FL

posted 3 months ago

Full-time - Senior

Remote - Tallahassee, FL

About the position

Cribl Inc is seeking a Staff Site Reliability Engineer to join our mission to unlock the value of all observability data. As a remote-first company, we empower our employees to do their best work, wherever they are. In this role, you will be part of a team of technical engineers committed to shipping only high-quality software while enjoying a fun and collaborative work environment. You will contribute to envisioning, creating, deploying, testing, and shipping Cribl products, which are designed to provide users with a new level of observability, intelligence, and control over their real-time data. This position offers a unique opportunity to be part of a team that is fundamentally changing technology and making a real impact in the industry. As a Staff Site Reliability Engineer, you will engage with various teams to improve service delivery and reliability across their entire lifecycle. You will measure and monitor all production systems with a focus on availability, latency, and overall system health. Your role will involve seeking out the causes of errors and instability in our production cloud services and driving teams towards better operational excellence. You will also engage with product and platform teams to improve and evolve systems by advocating for changes that enhance reliability, resilience, and observability. Additionally, you will help identify and reduce toil through creative innovation and automation, and you will have on-call responsibilities as part of your role.

Responsibilities

Engage with teams and improve service delivery and reliability across their entire lifecycle.
Measure and monitor all production systems with an eye towards availability, latency, and overall system health.
Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence.
Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability.
Help identify and drive down toil with creative innovation and automation.
Participate in on-call responsibilities.

Requirements

Extensive experience with enterprise scale continuous delivery environments.
8+ years of experience with a DevOps or SRE job title.
Development experience with JavaScript/Node.js/TypeScript in a Linux/Mac environment.
Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible.
Experience with sustainable incident response in a blameless environment.
Knowledge of cloud platforms (prefer AWS) and container + orchestration technologies.
Experience with APM and Observability tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry, etc.
Background in Linux Systems Engineering.
Experience with incident response related tools such as PagerDuty, FireHydrant, Blameless, etc.
Comfortable with a high level of autonomy and working with a distributed team.

Nice-to-haves

Knowledge of Cloud and application security.
Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
A love for high quality and a knack for testing.
Opinions about dashboards, metrics, and SLO's.

Benefits

Health insurance
Dental insurance
Vision insurance
Short-term disability insurance
Life insurance
Paid holidays
Paid time off
Fertility treatment benefit
401(k)
Equity
Eligibility for a discretionary company-wide bonus

Staff Site Reliability Engineer (SRE)

About the position

Responsibilities

Requirements

Nice-to-haves

Benefits

Tools

Career Hubs

Guides

Company