Cribl - Oklahoma City, OK

posted 3 months ago

Full-time - Senior
Remote - Oklahoma City, OK

About the position

Cribl is seeking a Staff Site Reliability Engineer (SRE) to join our mission of unlocking the value of all observability data. As a remote-first company, we empower our employees to do their best work from anywhere. In this role, you will be part of a team dedicated to delivering high-quality software while enjoying a collaborative and fun work environment. You will contribute to the design, development, deployment, testing, and shipping of Cribl products, which are essential for providing users with enhanced observability, intelligence, and control over their real-time data. As a Staff SRE, you will engage with various teams to improve service delivery and reliability throughout the entire lifecycle of our systems. Your responsibilities will include measuring and monitoring production systems for availability, latency, and overall health, as well as identifying and addressing the root causes of errors and instability in our cloud services. You will work closely with product and platform teams to advocate for changes that enhance reliability, resilience, and observability. Additionally, you will focus on reducing operational toil through innovative solutions and automation, and you will have on-call responsibilities to ensure system reliability. This position is ideal for individuals who are passionate about reliability and have strong opinions on improving systems. You will have the opportunity to be involved in all aspects of cloud operations, scaling, and high availability, making a significant impact on the technology we are building at Cribl.

Responsibilities

  • Engage with teams to improve service delivery and reliability across their entire lifecycle.
  • Measure and monitor all production systems with an eye towards availability, latency, and overall system health.
  • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence.
  • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability.
  • Help identify and drive down toil with creative innovation and automation.
  • Participate in on-call responsibilities.

Requirements

  • Extensive experience with enterprise scale continuous delivery environments.
  • 8+ years of experience with a DevOps or SRE job title.
  • Development experience with JavaScript/Node.js/TypeScript in a Linux/Mac environment.
  • Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible.
  • Experience with sustainable incident response in a blameless environment.
  • Knowledge of cloud platforms (preferably AWS) and container + orchestration technologies.
  • Experience with APM and Observability tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry, etc.
  • Background in Linux Systems Engineering.
  • Experience with incident response tools such as PagerDuty, FireHydrant, Blameless, etc.
  • Comfortable with a high level of autonomy and working with a distributed team.

Nice-to-haves

  • Knowledge of Cloud and application security.
  • Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
  • A love for high quality and a knack for testing.
  • Opinions about dashboards, metrics, and SLOs.

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Short-term disability insurance
  • Life insurance
  • Paid holidays
  • Paid time off
  • Fertility treatment benefit
  • 401(k) plan
  • Equity options
  • Eligibility for a discretionary company-wide bonus
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service