Cribl - Little Rock, AR

posted 3 months ago

Full-time - Senior
Remote - Little Rock, AR

About the position

Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission to unlock the value of all observability data. As the data engine for IT and Security, Cribl provides users with a new level of observability, intelligence, and control over their real-time data. This role is remote, allowing you to work from anywhere while being part of a collaborative engineering organization. You will contribute to envisioning, creating, deploying, testing, and shipping Cribl products, all while enjoying a fun and engaging work culture that includes sharing goat gifs and fostering a positive team environment. In this position, you will be part of a team of technical engineers committed to delivering high-quality software. You will engage with various teams to improve service delivery and reliability across their entire lifecycle. Your responsibilities will include measuring and monitoring production systems with a focus on availability, latency, and overall system health. You will also seek out the causes of errors and instability in our production cloud services, driving teams towards better operational excellence. As a Senior Site Reliability Engineer, you will help identify and reduce toil through creative innovation and automation. You will have on-call responsibilities and will be expected to engage with product and platform teams to lobby for changes that enhance reliability, resilience, and observability. This role is ideal for someone who is passionate about reliability, has strong opinions on improving systems, and desires to build consensus around ideas that drive operational excellence.

Responsibilities

  • Engage with teams and improve service delivery and reliability across their entire lifecycle
  • Measure and monitor all production systems with an eye towards availability, latency and overall system health
  • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
  • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
  • Help identify and drive down toil with creative innovation and automation
  • Participate in on-call responsibilities

Requirements

  • Extensive experience with enterprise scale continuous delivery environments
  • 5+ years of experience with a DevOps or SRE job title
  • Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
  • Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible
  • Experience with sustainable incident response in a blameless environment
  • Knowledge of cloud platforms (prefer AWS) and container + orchestration technologies
  • Experience with APM and Observability and related tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc.
  • Background in Linux Systems Engineering
  • Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc.
  • Comfortable with a high level of autonomy and working with a distributed team

Nice-to-haves

  • Knowledge of Cloud and application security
  • Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
  • A love for high quality and a knack for testing
  • Opinions about dashboards, metrics, and SLO's

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Short-term disability insurance
  • Life insurance
  • Paid holidays
  • Paid time off
  • Fertility treatment benefit
  • 401(k) plan
  • Equity
  • Eligibility for a discretionary company-wide bonus
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service