Cribl - Albany, NY
posted 2 months ago
Cribl is on a mission to unlock the value of all observability data, and we are looking for a Staff Site Reliability Engineer (SRE) to join our team. As a remote-first company, we empower our employees to do their best work from anywhere. In this role, you will be part of a dynamic engineering organization that is committed to delivering high-quality software while enjoying a fun and collaborative work environment. You will have the opportunity to work with some of the biggest names in the industry, helping them solve their most pressing data needs. As a Staff Site Reliability Engineer, you will engage with various teams to improve service delivery and reliability throughout the entire lifecycle of our products. Your responsibilities will include measuring and monitoring production systems to ensure availability, latency, and overall system health. You will actively seek out the causes of errors and instability in our production cloud services, driving teams towards operational excellence. Your role will also involve collaborating with product and platform teams to advocate for changes that enhance reliability, resilience, and observability. In addition to your technical expertise, you will be expected to identify and reduce toil through creative innovation and automation. On-call responsibilities will be part of your role, ensuring that you are actively involved in maintaining the reliability of our systems. If you are passionate about reliability and have strong opinions on how to improve processes, we want to hear from you!