Cribl - Saint Paul, MN
posted 2 months ago
Cribl Inc is seeking a Staff Site Reliability Engineer to join our mission to unlock the value of all observability data. As a remote-first company, we empower our employees to do their best work, wherever they are. In this role, you will be part of a team of technical engineers committed to shipping only high-quality software while enjoying a fun and collaborative work environment. You will contribute to envisioning, creating, deploying, testing, and shipping Cribl products, which are designed to provide users with a new level of observability, intelligence, and control over their real-time data. This position is not just about fixing things on the operational side; our SRE engineers are involved from conception to design to development and all the way through production and beyond. You will have the opportunity to provide your creative input into all things Cloud, Scaling, Reliability, High Availability, and much more. If reliability is your passion and you have strong opinions on how to improve systems, this role is for you. As an active member of our team, you will engage with various teams to improve service delivery and reliability across their entire lifecycle. You will measure and monitor all production systems with a focus on availability, latency, and overall system health. Your role will involve seeking out the causes of errors and instability in our production cloud services and driving teams towards better operational excellence. You will also engage with product and platform teams to improve and evolve systems by advocating for changes that enhance reliability, resilience, and observability. Additionally, you will help identify and reduce toil through creative innovation and automation, and you will have on-call responsibilities as part of your role.