New Relic, Inc.
Observability at New Relic Engineering Playbook
Pages
20
Time to read
24 mins
Publication
Language
English
Pages
20
Time to read
24 mins
Publication
Language
English
This white paper outlines New Relic's engineering playbook focused on achieving hyperscale and reliability through observability. It details how various teams within New Relic utilize their own observability platform to meet critical business objectives, such as enhancing developer productivity, achieving cloud cost savings, and maintaining high uptime. The document presents concrete examples of how internal teams, including Site Reliability Engineering (SRE) and platform engineering, leverage observability to ensure operational excellence. Key strategies discussed include measuring important metrics for system performance, implementing self-healing systems to reduce engineering toil, and utilizing automation for rapid issue mitigation. The paper emphasizes the importance of Service Level Objectives (SLOs) in balancing reliability and new functionality, and it describes the role of alerting and auto-scaling in maintaining system health. Through these practices, New Relic aims to continuously refine its platform while delivering reliable experiences for its customers.