Home/SRE & Ops
Topic

SRE & Ops

446 articles on SRE & Ops.

11,834 articles
SRE & Ops — Introducing Zelos: A ZooKeeper API leveraging Delos

Introducing Zelos: A ZooKeeper API leveraging Delos

Within large-scale services, durable storage, distributed leases, and coordination primitives such as distributed locks, semaphores, and events should be strongly consistent. At Meta, we have historically used Apache ZooKeeper as a centralized service for these primitives. However, as Meta’s workload has scaled, we’ve found ourselves pushing the limits of ZooKeeper’s capabilities. Modifying and tu

CWChris Wiltz·June 8, 2022SRE & Ops
SRE & Ops — Cache made consistent

Cache made consistent

Caches help reduce latency, scale read-heavy workloads, and save cost. They are literally everywhere. Caches run on your phone and in your browser. For example, CDNs and DNS are essentially geo-replicated caches. It’s thanks to many caches working behind the scenes that you can read this blog post right now. Phil Karlton famously said, “There […]

CWChris Wiltz·June 8, 2022SRE & Ops
SRE & Ops — GitHub Availability Report: May 2022

GitHub Availability Report: May 2022

In May, we experienced three distinct incidents resulting in significant impact to multiple services across GitHub.com. This report also sheds light into the billing incident that impacted Actions and Codespaces users in April.

JOJakub OleksyJakub Oleksy·June 1, 2022SRE & Ops
SRE & Ops — Logs on R2: slash your logging costs

Logs on R2: slash your logging costs

You shouldn’t have to make trade-offs between keeping logs that you need and managing tight budgets. R2’s low costs makes this decision easier and now you can use Logpush to store logs on R2.

TTanushreeTanushree·May 11, 2022SRE & Ops
SRE & Ops — Meta Open Source is transferring Jest to the OpenJS Foundation

Meta Open Source is transferring Jest to the OpenJS Foundation

Meta Open Source is officially transferring Jest, its open source JavaScript testing framework, to the OpenJS Foundation. With over 17 million weekly downloads and over 38,000 GitHub stars, Jest is the most used testing framework in the JavaScript ecosystem and is used by companies of all sizes, including Amazon, Google, Microsoft, and Stripe. We believe […]

CWChris Wiltz·May 11, 2022SRE & Ops
SRE & Ops — BellJar: A new framework for testing system recoverability at scale

BellJar: A new framework for testing system recoverability at scale

Building infrastructure that can easily recover from outages, particularly outages involving adjacent infrastructure, too often becomes a murky exploration of nuanced fate-sharing between systems. Untangling dependencies and uncovering side effects of unavailability has historically been time-consuming work. A lack of great tooling built for this, and the rarity of infrastructure outages, makes re

CWChris Wiltz·May 5, 2022SRE & Ops
SRE & Ops — How the Cinder JIT’s function inliner helps us optimize Instagram

How the Cinder JIT’s function inliner helps us optimize Instagram

Since Instagram runs one of the world’s largest deployments of the Django web framework, we have natural interest in finding ways to optimize Python so we can speed up our production application. As part of this effort, we’ve recently open-sourced Cinder, our Python runtime that is a fork of CPython. Cinder includes optimizations like immortal […]

CWChris Wiltz·May 2, 2022SRE & Ops
SRE & Ops — Slack’s Incident on 2-22-22

Slack’s Incident on 2-22-22

By Laura Nolan, with contributions from Glen D. Sanford, Jamie Scheinblum, and Chris Sullivan. Assessing conditions Slack experienced a major incident on February 22 this year, during which time many users were unable to connect to Slack, including the author — which certainly made my role as Incident Commander more challenging! This incident was a…

LNLaura Nolan·April 26, 2022SRE & Ops
SRE & Ops — How Meta enables de-identified authentication at scale

How Meta enables de-identified authentication at scale

Data minimization — collecting the minimum amount of data required to support our services — is one of our core principles at Meta as we continue developing new privacy-enhancing technologies (PETs). We are constantly seeking ways to improve privacy and protect user data on our family of products. Previously, we’ve approached data minimization by exploring […]

CWChris Wiltz·March 30, 2022SRE & Ops
SRE & Ops — Detecting silent errors in the wild: Combining two novel approaches to quickly detect silent data corruptions at scale

Detecting silent errors in the wild: Combining two novel approaches to quickly detect silent data corruptions at scale

Silent data corruptions (SDCs), data errors that go undetected by the larger system, are a widespread problem for large-scale infrastructure systems. Left undetected, these types of corruptions can cause data loss and propagate across the stack and manifest as application-level problems. Silent data corruptions (SDC) in hardware impact computational integrity for large-scale applications. Sources

CWChris Wiltz·March 17, 2022SRE & Ops
SRE & Ops — Using Terraform to Manage Infrastructure - Shopify

Using Terraform to Manage Infrastructure - Shopify

Large applications are often a mix of code your team has written and third-party applications your team needs to manage. These third-party applications could be things like AWS or Docker. In my team’s case, it’s Twilio TaskRouter. The configuration of these services may not change as often as your app code does, but when it does, the process is fraught with the potential for errors. This is becaus

SEShopify Engineering·March 17, 2022SRE & Ops
SRE & Ops — VESPA: Static profiling for binary optimization

VESPA: Static profiling for binary optimization

What the research is: Recent research has demonstrated that binary optimization is important for achieving peak performance for various applications. For instance, the state-of-the-art BOLT binary optimizer developed at Meta, which is part of the LLVM Compiler Project, significantly improves the performance of highly optimized binaries produced using compilers’ most aggressive optimizations, such

CWChris Wiltz·March 15, 2022SRE & Ops
SRE & Ops — Get full observability into your Cloudflare logs with New Relic

Get full observability into your Cloudflare logs with New Relic

Correlating Cloudflare logs across your stack in New Relic One is powerful for monitoring and debugging in order to keep services safe and reliable. We’re excited to have partnered with New Relic to create a direct integration that provides this visibility

TMTanushree, Mike Neville O NeillTanushree, Mike Neville O Neill·March 14, 2022SRE & Ops
SRE & Ops — Improving The CI/CD Flow For Your Application — Smashing Magazine

Improving The CI/CD Flow For Your Application — Smashing Magazine

Looking for ways to create a smooth CI/CD flow for your software? In this article, Tom Hastjarjanto shares a quick list of useful concepts that can be combined with GitHub Actions and NPM packages. To fully benefit from the setup and the release with maximum confidence, it is highly recommended to have a robust test suite that runs on integration.

THTom HastjarjantoTom Hastjarjanto·March 14, 2022SRE & Ops
SRE & Ops — An Introduction To AWS Cloud Development Kit (CDK) — Smashing Magazine

An Introduction To AWS Cloud Development Kit (CDK) — Smashing Magazine

In this article, Vivek Maskara introduces Amazon Web Services’ (AWS) Cloud Development Kit (CDK) which is increasingly becoming a popular tool for managing AWS-based infrastructure. We’ll take a closer look into CDK concepts, and then how to use the AWS CDK toolkit to deploy a sample application to an AWS account.

VMVivek MaskaraVivek Maskara·March 10, 2022SRE & Ops
SRE & Ops — Augmenting Flexible Paxos in LogDevice to improve read availability

Augmenting Flexible Paxos in LogDevice to improve read availability

We’ve improved read availability in LogDevice, Meta’s scalable distributed log storage system, by removing a fundamental trade-off in Flexible Paxos, the algorithm used to gain consensus among our distributed systems. At Meta’s scale, systems need to be reliable, even in the face of organic failures like power loss events, or when systems are undergoing hardware […]

CWChris Wiltz·March 7, 2022SRE & Ops
SRE & Ops — Incident Report: Spotify Outage on March 8, 2022

Incident Report: Spotify Outage on March 8, 2022

On March 8, we experienced a global outage triggered by issues in a cloud-hosted service discovery system used at Spotify. We were made aware of issues with login at 18:12 UTC / 13:12 ET and started implementing fixes to critical systems at 18:39 UTC / 13:39 ET. This outage affected our users and we apologize for the inconvenience it may have caused. Our service has now fully recovered.

SESpotify Engineering·March 1, 2022SRE & Ops
SRE & Ops — Internet is back in Tonga after 38 days of outage

Internet is back in Tonga after 38 days of outage

Tonga, the South Pacific archipelago nation (with 169 islands), was reconnected to the Internet this early morning (UTC) and is back online after successful repairs to the undersea cable that was damaged on Saturday, January 15, 2022, by the January 14, volcanic eruption

JTJoao TomeJoao Tome·February 22, 2022SRE & Ops
SRE & Ops — Balancing Safety and Velocity in CI/CD at Slack

Balancing Safety and Velocity in CI/CD at Slack

In 2021, we changed developer testing workflows for Webapp, Slack’s main monorepo, from predominantly testing before merging to a multi-tiered testing workflow after merging. This changed our previous definition of safety and developer workflows between testing and deploys. In this project, we aimed to ensure frequent, reliable, and high-quality releases to our customers for a…

CVCarlos Valdez·February 18, 2022SRE & Ops
SRE & Ops — Tonga’s likely lengthy Internet outage

Tonga’s likely lengthy Internet outage

The latest Internet outage, in the South Pacific country of Tonga (with 169 islands), is still ongoing. It started with the large eruption of Hunga Tonga–Hunga Haʻapai, an uninhabited volcanic island of the Tongan archipelago on Friday, January 14, 2022

JTJoao TomeJoao Tome·January 19, 2022SRE & Ops
SRE & Ops — FOQS: Making a distributed priority queue disaster-ready

FOQS: Making a distributed priority queue disaster-ready

Facebook Ordered Queueing Service (FOQS) is a fully managed, distributed priority queueing service used for reliable message delivery among many services. FOQS has evolved from a regional deployment into a geo-distributed, global deployment to help ensure that data stored within logical queues is highly available, even through large-scale disaster scenarios. Migrating to a global architecture

MEMeta Engineering·January 18, 2022SRE & Ops
SRE & Ops — That Old Certificate Expired and Started an Outage. This is What Happened Next - Shopify

That Old Certificate Expired and Started an Outage. This is What Happened Next - Shopify

In distributed systems, there’s plenty of occasions for things to go wrong. This is why resiliency and redundancy are important. But no matter the systems you put in place, no matter whether you did or didn’t touch your deployments, issues might arise. It makes it critical to acknowledge the near misses: the situations where something could have gone wrong and the situations where something did, b

SEShopify Engineering·January 12, 2022SRE & Ops
SRE & Ops — Power Loss Siren: Making Meta resilient to power loss events

Power Loss Siren: Making Meta resilient to power loss events

There are thousands of distributed services running on millions of servers in Meta’s data centers. Part of ensuring the reliability of those services means making them resilient to power loss events as our data center fleet grows. To help increase resiliency, we built the Power Loss Siren (PLS) — a rack level, low latency, distributed […]

CWChris Wiltz·December 16, 2021SRE & Ops
SRE & Ops — A Simple Kubernetes Admission Webhook

A Simple Kubernetes Admission Webhook

While adding a recent feature to our Kubernetes compute platform, we had the need to mutate newly-created pods based on annotations set by users. The mutation needed to follow simple business rules, and didn’t need to keep track of any state. Surely there must be a canonical solution to this simple problem? Well, sort of.…

CLClément Labbe·December 14, 2021SRE & Ops
SRE & Ops — SLICK: Adopting SLOs for improved reliability

SLICK: Adopting SLOs for improved reliability

We would like to thank Peter Tang for all his work on SLICK, and for helping us write this post! To support the people and communities who use our apps and products, we need to stay in constant contact with them. We want to provide the experiences we offer reliably. We also need to establish […]

MEMeta Engineering·December 13, 2021SRE & Ops
SRE & Ops — GraphQL global ID migration update

GraphQL global ID migration update

All newly created GraphQL objects now have IDs that conform to a new format, which we refer to as “next IDs.” Learn how to migrate older IDs to the new format and why we’re making the change.

AHAndrew HoglundAndrew Hoglund·November 16, 2021SRE & Ops
SRE & Ops — OCP Summit 2021: Open networking hardware lays the groundwork for the metaverse

OCP Summit 2021: Open networking hardware lays the groundwork for the metaverse

Open infrastructure technologies and networking hardware will play an important role as we build new technologies for the metaverse, where billions of people will someday come together in virtual spaces. As we head toward the next major computing platform with a continued spirit of embracing openness and disaggregation, we’re announcing two new milestones for our […]

CWChris Wiltz·November 9, 2021SRE & Ops
SRE & Ops — Kangaroo: A new flash cache optimized for tiny objects

Kangaroo: A new flash cache optimized for tiny objects

What the research is: Kangaroo is a new flash cache that enables more efficient caching of tiny objects (objects that are ~100 bytes or less) and overcomes the challenges presented by existing flash cache designs. Since Kangaroo is implemented within CacheLib, Facebook’s open source caching engine, developers can use Kangaroo through CacheLib’s API to build […]

CWChris Wiltz·October 26, 2021SRE & Ops
SRE & Ops — Autonomous testing of services at scale

Autonomous testing of services at scale

Enabling developers to prototype, test, and iterate on new features quickly is important to Facebook’s success. To do this effectively, it’s key to have a stable infrastructure that doesn’t introduce unnecessary friction. This gets significantly more challenging when the infrastructure in question must also scale to support more than 3 billion people around the world, […]

MEMeta Engineering·October 20, 2021SRE & Ops
SRE & Ops — Tunnel: Cloudflare’s Newest Homeowner

Tunnel: Cloudflare’s Newest Homeowner

Starting today, users who deploy and manage Cloudflare Tunnel at scale now have easier visibility into their Tunnel’s respective status, routes, uptime, connectors, cloudflared version, and much more through our new UI in the Cloudflare for Teams Dashboard.

AAbeAbe·October 18, 2021SRE & Ops
SRE & Ops — Infrastructure Observability for Changing the Spend Curve

Infrastructure Observability for Changing the Spend Curve

Slack is an integral part of where work happens for teams across the world, and our work in the Core Development Engineering department supports engineers throughout Slack that develop, build, test, and release high-quality services to Slack’s customers. In this article, we share how teams at Slack evolved our internal tooling and made infrastructure bets.…

FCFrank Chen·October 7, 2021SRE & Ops
SRE & Ops — Changing the Wheels on a Moving Bus — Spotify’s Event Delivery Migration

Changing the Wheels on a Moving Bus — Spotify’s Event Delivery Migration

At Spotify, data rules all. We log a variety of data, from listening history, to results of A/B testing, to page load times so we can analyze and improve the Spotify service. We instrument and log data across every surface that is running Spotify code through a system called the Event Delivery Infrastructure (EDI). Throughout this blog post we make a distinction between the internal users of the E

FSFlavio Santos and Robert Stephenson·October 1, 2021SRE & Ops