Home/SRE & Ops
Topic

SRE & Ops

446 articles on SRE & Ops.

11,834 articles
SRE & Ops — Technology Lifecycle

Technology Lifecycle

This blog post discusses the strategies that Slack uses to manage the lifecycle (development, support, and eventual retirement) of infrastructure projects, through the lens of the migration through three successive internal “platform” offerings. Our challenges Circa 2020, our Cloud Engineering team (now evolved into multiple teams responsible for narrower aspects) was responsible for managing our…

TBTricia Bogen·March 21, 2023SRE & Ops
SRE & Ops — Introducing Velox: An open source unified execution engine

Introducing Velox: An open source unified execution engine

Meta is introducing Velox, an open source unified execution engine aimed at accelerating data management systems and streamlining their development. Velox is under active development. Experimental results from our paper published at the International Conference on Very Large Data Bases (VLDB) 2022 show how Velox improves efficiency and consistency in data management systems. Velox helps […] Read M

CWChris Wiltz·March 9, 2023SRE & Ops
SRE & Ops — How Cloudflare runs Prometheus at scale

How Cloudflare runs Prometheus at scale

Here at Cloudflare we run over 900 instances of Prometheus with a total of around 4.9 billion time series. Operating such a large Prometheus deployment doesn’t come without challenges . In this blog post we’ll cover some of the issues we hit and how we solved them

LLukaszLukasz·March 3, 2023SRE & Ops
SRE & Ops — GitHub Actions Importer is now generally available

GitHub Actions Importer is now generally available

We’re excited to announce the general availability of GitHub Actions Importer. GitHub Actions Importer helps you plan, forecast, and automate migrations from Azure DevOps, CircleCI, GitLab, Jenkins, and Travis CI…

DGDawit GebregziabherDawit Gebregziabher·March 1, 2023SRE & Ops
SRE & Ops — Building a cross-platform runtime for AR

Building a cross-platform runtime for AR

Meta’s augmented reality (AR) platform is one of the largest in the world, helping the billions of people on Meta’s apps experience AR every day and giving hundreds of thousands of creators a means to express themselves Meta’s AR tools are unique because they can be used on a wide variety of devices — from […]

MEMeta Engineering·February 13, 2023SRE & Ops
SRE & Ops — The evolution of Facebook’s iOS app architecture

The evolution of Facebook’s iOS app architecture

Facebook for iOS (FBiOS) is the oldest mobile codebase at Meta. Since the app was rewritten in 2012, it has been worked on by thousands of engineers and shipped to billions of users, and it can support hundreds of engineers iterating on it at a time. After years of iteration, the Facebook codebase does not […]

CWChris Wiltz·February 6, 2023SRE & Ops
SRE & Ops — Enabling branch deployments through IssueOps with GitHub Actions

Enabling branch deployments through IssueOps with GitHub Actions

What if developers want to leverage branch deployments but don’t have a full ChatOps stack integrated with their repositories? We wanted to set out to find a way for all developers to be able to take advantage of branch deployments with ease, right from their GitHub repository, and so the branch-deploy Action was born!

GBGrant BirkinbineGrant Birkinbine·February 2, 2023SRE & Ops
SRE & Ops — Incident Report: Spotify Outage on January 14, 2023

Incident Report: Spotify Outage on January 14, 2023

On January 14, between 00:15 UTC and 03:45 UTC, Spotify suffered an outage. The impact was small at first and increased over the course of an hour until most functionality (including playback) was not working. Spotify engineers were first notified of the problem at 00:40 UTC, and our incident response team was immediately assembled. Due to the nature of the incident, triage took longer than one wo

SESpotify Engineering·February 1, 2023SRE & Ops
SRE & Ops — Asynchronous computing at Meta: Overview and learnings

Asynchronous computing at Meta: Overview and learnings

We’ve made architecture changes to Meta’s event driven asynchronous computing platform that have enabled easy integration with multiple event-sources. We’re sharing our learnings from handling various workloads and how to tackle trade offs made with certain design choices in building the platform. Asynchronous computing is a paradigm where the user does not expect a workload […]

CWChris Wiltz·January 31, 2023SRE & Ops
SRE & Ops — Tulip: Modernizing Meta’s data platform

Tulip: Modernizing Meta’s data platform

The technical journey discusses the motivations, challenges, and technical solutions employed for warehouse schematization, especially a change to the wire serialization format employed in Meta’s data platform for data interchange related to Warehouse Analytics Logging. Here, we discuss the engineering, scaling, and nontechnical challenges of modernizing Meta’s exabyte-scale data platform by migra

CWChris Wiltz·January 26, 2023SRE & Ops
SRE & Ops — Cloudflare incident on January 24, 2023

Cloudflare incident on January 24, 2023

Several Cloudflare services became unavailable for 121 minutes on January 24th, 2023 due to an error releasing code that manages service tokens. The incident degraded a wide range of Cloudflare products

KSKenny, SamKenny, Sam·January 25, 2023SRE & Ops
SRE & Ops — Open-sourcing Anonymous Credential Service

Open-sourcing Anonymous Credential Service

Meta has open-sourced Anonymous Credential Service (ACS), a highly available multitenant service that allows clients to authenticate in a de-identified manner. ACS enhances privacy and security while also being compute-conscious. By open-sourcing and fostering a community for ACS, we believe we can accelerate the pace of innovation in de-identified authentication. Data minimization — collecting th

CWChris Wiltz·December 12, 2022SRE & Ops
SRE & Ops — Enabling static analysis of SQL queries at Meta

Enabling static analysis of SQL queries at Meta

UPM is our internal standalone library to perform static analysis of SQL code and enhance SQL authoring. UPM takes SQL code as input and represents it as a data structure called a semantic tree. Infrastructure teams at Meta leverage UPM to build SQL linters, catch user mistakes in SQL code, and perform data lineage analysis […]

MEMeta Engineering·November 30, 2022SRE & Ops
SRE & Ops — Retrofitting null-safety onto Java at Meta

Retrofitting null-safety onto Java at Meta

We developed a new static analysis tool called Nullsafe that is used at Meta to detect NullPointerException (NPE) errors in Java code. Interoperability with legacy code and gradual deployment model were key to Nullsafe’s wide adoption and allowed us to recover some null-safety properties in the context of an otherwise null-unsafe language in a multimillion-line […]

CWChris Wiltz·November 22, 2022SRE & Ops
SRE & Ops — Sapling: Source control that’s user-friendly and scalable

Sapling: Source control that’s user-friendly and scalable

Sapling is a new Git-compatible source control client. Sapling emphasizes usability while also scaling to the largest repositories in the world. ReviewStack is a demonstration code review UI for GitHub pull requests that integrates with Sapling to make reviewing stacks of commits easy. You can get started using Sapling today. Source control is one of […]

CWChris Wiltz·November 15, 2022SRE & Ops
SRE & Ops — Tulip: Schematizing Meta’s data platform

Tulip: Schematizing Meta’s data platform

We’re sharing Tulip, a binary serialization protocol supporting schema evolution. Tulip assists with data schematization by addressing protocol reliability and other issues simultaneously. It replaces multiple legacy formats used in Meta’s data platform and has achieved significant performance and efficiency gains. There are numerous heterogeneous services, such as warehouse data storage and vario

CWChris Wiltz·November 9, 2022SRE & Ops
SRE & Ops — Improving Instagram notification management with machine learning and causal inference

Improving Instagram notification management with machine learning and causal inference

We’re sharing how Meta is applying statistics and machine learning (ML) to improve notification personalization and management on Instagram – particularly on daily digest push notifications. By using causal inference and ML to identify highly active users who are likely to see more content organically, we have been able to reduce the number of notifications […]

CWChris Wiltz·October 31, 2022SRE & Ops
SRE & Ops — How We Use Terraform At Slack

How We Use Terraform At Slack

At Slack, we use Terraform for managing our Infrastructure, which runs on AWS, DigitalOcean, NS1, and GCP. Even though most of our infrastructure is running on AWS, we have chosen to use Terraform as opposed to using an AWS-native service such as CloudFormation so that we can use a single tool across all of our…

AGArchie Gunasekara·October 25, 2022SRE & Ops
SRE & Ops — OCP Summit 2022: Open hardware for AI infrastructure

OCP Summit 2022: Open hardware for AI infrastructure

At OCP Summit 2022, we’re announcing Grand Teton, our next-generation platform for AI at scale that we’ll contribute to the OCP community. We’re also sharing new innovations designed to support data centers as they advance to support new AI technologies: A new, more efficient version of Open Rack. Our Air-Assisted Liquid Cooling (AALC) – design. […]

CWChris Wiltz·October 18, 2022SRE & Ops
SRE & Ops — Scaling data ingestion for machine learning training at Meta

Scaling data ingestion for machine learning training at Meta

Many of Meta’s products, such as search, ads ranking and Marketplace, utilize AI models to continuously improve user experiences. As the performance of hardware we use to support training infrastructure increases, we need to scale our data ingestion infrastructure accordingly to handle workloads more efficiently. GPUs, which are used for training infrastructure, tend to double […]

MEMeta Engineering·September 19, 2022SRE & Ops
SRE & Ops — How thermal simulation helps optimize Meta’s data centers

How thermal simulation helps optimize Meta’s data centers

Data center optimization has always played an important role at Meta. By optimizing our data centers’ environmental controls, we can reduce our environmental impact while ensuring that people can always depend on our products. With most other complex systems, optimization of energy consumption is a trial-and-error process. But experimenting on any component of a live […]

MEMeta Engineering·September 14, 2022SRE & Ops
SRE & Ops — MemLab: An open source framework for finding JavaScript memory leaks

MemLab: An open source framework for finding JavaScript memory leaks

We’ve open-sourced MemLab, a JavaScript memory testing framework that automates memory leak detection. Finding and addressing the root cause of memory leaks is important for delivering a quality user experience on web applications. MemLab has helped engineers and developers at Meta improve user experience and make significant improvements in memory optimization. We hope it will […]

CWChris Wiltz·September 12, 2022SRE & Ops
SRE & Ops — GitHub Availability Report: August 2022

GitHub Availability Report: August 2022

In August, we experienced one incident resulting in significant impact to Codespaces. We’re still investigating that incident and will include it in next month’s report. This report also sheds light into an incident that impacted Codespaces in July.

JOJakub OleksyJakub Oleksy·September 7, 2022SRE & Ops
SRE & Ops — Open-sourcing TAOBench: An end-to-end social network benchmark

Open-sourcing TAOBench: An end-to-end social network benchmark

What the research is: The continued emergence of large social network applications has introduced a scale of data and query volume that challenges the limits of existing data stores. However, few benchmarks accurately simulate these request patterns, leaving researchers in short supply of tools to evaluate and improve upon these systems. To address this issue, […]

CWChris Wiltz·September 7, 2022SRE & Ops
SRE & Ops — Viewing the world as a computer: Global capacity management

Viewing the world as a computer: Global capacity management

Meta currently operates 14 data centers around the world. This rapidly expanding global data center footprint poses new challenges for service owners and for our infrastructure management systems. Systems like Twine, which we use to scale cluster management, and RAS, which handles perpetual region-wide resource allocation, have provided the abstractions and automation necessary for service […] Rea

CWChris Wiltz·September 6, 2022SRE & Ops
SRE & Ops — Improving Meta’s SLO workflows with data annotations

Improving Meta’s SLO workflows with data annotations

When we focus on minimizing errors and downtime here at Meta, we place a lot of attention on service-level indicators (SLIs) and service-level objectives (SLOs). Consider Instagram, for example. There, SLIs represent metrics from different product surfaces, like the volume of error response codes to certain endpoints, or the number of successful media uploads. Based […]

CWChris Wiltz·August 29, 2022SRE & Ops
SRE & Ops — GitHub Pages now uses Actions by default

GitHub Pages now uses Actions by default

As GitHub Pages, home to 16 million websites, approaches its 15th anniversary, we’re excited to announce that all sites now build and deploy with GitHub Actions.

CPChris Patterson, Yoann ChaudetChris Patterson, Yoann Chaudet·August 10, 2022SRE & Ops
SRE & Ops — Programming languages endorsed for server-side use at Meta

Programming languages endorsed for server-side use at Meta

Supporting a programming language at Meta is a very careful and deliberate decision. We’re sharing our internal programming language guidance that helps our engineers and developers choose the best language for their projects. Rust is the latest addition to Meta’s list of supported server-side languages. At Meta, we use many different programming languages for a […]

CWChris Wiltz·July 27, 2022SRE & Ops
SRE & Ops — It’s time to leave the leap second in the past

It’s time to leave the leap second in the past

The leap second concept was first introduced in 1972 by the International Earth Rotation and Reference Systems Service (IERS) in an attempt to periodically update Coordinated Universal Time (UTC) due to imprecise observed solar time (UT1) and the long-term slowdown in the Earth’s rotation. This periodic adjustment mainly benefits scientists and astronomers as it allows […]

CWChris Wiltz·July 25, 2022SRE & Ops
SRE & Ops — Spin Infrastructure Adventures: Containers, Systemd, and CGroups - Shopify

Spin Infrastructure Adventures: Containers, Systemd, and CGroups - Shopify

The Spin infrastructure team works hard at improving the stability of the system. In February 2022 we moved to Container Optimized OS (COS), the Google maintained operating system for their Kubernetes Engine SaaS offering. A month later we turned on multi-cluster to allow for increased scalability as more users came on board. Recently, we’ve increased default resources allotted to instances dramat

SEShopify Engineering·July 15, 2022SRE & Ops
SRE & Ops — Owl: Distributing content at Meta scale

Owl: Distributing content at Meta scale

Being able to distribute large, widely -consumed objects (so-called hot content) efficiently to hosts is becoming increasingly important within Meta’s private cloud. These are commonly distributed content types such as executables, code artifacts, AI models, and search indexes that help enable our software systems. Owl is a new system for high-fanout distribution of large data […]

CWChris Wiltz·July 14, 2022SRE & Ops
SRE & Ops — Using Nothing But Docker For Projects — Smashing Magazine

Using Nothing But Docker For Projects — Smashing Magazine

Using only Docker to build and run applications and commands removes the need for previous knowledge in some tool or programming language. Also, it avoids the necessity to install new modules and dependencies directly to the system, which makes development machine-independent.

RARaphael AmoedoRaphael Amoedo·July 14, 2022SRE & Ops
SRE & Ops — Stuff The Internet Says On Scalability For July 11th, 2022 - High Scalability -

Stuff The Internet Says On Scalability For July 11th, 2022 - High Scalability -

Never fear, HighScalability is here! Every cell a universe. Most detailed image of a human cell to date. @microscopicture Other images considered: one byte of RAM in 1946; visual guide on troubleshooting Kubernetes; Cloudflare using lava lamps to generate cryptographic keys; 5MB of data looked like in 1966 My Stuff: * Love this Stuff? I need your support on Patreon to help keep this stuff going. *

HSHigh Scalability·July 11, 2022SRE & Ops
SRE & Ops — GitHub Availability Report: June 2022

GitHub Availability Report: June 2022

In June, we experienced four incidents resulting in significant impact to multiple GitHub.com services. This report also sheds light into an incident that impacted several GitHub.com services in May.

JOJakub OleksyJakub Oleksy·July 6, 2022SRE & Ops
SRE & Ops — Transparent memory offloading: more memory at a fraction of the cost and power

Transparent memory offloading: more memory at a fraction of the cost and power

Transparent memory offloading (TMO) is Meta’s data center solution for offering more memory at a fraction of the cost and power of existing technologies In production since 2021, TMO saves 20 percent to 32 percent of memory per server across millions of servers in our data center fleet We are witnessing massive growth in the […]

MEMeta Engineering·June 20, 2022SRE & Ops