Home/SRE & Ops
Topic

SRE & Ops

446 articles on SRE & Ops.

11,834 articles
SRE & Ops — Journey to 1000 models: Scaling Instagram’s recommendation system

Journey to 1000 models: Scaling Instagram’s recommendation system

In this post, we explore how Instagram has successfully scaled its algorithm to include over 1000 ML models without sacrificing recommendation quality or reliability. We delve into the intricacies of managing such a vast array of models, each with its own performance characteristics and product goals. We share insights and lessons learned along the way—from […]

CWChris Wiltz·May 21, 2025SRE & Ops
SRE & Ops — Meta’s Full-stack HHVM optimizations for GenAI

Meta’s Full-stack HHVM optimizations for GenAI

As Meta has launched new, innovative products leveraging generative AI (GenAI), we need to make sure the underlying infrastructure components evolve along with it. Applying infrastructure knowledge and optimizations have allowed us to adapt to changing product requirements, delivering a better product along the way. Ultimately, our infrastructure systems need to balance our need to […]

CWChris Wiltz·May 20, 2025SRE & Ops
SRE & Ops — Introducing Pyrefly: A new type checker and IDE experience for Python

Introducing Pyrefly: A new type checker and IDE experience for Python

Today we are announcing an alpha version of Pyrefly, an open source Python type checker and IDE extension crafted in Rust. Pyrefly is a static type checker that analyzes Python code to ensure type consistency and help you catch errors throughout your codebase before your code runs. It also supports IDE integration and CLI usage […]

CWChris Wiltz·May 15, 2025SRE & Ops
SRE & Ops — Enhancing the Python ecosystem with type checking and free threading

Enhancing the Python ecosystem with type checking and free threading

Meta and Quansight have improved key libraries in the Python Ecosystem. There is plenty more to do and we invite the community to help with our efforts. We’ll look at two key efforts in Python’s packaging ecosystem to make packages faster and easier to use: 🚀 Unlock performance wins for developers through free-threaded Python – […]

CWChris Wiltz·May 5, 2025SRE & Ops
SRE & Ops — Expanding observability on Vercel

Expanding observability on Vercel

The Vercel Marketplace adds new native integrations from Sentry, Checkly, and Dash0. Use the tools you already trust to monitor, measure, and debug your apps with integrated billing, single sign-on, and access to provider dashboards from Vercel.

VVercel·April 8, 2025SRE & Ops
SRE & Ops — Mobile GraphQL at Meta in 2025

Mobile GraphQL at Meta in 2025

Mobile GraphQL is a framework used at Meta for fetching data in mobile applications using GraphQL, a strongly-typed, declarative query language. At Meta it handles data fetching for apps like Facebook and Instagram. Sabrina, a software engineer on Meta’s Mobile GraphQL Platform Team, joins Pascal Hartig on the Meta Tech podcast to discuss the evolution […]

CWChris Wiltz·March 31, 2025SRE & Ops
SRE & Ops — Impromptu disaster recovery

Impromptu disaster recovery

Background im-promp-tu ( im-ˈpräm(p)-(ˌ)tü ) made, done, or formed on or as if on the spur of the moment: improvised composed or uttered without previous preparation: extemporaneous Merriam-Webster On March 18th, 2025, I thought I would look into self-hosted project management solutions — something kanban-y, but.. better? This one does not spark joy. After discovering that Teamhood was awesome (an

AWAmos WengerAmos Wenger·March 28, 2025SRE & Ops
SRE & Ops — Cloudflare incident on March 21, 2025

Cloudflare incident on March 21, 2025

On March 21, 2025, multiple Cloudflare services, including R2 object storage experienced an elevated rate of error responses. Here’s what caused the incident, the impact, and how we are making sure it doesn’t happen again.

PPhillipPhillip·March 25, 2025SRE & Ops
SRE & Ops — A case for QLC SSDs in the data center

A case for QLC SSDs in the data center

The growth of data and need for increased power efficiency are leading to innovative storage solutions. HDDs have been growing in density, but not performance, and TLC flash remains at a price point that is restrictive for scaling. QLC technology addresses these challenges by forming a middle tier between HDDs and TLC SSDs. QLC […]

CWChris Wiltz·March 4, 2025SRE & Ops
SRE & Ops — Cloudflare incident on February 6, 2025

Cloudflare incident on February 6, 2025

On Thursday, February 6, 2025, we experienced an outage with our object storage service (R2) and products that rely on it. Here's what happened and what we're doing to fix this going forward.

SJSilverlock, JavierSilverlock, Javier·February 7, 2025SRE & Ops
SRE & Ops — Indexing code at scale with Glean

Indexing code at scale with Glean

We’re sharing details about Glean, Meta’s open source system for collecting, deriving, and working with facts about source code. In this blog post we’ll talk about why a system like Glean is important, explain the rationale for Glean’s design, and run through some of the ways we’re using Glean to supercharge our developer tooling at […]

CWChris Wiltz·December 19, 2024SRE & Ops
SRE & Ops — Cloudflare incident on November 14, 2024, resulting in lost logs

Cloudflare incident on November 14, 2024, resulting in lost logs

On November 14, 2024, Cloudflare experienced a Cloudflare Logs outage, impacting the majority of customers using these products. During the ~3.5 hours that these services were impacted, about 55% of the logs we normally send to customers were not sent and were lost. The details of what went wrong and why are interesting both for customers and practitioners.

JTJamie, Tom WalwynJamie, Tom Walwyn·November 26, 2024SRE & Ops
SRE & Ops — Sequence learning: A paradigm shift for personalized ads recommendations

Sequence learning: A paradigm shift for personalized ads recommendations

AI plays a fundamental role in creating valuable connections between people and advertisers within Meta’s family of apps. Meta’s ad recommendation engine, powered by deep learning recommendation models (DLRMs), has been instrumental in delivering personalized ads to people. Key to this success was incorporating thousands of human-engineered signals or features in the DLRM-based recommendation syst

CWChris Wiltz·November 19, 2024SRE & Ops
SRE & Ops — There’s No Such Thing as a Free Lunch!

There’s No Such Thing as a Free Lunch!

Incident Management takes time Incidents need responders that are trained and experienced. At Slack, training is a foundation of our incident management program. Self-service training and live courses based mainly on prepared content are one piece of the puzzle, but there can be a missing piece in many organizations. How can staff get practical experience…

SNScott Nelson WindelsScott Nelson Windels·November 14, 2024SRE & Ops
SRE & Ops — Fearless SSH: short-lived certificates bring Zero Trust to infrastructure

Fearless SSH: short-lived certificates bring Zero Trust to infrastructure

Access for Infrastructure, BastionZero’s integration into Cloudflare One, will enable organizations to apply Zero Trust controls to their servers, databases, Kubernetes clusters, and more. Today we’re announcing short-lived SSH access as the first available feature of this integration.

GAGoldbe, Ann Ming SamborskiGoldbe, Ann Ming Samborski·October 23, 2024SRE & Ops
SRE & Ops — OCP Summit 2024: The open future of networking hardware for AI

OCP Summit 2024: The open future of networking hardware for AI

At Open Compute Project Summit (OCP) 2024, we’re sharing details about our next-generation network fabric for our AI training clusters. We’ve expanded our network hardware portfolio and are contributing two new disaggregated network fabrics and a new NIC to OCP. We look forward to continued collaboration with OCP to open designs for racks, servers, storage […]

CWChris Wiltz·October 15, 2024SRE & Ops
SRE & Ops — Meta’s open AI hardware vision

Meta’s open AI hardware vision

At the Open Compute Project (OCP) Global Summit 2024, we’re showcasing our latest open AI hardware designs with the OCP community. These innovations include a new AI platform, cutting-edge open rack designs, and advanced network fabrics and components. By sharing our designs, we hope to inspire collaboration and foster innovation. If you’re passionate about building […]

CWChris Wiltz·October 15, 2024SRE & Ops
SRE & Ops — Leveraging Kubernetes virtual machines at Cloudflare with KubeVirt

Leveraging Kubernetes virtual machines at Cloudflare with KubeVirt

The Kubernetes team runs several multi-tenant clusters across Cloudflare’s core data centers. When multi-tenant cluster isolation is too limiting for an application, we use KubeVirt. KubeVirt is a cloud-native solution that enables our developers to run virtual machines alongside containers.

JCJustin CichraJustin Cichra·October 8, 2024SRE & Ops
SRE & Ops — Cloudflare incident on September 17, 2024

Cloudflare incident on September 17, 2024

On September 17, 2024, during planned routine maintenance, Cloudflare stopped announcing 15 IPv4 prefixes, affecting some Business plan websites for approximately one hour. During this time, IPv4 traffic for these customers would not have reached Cloudflare and users attempting to connect to websites using addresses within those prefixes would have received errors.

JAJoe AbleyJoe Abley·September 20, 2024SRE & Ops
SRE & Ops — Advancing Our Chef Infrastructure

Advancing Our Chef Infrastructure

At Slack, we manage tens of thousands of EC2 instances that host a variety of services, including our Vitess databases, Kubernetes workers, and various components of the Slack application. The majority of these instances run on some version of Ubuntu, while a portion operates on Amazon Linux. With such a vast infrastructure, the critical question…

AGArchie Gunasekara·September 17, 2024SRE & Ops
SRE & Ops — Simulator-based reinforcement learning for data center cooling optimization

Simulator-based reinforcement learning for data center cooling optimization

We’re sharing more about the role that reinforcement learning plays in helping us optimize our data centers’ environmental controls. Our reinforcement learning-based approach has helped us reduce energy consumption and water usage across various weather conditions in our data centers. Meta is revamping its new data center design to optimize for artificial intelligence and the […]

CWChris Wiltz·September 10, 2024SRE & Ops
SRE & Ops — RETINAS: Real-Time Infrastructure Accounting for Sustainability

RETINAS: Real-Time Infrastructure Accounting for Sustainability

We are introducing a new metric— real-time server fleet utilization effectiveness —as part of the RETINAS initiative to help reduce emissions and achieve net zero emissions across our value chain in 2030. This new metric allows us to measure server resource usage (e.g., compute, storage) and efficiency in our large-scale data center server fleet in […]

CWChris Wiltz·August 26, 2024SRE & Ops
SRE & Ops — DCPerf: An open source benchmark suite for hyperscale compute applications

DCPerf: An open source benchmark suite for hyperscale compute applications

We are open-sourcing DCPerf, a collection of benchmarks that represents the diverse categories of workloads that run in data center cloud deployments. We hope that DCperf can be used more broadly by academia, the hardware industry, and internet companies to design and evaluate future products. DCPerf is available now on GitHub. Hyperscale and cloud datacenter […]

CWChris Wiltz·August 5, 2024SRE & Ops
SRE & Ops — A recent spate of Internet disruptions

A recent spate of Internet disruptions

Cloudflare Radar is constantly monitoring the Internet for widespread disruptions. Here we examine several recent noteworthy disruptions detected in the first month of Q3, including traffic anomalies observed in Bangladesh, Syria, Pakistan, and Venezuela.

DBDavid BelsonDavid Belson·August 1, 2024SRE & Ops
SRE & Ops — AI Lab: The secrets to keeping machine learning engineers moving fast

AI Lab: The secrets to keeping machine learning engineers moving fast

The key to developer velocity across AI lies in minimizing time to first batch (TTFB) for machine learning (ML) engineers. AI Lab is a pre-production framework used internally at Meta. It allows us to continuously A/B test common ML workflows – enabling proactive improvements and automatically preventing regressions on TTFB. AI Lab prevents TTFB regressions […]

CWChris Wiltz·July 16, 2024SRE & Ops
SRE & Ops — Meta’s approach to machine learning prediction robustness

Meta’s approach to machine learning prediction robustness

Meta’s advertising business leverages large-scale machine learning (ML) recommendation models that power millions of ads recommendations per second across Meta’s family of apps. Maintaining reliability of these ML systems helps ensure the highest level of service and uninterrupted benefit delivery to our users and advertisers. To minimize disruptions and ensure our ML systems are intrinsically

CWChris Wiltz·July 10, 2024SRE & Ops
SRE & Ops — The key to a happy Rust/C++ relationship

The key to a happy Rust/C++ relationship

The history of Rust at Meta goes all the way back to 2016, when we first started using it for source control. Today, it has been widely embraced at Meta and is one of our primary supported server-side languages (along with C++, Python, and Hack). But that doesn’t mean there weren’t any growing pains. Aida […]

CWChris Wiltz·June 25, 2024SRE & Ops
SRE & Ops — Leveraging AI for efficient incident response

Leveraging AI for efficient incident response

We’re sharing how we streamline system reliability investigations using a new AI-assisted root cause analysis system. The system uses a combination of heuristic-based retrieval and large language model-based ranking to speed up root cause identification during investigations. Our testing has shown this new system achieves 42% accuracy in identifying root causes for investigations at their […] Read

CWChris Wiltz·June 24, 2024SRE & Ops
SRE & Ops — PVF: A novel metric for understanding AI systems’ vulnerability against SDCs in model parameters

PVF: A novel metric for understanding AI systems’ vulnerability against SDCs in model parameters

We’re introducing parameter vulnerability factor (PVF), a novel metric for understanding and measuring AI systems’ vulnerability against silent data corruptions (SDCs) in model parameters. PVF can be tailored to different AI models and tasks, adapted to different hardware faults, and even extended to the training phase of AI models. We’re sharing results of our own […]

CWChris Wiltz·June 19, 2024SRE & Ops
SRE & Ops — Advanced Rollout Techniques: Custom Strategies for Stateful Apps in Kubernetes

Advanced Rollout Techniques: Custom Strategies for Stateful Apps in Kubernetes

In a previous blog post—A Simple Kubernetes Admission Webhook—I discussed the process of creating a Kubernetes webhook without relying on Kubebuilder. At Slack, we use this webhook for various tasks, like helping us support long-lived Pods (see Supporting Long-Lived Pods), and today, I delve once more into the topic of long-lived Pods, focusing on our…

CLClément Labbe·June 15, 2024SRE & Ops
SRE & Ops — How Meta trains large language models at scale

How Meta trains large language models at scale

As we continue to focus our AI research and development on solving increasingly complex problems, one of the most significant and challenging shifts we’ve experienced is the sheer scale of computation required to train large language models (LLMs). Traditionally, our AI model training has involved a training massive number of models that required a comparatively […]

CWChris Wiltz·June 12, 2024SRE & Ops
SRE & Ops — Maintaining large-scale AI capacity at Meta

Maintaining large-scale AI capacity at Meta

Meta is currently operating many data centers with GPU training clusters across the world. Our data centers are the backbone of our operations, meticulously designed to support the scaling demands of compute and storage. A year ago, however, as the industry reached a critical inflection point due to the rise of artificial intelligence (AI), we […]

CWChris Wiltz·June 12, 2024SRE & Ops
SRE & Ops — Unlocking the power of mixed reality devices with MobileConfig

Unlocking the power of mixed reality devices with MobileConfig

MobileConfig enables developers to centrally manage a mobile app’s configuration parameters in our data centers. Once a parameter value is changed on our central server, billions of app devices automatically fetch and apply the new value without app updates. These remotely managed configuration parameters serve various purposes such as A/B testing, feature rollout, and app […]

CWChris Wiltz·June 11, 2024SRE & Ops
SRE & Ops — Serverless Jupyter Notebooks at Meta

Serverless Jupyter Notebooks at Meta

At Meta, Bento, our internal Jupyter notebooks platform, is a popular tool that allows our engineers to mix code, text, and multimedia in a single document. Use cases run the entire spectrum from what we call “lite” workloads that involve simple prototyping to heavier and more complex machine learning workflows. However, even though the lite […]

CWChris Wiltz·June 10, 2024SRE & Ops
SRE & Ops — Adopting OpenTelemetry for our logging pipeline

Adopting OpenTelemetry for our logging pipeline

Recently, Cloudflare's Observability team undertook an effort to migrate our existing syslog-ng backed logging infrastructure to instead being backed by OpenTelemetry Collectors. In this post, we detail the process that we undertook, and the difficulties we faced along the way

CCloudflare·June 3, 2024SRE & Ops
SRE & Ops — Introducing no-code migrations to Stripe Billing

Introducing no-code migrations to Stripe Billing

Last month we announced a new Billing migration toolkit designed to address your top pain points: engineering resourcing, the risk of disruptive errors, and migration time. Here’s how the migration toolkit addresses each.

DCDivyanti ChauhanDivyanti Chauhan·May 28, 2024SRE & Ops
SRE & Ops — Composable data management at Meta

Composable data management at Meta

In recent years, Meta’s data management systems have evolved into a composable architecture that creates interoperability, promotes reusability, and improves engineering efficiency. We’re sharing how we’ve achieved this, in part, by leveraging Velox, Meta’s open source execution engine, as well as work ahead as we continue to rethink our data management systems. Data is at […]

CWChris Wiltz·May 22, 2024SRE & Ops
SRE & Ops — Behind the scenes of Threads for web

Behind the scenes of Threads for web

When Threads first launched one of the top feature requests was for a web client. In this episode of the Meta Tech Podcast, Pascal Hartig (@passy) sits down with Ally C. and Kevin C., two engineers on the Threads Web Team that delivered the basic version of Threads for web in just under three months. […]

CWChris Wiltz·May 14, 2024SRE & Ops