
What the research is: A scalable service management platform for Facebook’s stream processing service. Turbine is designed to bridge the gap between the capabilities of existing general-purpose cluster management frameworks like Tupperware and Facebook’s stream processing requirements. In production for several years now, Turbine has enabled a boom in stream processing at Facebook. Turbine is

Starting at 1531 UTC and lasting until 1952 UTC, the Cloudflare Dashboard and API were unavailable because of the disconnection of multiple, redundant fibre connections from one of our two core data centers.
JG
John Graham Cumming·April 16, 2020SRE & Ops 
We recently migrated hundreds of ZooKeeper instances from individual server instances to Kubernetes without downtime. Our approach used powerful Kubernetes features like endpoints to ease the process, so we’re sharing the high level outline of the approach for anyone who wants to follow in our footsteps. See the end for important networking prerequisites. Click to read more
FMFrancesca McCaffrey·April 14, 2020SRE & Ops 
Over the course of ten weeks, our team of three interns (two engineering, one product management) went from a problem statement to a new feature, which is still working in production for all Cloudflare customers.

As a large portion of Internet access shifted from office-focused areas, like city centers and business parks, towards more residential areas like suburbs and outlying towns, we wanted to find out just precisely how broad this geographical traffic migration was.

Writing is a crucial skill every software developer should cultivate. And writing on your own technical blog can have immense benefits to your career as a software developer and help you cultivate your skills and expertise. Creating and hosting a technical blog provides an opportunity to do just that. In this article, Zara Cooper will take a look at how to deploy a blog for free and with minimal e

In-depth analysis of February service disruptions that impacted GitHub services.

UPDATE: To continue our support of this public NTP service, we have open-sourced our collection of NTP libraries on GitHub. Almost all of the billions of devices connected to the internet have onboard clocks, which need to be accurate to properly perform their functions. Many clocks contain inaccurate internal oscillators, which can cause seconds of […]


![SRE & Ops — [CSP] Unsafe-inline and nonce deployment](/covers/64a9442974.webp?v=8564865)




The refined UI for Build and Serverless Function logs makes consuming logs a pleasure.

Replicated is a 5-year old infrastructure software company with a focus on enabling a new model of enterprise software delivery that we call Kubernetes Off-The-Shelf (KOTS) Software.


Facebook’s codebase changes frequently each day as engineers develop new features and optimizations for our apps. If not handled properly, each of these changes could potentially regress performance for billions of people around the world. At each step in the development process, we apply a suite of automated regression detection tools to mitigate these risks […]

We recently hosted the third annual ZooKeeper Meetup@Facebook in Menlo Park, featuring technical talks focused on new performance, scalability, and security features implemented at large-scale companies with complex distributed systems. Events such as these are important for building community among peers across the industry. Speakers from Cloudera, Facebook, Salesforce, San José State University,

Hear from John Allspaw – former CTO of Etsy and one of the leading members of the DevOps movement – as he explains how Resilience Engineering techniques can help us understand incidents in online software worlds in a broader and deeper way.
SESpotify Engineering·February 1, 2020SRE & Ops 
The root cause of our recent service outage and our next steps

WordPress adoption is massive. So why would a WordPress site consider moving to JAMstack? In this technical case study, Sarah Drasner will cover what an actual WordPress migration looks like, using Smashing Magazine itself! She’ll talk through the gains and losses, the things she wishes she knew earlier, and what she was surprised by. Let’s dig in!

Apache Airflow is a tool for describing, executing, and monitoring workflows. At Slack, we use Airflow to orchestrate and manage our data warehouse workflows, which includes product and business metrics and also is used for different engineering use-cases (e.g. search and offline indexing). For two years we’ve been running Airflow 1.8, and it was time for…
SESlack Engineering·January 15, 2020SRE & Ops 
Last year was a busy one for our open source engineers. In 2019 we released 170 new open source projects, bringing our portfolio to a total of 579 active repositories. While it’s important for our internal engineers to contribute to these projects (and they certainly do — with more than 82,000 commits this year), we […]


A checklist for developers to make sure their passkey implementations are following all the best practices.