Home/SRE & Ops
Topic

SRE & Ops

446 articles on SRE & Ops.

11,834 articles
SRE & Ops — Turbine: Facebook’s service management platform for stream processing

Turbine: Facebook’s service management platform for stream processing

What the research is: A scalable service management platform for Facebook’s stream processing service. Turbine is designed to bridge the gap between the capabilities of existing general-purpose cluster management frameworks like Tupperware and Facebook’s stream processing requirements. In production for several years now, Turbine has enabled a boom in stream processing at Facebook. Turbine is

MEMeta Engineering·April 21, 2020SRE & Ops
SRE & Ops — HubSpot Migrating ZooKeeper into Kubernetes - High Scalability -

HubSpot Migrating ZooKeeper into Kubernetes - High Scalability -

We recently migrated hundreds of ZooKeeper instances from individual server instances to Kubernetes without downtime. Our approach used powerful Kubernetes features like endpoints to ease the process, so we’re sharing the high level outline of the approach for anyone who wants to follow in our footsteps. See the end for important networking prerequisites. Click to read more

FMFrancesca McCaffrey·April 14, 2020SRE & Ops
SRE & Ops — Create Your Free Developer Blog Using Hugo And Firebase — Smashing Magazine

Create Your Free Developer Blog Using Hugo And Firebase — Smashing Magazine

Writing is a crucial skill every software developer should cultivate. And writing on your own technical blog can have immense benefits to your career as a software developer and help you cultivate your skills and expertise. Creating and hosting a technical blog provides an opportunity to do just that. In this article, Zara Cooper will take a look at how to deploy a blog for free and with minimal e

ZCZara CooperZara Cooper·April 6, 2020SRE & Ops
SRE & Ops — Building a more accurate time service at Facebook scale

Building a more accurate time service at Facebook scale

UPDATE: To continue our support of this public NTP service, we have open-sourced our collection of NTP libraries on GitHub. Almost all of the billions of devices connected to the internet have onboard clocks, which need to be accurate to properly perform their functions. Many clocks contain inaccurate internal oscillators, which can cause seconds of […]

MEMeta Engineering·March 18, 2020SRE & Ops
SRE & Ops — How Replicated Developers Develop Remotely

How Replicated Developers Develop Remotely

Replicated is a 5-year old infrastructure software company with a focus on enabling a new model of enterprise software delivery that we call Kubernetes Off-The-Shelf (KOTS) Software.

CCloudflare·March 10, 2020SRE & Ops
SRE & Ops — Preventing performance regressions with Health Compass and Incident Tracker

Preventing performance regressions with Health Compass and Incident Tracker

Facebook’s codebase changes frequently each day as engineers develop new features and optimizations for our apps. If not handled properly, each of these changes could potentially regress performance for billions of people around the world. At each step in the development process, we apply a suite of automated regression detection tools to mitigate these risks […]

MEMeta Engineering·March 5, 2020SRE & Ops
SRE & Ops — ZooKeeper Meetup@Facebook: Advancing the state of distributed coordination

ZooKeeper Meetup@Facebook: Advancing the state of distributed coordination

We recently hosted the third annual ZooKeeper Meetup@Facebook in Menlo Park, featuring technical talks focused on new performance, scalability, and security features implemented at large-scale companies with complex distributed systems. Events such as these are important for building community among peers across the industry. Speakers from Cloudera, Facebook, Salesforce, San José State University,

MEMeta Engineering·February 6, 2020SRE & Ops
SRE & Ops — How Smashing Magazine Manages Content: Migration From WordPress To JAMstack

How Smashing Magazine Manages Content: Migration From WordPress To JAMstack

WordPress adoption is massive. So why would a WordPress site consider moving to JAMstack? In this technical case study, Sarah Drasner will cover what an actual WordPress migration looks like, using Smashing Magazine itself! She’ll talk through the gains and losses, the things she wishes she knew earlier, and what she was surprised by. Let’s dig in!

SSarahdrasner·January 28, 2020SRE & Ops
SRE & Ops — Reliably Upgrading Apache Airflow at Slack’s Scale

Reliably Upgrading Apache Airflow at Slack’s Scale

Apache Airflow is a tool for describing, executing, and monitoring workflows. At Slack, we use Airflow to orchestrate and manage our data warehouse workflows, which includes product and business metrics and also is used for different engineering use-cases (e.g. search and offline indexing). For two years we’ve been running Airflow 1.8, and it was time for…

SESlack Engineering·January 15, 2020SRE & Ops
SRE & Ops — Open source year in review

Open source year in review

Last year was a busy one for our open source engineers. In 2019 we released 170 new open source projects, bringing our portfolio to a total of 579 active repositories. While it’s important for our internal engineers to contribute to these projects (and they certainly do — with more than 82,000 commits this year), we […]

MEMeta Engineering·January 13, 2020SRE & Ops