Home/SRE & Ops
Topic

SRE & Ops

446 articles on SRE & Ops.

11,834 articles
SRE & Ops — GitHub in the enterprise

GitHub in the enterprise

During the last year alone, over 56 million developers created more than 60 million new repos and made more than 1.9 billion contributions on GitHub. These developers are building the…

EBErica BresciaErica Brescia·December 9, 2020SRE & Ops
SRE & Ops — How Facebook keeps its large-scale infrastructure hardware up and running

How Facebook keeps its large-scale infrastructure hardware up and running

Facebook’s services rely on fleets of servers in data centers all over the globe — all running applications and delivering the performance our services need. This is why we need to make sure our server hardware is reliable and that we can manage server hardware failures at our scale with as little disruption to our […]

CWChris Wiltz·December 9, 2020SRE & Ops
SRE & Ops — The State of Ruby Static Typing at Shopify - Shopify

The State of Ruby Static Typing at Shopify - Shopify

Shopify changes a lot. We merge around 400 commits to the main branch daily and deploy a new version of our core monolith 40 times a day. The Monolith is also big: 37,000 Ruby files, 622,000 methods, more than 2,000,000 calls. At this scale with a dynamic language, even with the most rigorous review process and over 150,000 automated tests, it’s a challenge to ensure everything works properly. Dev

SEShopify Engineering·December 7, 2020SRE & Ops
SRE & Ops — Your DevOps guide to GitHub Universe

Your DevOps guide to GitHub Universe

GitHub Universe is almost here. For more on what to expect from this year’s stream, we sat down with virtual host, Brian Douglas, for a quick Q&A on GitHub Actions,…

GMGrace MadlingerGrace Madlinger·December 4, 2020SRE & Ops
SRE & Ops — The evolving role of operations in DevOps

The evolving role of operations in DevOps

GitHub’s team delves into answering the question “what are operations roles in the development and operations (DevOps) environments”. From automating the role of QA in DevOps and more for smaller, faster delivery cycles.

JMJared MurrellJared Murrell·December 3, 2020SRE & Ops
SRE & Ops — Vouching for Docker Images - Shopify

Vouching for Docker Images - Shopify

If you were using computers in the ‘90s and the early 2000s, you probably had the experience of installing a piece of software you downloaded from the internet, only to discover that someone put some nasty into it, and now you’re dragging your computer to IT to beg them to save your data. To remedy this, software developers started “signing” their software in a way that proved both who they were a

SEShopify Engineering·December 1, 2020SRE & Ops
SRE & Ops — A Byzantine failure in the real world

A Byzantine failure in the real world

At Cloudflare, we are always on the lookout for Single Points of Failure. In this post, we explore the role a failure mode known as a Byzantine fault played in a a real-world incident.

TLTom Lianza, Chris SnookTom Lianza, Chris Snook·November 27, 2020SRE & Ops
SRE & Ops — Migrating Slack Airflow to Python 3 Without Disruption

Migrating Slack Airflow to Python 3 Without Disruption

Last year, we migrated Airflow from 1.8 to 1.10 at Slack (see here) and we did a “Big bang” upgrade because of the constraints we had. This year, due to Python 2 reaching end of life, we again had a major migration of Airflow from Python 2 to 3 and we wanted to put our…

ASAshwin Sai Shankar·November 19, 2020SRE & Ops
SRE & Ops — Don't put data science notebooks into production

Don't put data science notebooks into production

We've come across many clients who are interested in taking the computational notebooks developed by their data scientists, and putting them directly into the codebase of production applications. Data science ideas do need to move out of notebooks and into production, but trying to deploy that notebooks as a code artifact breaks a multitude of good software practices. Predictably, that results in

DJDavid Johnston·November 18, 2020SRE & Ops
SRE & Ops — Automated Origin CA for Kubernetes

Automated Origin CA for Kubernetes

Today we're releasing origin-ca-issuer, an extension to cert-manager integrating with Cloudflare Origin CA to easily create and renew certificates for your account's domains.

TSTerin StockTerin Stock·November 13, 2020SRE & Ops
SRE & Ops — Building a ubiquitous shared infrastructure using Twine

Building a ubiquitous shared infrastructure using Twine

What the research is: Twine is our homegrown cluster management system, which has been running in production for the past decade. A cluster management system allocates workloads to machines and manages the life cycle of machines, containers, and workloads. Kubernetes is a prominent example of an open source cluster management system. Twine has helped convert […]

MEMeta Engineering·November 11, 2020SRE & Ops
SRE & Ops — Getting started with DevOps automation

Getting started with DevOps automation

This is the second post in our series on DevOps fundamentals. For a guide to what DevOps is and answers to common DevOps myths check out part one. What role…

JMJared MurrellJared Murrell·October 29, 2020SRE & Ops
SRE & Ops — Introducing resctl-demo: Better resource control with simulation

Introducing resctl-demo: Better resource control with simulation

What it is: We all want our workloads to run on our machines as efficiently as possible. With the Facebook resource control demo (resctl-demo), developers can simulate system resource conflicts and test ways to resolve them. It’s like going on a guided tour through a system’s resource control, complete with live demos and detailed explanations. […]

CWChris Wiltz·October 14, 2020SRE & Ops
SRE & Ops — Nemo: Data discovery at Facebook

Nemo: Data discovery at Facebook

Large-scale companies serve millions or even billions of people who depend on the services these companies provide for their everyday needs. To keep these services running and delivering meaningful experiences, the teams behind them need to find the most relevant and accurate information quickly so that they can make informed decisions and take action. Finding […]

MEMeta Engineering·October 9, 2020SRE & Ops
SRE & Ops — GitHub Welcomes the OpenJDK Project!

GitHub Welcomes the OpenJDK Project!

Earlier this month we were thrilled to welcome the OpenJDK Community to GitHub. The communities migration effort, codenamed Project ‘Skara’, brought JDK 16 main-line development into GitHub. The JDK project is at the…

MWMartin WoodwardMartin Woodward·September 30, 2020SRE & Ops
SRE & Ops — Lightning Q&A: DevSecOps in five with Maya Kaczorowski

Lightning Q&A: DevSecOps in five with Maya Kaczorowski

In this interview, we dig deeper with Maya Kaczorowski on what DevSecOps is, and how to apply it. It’s a mindset shift in how development teams think about security. DevSecOps is about making all parties who are part of the application development lifecycle accountable for security of the application.

GMGrace MadlingerGrace Madlinger·September 24, 2020SRE & Ops
SRE & Ops — Mark Harman elected Fellow of the Royal Academy of Engineering

Mark Harman elected Fellow of the Royal Academy of Engineering

The U.K.’s Royal Academy of Engineering has elected Facebook Research Scientist Mark Harman as a Fellow for his achievements in academia and industry, including his work on search-based software engineering (SBSE), intelligent software testing tools, and web-enabled simulation (WES) approaches. Election to the Academy is by invitation only, and it is one of the highest […]

MEMeta Engineering·September 21, 2020SRE & Ops
SRE & Ops — The next decade: How Facebook is stepping up the fight against climate change

The next decade: How Facebook is stepping up the fight against climate change

In 2018, we set an ambitious goal of achieving a 75 percent absolute reduction in operational greenhouse gas (GHG) emissions and to support our global operations with 100 percent renewable energy by the end of 2020. On both counts, we are well on track: These commitments have spurred the construction of over 5,400 megawatts (MW) […]

MEMeta Engineering·September 14, 2020SRE & Ops
SRE & Ops — An Introduction To Running Lighthouse Programmatically — Smashing Magazine

An Introduction To Running Lighthouse Programmatically — Smashing Magazine

Being able to run Google’s Lighthouse analysis suite programmatically provides a lot of advantages, especially for larger or more complex web applications. Using Lighthouse programmatically allows engineers to set up quality monitoring for sites that need more customization than straightforward applications of Lighthouse (such as Lighthouse CI) allow. This article contains a brief introduction to

KBKaty BowmanKaty Bowman·September 11, 2020SRE & Ops
SRE & Ops — Fault tolerance through optimal workload placement

Fault tolerance through optimal workload placement

As our infrastructure has expanded, we have seen an exponential growth in failures that affect a subset of capacity in a data center. These may stem from software or firmware errors, or issues in the mechanical or electrical equipment in the data center. As our data center footprint grows, so does the frequency of such […]

CWChris Wiltz·September 8, 2020SRE & Ops
SRE & Ops — Containerizing ZooKeeper with Twine: Powering container orchestration from within

Containerizing ZooKeeper with Twine: Powering container orchestration from within

Hardware fails, networks partition, and humans break things. The job of our infrastructure engineers is to abstract these realities away and provide a reliable, stable production environment nonetheless. Two of the technologies we deploy in this pursuit are Twine, our internal container orchestrator, and Apache ZooKeeper. Twine maintains service availability by managing services in containers

MEMeta Engineering·August 31, 2020SRE & Ops
SRE & Ops — August 30th 2020: Analysis of CenturyLink/Level(3) outage

August 30th 2020: Analysis of CenturyLink/Level(3) outage

Today CenturyLink/Level(3), a major ISP and Internet bandwidth provider, experienced a significant outage that impacted some of Cloudflare’s customers as well as a significant number of other services and providers across the Internet.

MPMatthew PrinceMatthew Prince·August 30, 2020SRE & Ops
SRE & Ops — A Bit on CI/CD

A Bit on CI/CD

I'd say "website" fits better than "mobile app" but I like this framing from Max Lynch:

CCChris Coyier·August 26, 2020SRE & Ops
SRE & Ops — Scaling services with Shard Manager

Scaling services with Shard Manager

Over the years, as we’ve expanded in scale and functionalities, Facebook has evolved from a basic web server architecture into a complex one with thousands of services working behind the scenes. It’s no trivial task to scale the wide range of back-end services needed for Facebook’s products. And we found that many of our teams […]

CWChris Wiltz·August 24, 2020SRE & Ops
SRE & Ops — Introducing Deploy Buttons

Introducing Deploy Buttons

Deploy Buttons help you deploy a project to the Workers Platform without even needing to set up a local development environment. Now, it’s as easy as clicking a Deploy Button and three short steps to deploy using our new web-based deploy tool.

DSDavid SongDavid Song·August 20, 2020SRE & Ops
SRE & Ops — How We Improved Developer Productivity for Our DevOps Teams

How We Improved Developer Productivity for Our DevOps Teams

Across Spotify, our teams diligently strive to fulfill our mission to “unlock the potential of human creativity by giving millions of creative artists the opportunity to live off their work, and billions of fans the opportunity to enjoy and be inspired by it”. As product managers in the Platform Developer Experience (PDX) Tribe, part of Spotify’s Technology Infrastructure Group, we focus on unlock

MJMaria Jernström and Jason Palmer·August 1, 2020SRE & Ops
SRE & Ops — The Migration of Legacy Applications to Workers

The Migration of Legacy Applications to Workers

As Cloudflare Workers, and other Serverless platforms, continue to drive down costs while making it easier for developers to stand up globally scaled applications, the migration of legacy applications is becoming increasingly common.

CCloudflare·July 28, 2020SRE & Ops
SRE & Ops — Scalable data classification for security and privacy

Scalable data classification for security and privacy

What the research is: We’ve built a data classification system that uses multiple data signals, a scalable system architecture, and machine learning to detect semantic types within Facebook at scale. This is important in situations where it’s necessary to detect where an organization’s data is stored in many different formats across various data stores. In […]

MEMeta Engineering·July 21, 2020SRE & Ops
SRE & Ops — Introducing the GitHub Availability Report

Introducing the GitHub Availability Report

What is the Availability Report? Historically, GitHub has published post-incident reviews for major incidents that impact service availability. Whether we’re sharing new investments to infrastructure or detailing site downtimes, our…

KBKeith BallingerKeith Ballinger·July 8, 2020SRE & Ops
SRE & Ops — Introducing our 2019 Sustainability Report

Introducing our 2019 Sustainability Report

As the world navigates the COVID-19 pandemic, companies are focusing not only on their business strategy, but also on their approach to key environmental and social issues. One of several areas we are focusing on now is our commitment to sustainable business practices and reporting. Today, we are taking another important step toward increased transparency […]

MEMeta Engineering·July 7, 2020SRE & Ops
SRE & Ops — Retrie: Haskell refactoring made easy

Retrie: Haskell refactoring made easy

What’s new: We’ve open-sourced Retrie, a code refactoring tool for Haskell that makes codemodding faster, easier, and safer. Using Retrie, developers can efficiently rewrite large codebases (exceeding 1 million lines), express rewrites as equations in Haskell syntax instead of regular expressions, and avoid large classes of codemodding errors. Retrie’s features include the ability to rewrite […] R

CWChris Wiltz·July 6, 2020SRE & Ops
SRE & Ops — Leveraging Mobile Infrastructure with Data-Driven Decisions

Leveraging Mobile Infrastructure with Data-Driven Decisions

TL;DR The pursuit wasn’t always easy, but “putting data first” has helped Spotify dramatically improve the infrastructure team’s decision-making process while improving company productivity. Raul Herbster describes the benefits, especially for mobile DevOps teams.

RHRaul Herbster·July 1, 2020SRE & Ops
SRE & Ops — All Hands on Deck

All Hands on Deck

This story speaks to the process behind incident response at Slack and uses the May 12th, 2020 outage as an example. For a deeper technical review of the same outage, read Laura Nolan’s post, “A Terrible, Horrible, No-Good, Very Bad Day at Slack” Slack is a critical tool for millions of people, so it’s natural when…

RRoss·June 29, 2020SRE & Ops
SRE & Ops — A Terrible, Horrible, No-Good, Very Bad Day at Slack

A Terrible, Horrible, No-Good, Very Bad Day at Slack

This story describes the technical details of the problems that caused the Slack downtime on May 12th, 2020. To learn more about the process behind incident response for same outage, read Ryan Katkov’s post, “All Hands on Deck”. On May 12, 2020, Slack had our first significant outage in a long time. We published a summary…

RRoss·June 29, 2020SRE & Ops
SRE & Ops — GitHub Action Hero: Victoria Drake

GitHub Action Hero: Victoria Drake

GitHub Actions allows you to automate your workflow. With GitHub Actions, you can deploy to any cloud, build containers, automate messages, and do so much more. Use any tool you…

MDMichelle DukeMichelle Duke·June 26, 2020SRE & Ops
SRE & Ops — Dependabot now updates your Actions workflows

Dependabot now updates your Actions workflows

GitHub Actions makes it easy to automate all your software workflows, from continuous integration and delivery to issue triage and more. Whether you want to build a container, deploy a…

AMAlex Mullans·June 25, 2020SRE & Ops
SRE & Ops — Accelerometer and SoftSKU: Improving hardware platform performance for diverse microservices

Accelerometer and SoftSKU: Improving hardware platform performance for diverse microservices

What the research is: New strategies to improve the performance of hardware platforms running Facebook’s microservices. SoftSKU is a novel mechanism that tailors an existing server processor to optimize it for a specific microservice without requiring any additional hardware. Accelerometer is an analytical model that predicts gains from these optimizations even before the custom hardware […] Read

MEMeta Engineering·May 11, 2020SRE & Ops
SRE & Ops — Releasing kubectl support in Access

Releasing kubectl support in Access

Starting today, you can use Cloudflare Access and Argo Tunnel to securely manage your Kubernetes cluster with the kubectl command-line tool. SSO requirements and a zero-trust model to your Kubernetes management in under 30 minutes.

CCloudflare·April 27, 2020SRE & Ops