Home/Chris Wiltz
Author

Chris Wiltz

281 articles by Chris Wiltz.

SRE & Ops — ZGateway: Learnings from Putting a Proxy in Front of ZippyDB

ZGateway: Learnings from Putting a Proxy in Front of ZippyDB

We’re introducing ZGateway, the proxy we are using to unify traffic through ZippyDB, Meta’s most widely-used key value store. As a bonus, it also enables admission control, load balancing, cross-region resilience, and richer operations. ZippyDB is the most widely used key value store at Meta, backing product metadata, counters, and configuration, and can serve billions […]

CWChris Wiltz·September 3, 2026SRE & Ops
AI & ML — An Organizational Second Brain: Building an AI That Learns From Experts

An Organizational Second Brain: Building an AI That Learns From Experts

We’ve built an AI agent that acts as a secondary expert for a given domain, making deep specialist knowledge readily available and preserved for anyone in an organization to access, share, and build upon. This is not a typical domain-specific agent. Its novelty comes from integrating two layers: A structured, auditable knowledge architecture separates what […]

CWChris Wiltz·September 2, 2026AI & ML
SRE & Ops — MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet

MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet

Training and serving frontier AI models depends on fast, reliable networks that move data between GPUs without wasting compute cycles. To meet this challenge at scale, Meta designed MetaRoCE – a clean-sheet RDMA transport protocol purpose-built for AI workloads on commodity Ethernet. We’re releasing the MetaRoCE specification, a reference software implementation and a compliance test […]

CWChris Wiltz·August 24, 2026SRE & Ops
SRE & Ops — MTIA 300: Meta’s First Training Chip with Built-in NICs and Communication-Offloading Engines

MTIA 300: Meta’s First Training Chip with Built-in NICs and Communication-Offloading Engines

MTIA 300 is the first of Meta’s family of in-house training and inference accelerators optimized for training ranking and recommendation models. We’re sharing how MTIA 300’s built-in NIC chiplets allow it to meet the communication needs associated with training recommendation models with superior performance over general-purpose GPUs. By co-designing MTIA’s communication library, HCCL, alongside t

CWChris Wiltz·August 24, 2026SRE & Ops
Security — How We’re Building Scam Alert on WhatsApp With End-to-End Encryption and Verifiability Guarantees

How We’re Building Scam Alert on WhatsApp With End-to-End Encryption and Verifiability Guarantees

WhatsApp is committed to helping people stay safe while protecting the privacy of their messages. As scam tactics evolve — from impersonation to social engineering to AI-generated lures — we’re always evolving as well, so that our protections stay ahead of scammers while protecting people’s personal messages with end-to-end encryption. Today, we’re sharing an early […]

CWChris Wiltz·August 12, 2026Security
SRE & Ops — From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

Every day, Meta’s recommendation platforms handle billions of user interactions, generating rich temporal signals that capture individual preferences and intent across products, ads, and content. In our 2024 post on sequence learning for ads recommendations, we showed how modeling the order and timing of user actions (rather than relying on static, manually engineered sparse features) […] Read Mor

CWChris Wiltz·August 5, 2026SRE & Ops
AI & ML — GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in […]

CWChris Wiltz·August 3, 2026AI & ML
SRE & Ops — Meta’s AI Storage Blueprint at Scale

Meta’s AI Storage Blueprint at Scale

Over the past several years, model capabilities and training dataset sizes have experienced exponential growth. During the past year or so, the time between new-frontier-model releases has gone down from months to weeks. Reliable and fast access to storage is important to both the speed and computational cost of this AI innovation. If AI is […]

CWChris Wiltz·July 1, 2026SRE & Ops
AI & ML — 10 Years of Meta’s Commitment to Python

10 Years of Meta’s Commitment to Python

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it. Python is one of the world’s most influential programming languages, and we use it across our engineering stack, from […]

CWChris Wiltz·June 30, 2026AI & ML
Career & Teams — How Meta Engineered Ultra-Narrow Batteries for AI Glasses

How Meta Engineered Ultra-Narrow Batteries for AI Glasses

Smart glasses like the Ray-Ban Meta and Oakley Meta Vanguards need to pack enough energy to power features like cameras, speakers, AI workloads, and even a display. But it all has to fit into the glasses’ temple arms. So how do you place a battery with enough power to run a pair of smart glasses […]

CWChris Wiltz·June 23, 2026Career & Teams
Frontend — Adopting AV1 for Real-Time Communication (RTC) at Scale

Adopting AV1 for Real-Time Communication (RTC) at Scale

Adopting AV1 for real-time communication at Meta has been a multi-year effort spanning codec selection, device eligibility, rate control, and error resilience. We’re sharing the technical and operational challenges while deploying AV1 and expanding coverage, and how we addressed them for real-time communication. We’re presenting several technologies for improving AV1 call quality, including rate c

CWChris Wiltz·June 22, 2026Frontend
SRE & Ops — Lights Out, Systems On: Validating Instant Power Loss Readiness

Lights Out, Systems On: Validating Instant Power Loss Readiness

We’re introducing Instantaneous PowerLoss Storm, a new testing paradigm within Meta’s infrastructure for handling and mitigating instant or zero-notice power loss in our data centers. We’re sharing: how we built readiness to tolerate instant failures into our existing systems with defense-in-depth strategies; tradeoffs made in implementing it, and how we validated our readiness. Disaster preparedn

CWChris Wiltz·June 3, 2026SRE & Ops
AI & ML — SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

SilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation Systems

We’re introducing SilverTorch, a reimagining of recommendation systems that unifies all retrieval components for user generated content under a unified architecture. SilverTorch shows up to 23.7x higher throughput compared to the state-of-the-art approaches. It’s also showing 20.9x more compute cost efficiency compared to a CPU-based solution while also improving accuracy. Our research paper, “Sil

CWChris Wiltz·May 26, 2026AI & ML
Frontend — Reel Friends: Building Social Discovery that Scales to Billions

Reel Friends: Building Social Discovery that Scales to Billions

On its face the new Friend Bubbles feature looks simple enough. It highlights Reels your friends have watched and reacted to. But sometimes the features that seem the most straightforward require the deepest engineering work. On this episode of the Meta Tech Podcast, Pascal Hartig chats with Subasree and Joseph, two software engineers from the Facebook […]

CWChris Wiltz·May 13, 2026Frontend
SRE & Ops — Migrating Data Ingestion Systems at Meta Scale

Migrating Data Ingestion Systems at Meta Scale

Meta’s data ingestion system, which our engineering teams leverage for up-to-date snapshots of the social graph, has recently undergone a significant revamp to enhance its reliability at scale. Moving from our legacy system to our new architecture required a large-scale migration of our entire data ingestion system. We’re sharing the solutions and strategies that enabled […]

CWChris Wiltz·May 12, 2026SRE & Ops
Frontend — Labyrinth 1.1: Making End-to-End Encrypted Backups Even More Reliable

Labyrinth 1.1: Making End-to-End Encrypted Backups Even More Reliable

We’re rolling out version 1.1 of Labyrinth, the encrypted storage system and protocol that secures messages and history on Messenger. Labyrinth 1.1 enhances the reliability of end-to-end encrypted backups with a new sub-protocol that helps messages survive the loss of a device, a switched device, and long gaps between sign-ins. Read our updated white paper, […]

CWChris Wiltz·May 11, 2026Frontend
SRE & Ops — How Meta Is Strengthening End-to-End Encrypted Backups

How Meta Is Strengthening End-to-End Encrypted Backups

The HSM-based Backup Key Vault Meta’s HSM-based Backup Key Vault provides the foundation for end-to-end encrypted backups for WhatsApp and Messenger. The system allows people to protect their backed-up message history with a recovery code, ensuring that the recovery code is stored in tamper-resistant hardware security modules (HSMs) and is inaccessible to Meta, cloud storage […]

CWChris Wiltz·May 1, 2026SRE & Ops
AI & ML — Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

Modernizing the Facebook Groups Search to Unlock the Power of Community Knowledge

We’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them. We’ve adopted a new hybrid retrieval architecture and implemented automated model-based evaluation to address the major friction points people experience when searching community content. Under this new framework, we’ve made tangib

CWChris Wiltz·April 21, 2026AI & ML
SRE & Ops — Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

We’re sharing insights into Meta’s Capacity Efficiency Program, where we’ve built an AI agent platform that helps automate finding and fixing performance issues throughout our infrastructure. By leveraging encoded domain expertise across a unified, standardized tool interface these agents help save power and free up engineers’ time away from addressing performance issues to innovating on […] Read

CWChris Wiltz·April 16, 2026SRE & Ops
Security — Post-Quantum Cryptography Migration at Meta: Framework, Lessons, and Takeaways

Post-Quantum Cryptography Migration at Meta: Framework, Lessons, and Takeaways

We’re sharing lessons learned from Meta’s post-quantum cryptography (PQC) migration to help other organizations strengthen their resilience as industry transitions to post-quantum cryptography standards. We’re proposing the idea of PQC Migration Levels to help teams within organizations manage the complexity of PQC migration for their various use cases. By outlining Meta’s approach to this work […

CWChris Wiltz·April 16, 2026Security
SRE & Ops — Escaping the Fork: How Meta Modernized WebRTC Across 50+ Use Cases

Escaping the Fork: How Meta Modernized WebRTC Across 50+ Use Cases

At Meta, WebRTC powers real-time audio and video across various platforms. But forking a large open-source project like WebRTC within our monorepo presents unique challenges – over time, an internal fork can drift behind upstream, cutting itself off from community upgrades. We’re sharing how we escaped this “forking trap” – from building a dual-stack architecture […]

CWChris Wiltz·April 9, 2026SRE & Ops
SRE & Ops — Trust But Canary: Configuration Safety at Scale

Trust But Canary: Configuration Safety at Scale

As AI increases developer speed and productivity it also increases the need for safeguards. On this episode of the Meta Tech Podcast, Pascal Hartig sits down with Ishwari and Joe from Meta’s Configurations team to discuss how Meta makes config rollouts safe at scale. Listen in to learn about canarying and progressive rollouts, the health checks […]

CWChris Wiltz·April 8, 2026SRE & Ops
SRE & Ops — How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

AI coding assistants are powerful but only as good as their understanding of your codebase. When we pointed AI agents at one of Meta’s large-scale data processing pipelines – spanning four repositories, three languages, and over 4,100 files – we quickly found that they weren’t making useful edits quickly enough. We fixed this by building […]

CWChris Wiltz·April 6, 2026SRE & Ops
SRE & Ops — KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

This is the second post in the Ranking Engineer Agent blog series exploring the autonomous AI capabilities accelerating Meta’s Ads Ranking innovation. The previous post introduced Ranking Engineer Agent’s ML exploration capability, which autonomously designs, executes, and analyzes ranking model experiments. This post covers how to optimize the low-level infrastructure that makes those models run

CWChris Wiltz·April 2, 2026SRE & Ops
AI & ML — Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

Meta Adaptive Ranking Model: Bending the Inference Scaling Curve to Serve LLM-Scale Models for Ads

Meta continues to lead the industry in utilizing groundbreaking AI Recommendation Systems (RecSys) to deliver better experiences for people, and better results for advertisers. To reach the next frontier of performance, we are scaling Meta’s Ads Recommender runtime models to LLM-scale & complexity to further a deeper understanding of people’s interests and intent. This increase […]

CWChris Wiltz·March 31, 2026AI & ML
SRE & Ops — AI for American-Produced Cement and Concrete

AI for American-Produced Cement and Concrete

Meta is continuing its long-term roadmap to help the construction industry leverage AI to produce high-quality and more sustainable concrete mixes, as well as those exclusively produced in the United States. Concurrent with the 2026 American Concrete Institute (ACI) Spring Convention, Meta is releasing a new AI model for designing concrete mixes – Bayesian Optimization […]

CWChris Wiltz·March 30, 2026SRE & Ops
AI & ML — Friend Bubbles: Enhancing Social Discovery on Facebook Reels

Friend Bubbles: Enhancing Social Discovery on Facebook Reels

Friend bubbles in Facebook Reels highlight Reels your friends have liked or reacted to, helping you discover new content and making it easier to connect over shared interests. This article explains the technical architecture behind friend bubbles, including how machine learning estimates relationship strength and ranks content your friends have interacted with to create more […]

CWChris Wiltz·March 18, 2026AI & ML
SRE & Ops — Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation

Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation

Meta’s Ranking Engineer Agent (REA) autonomously executes key steps across the end-to-end machine learning (ML) lifecycle for ads ranking models. This post covers REA’s ML experimentation capabilities: autonomously generating hypotheses, launching training jobs, debugging failures, and iterating on results. Future posts will cover additional REA capabilities. REA reduces the need for manual interv

CWChris Wiltz·March 17, 2026SRE & Ops
Frontend — Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps

Patch Me If You Can: AI Codemods for Secure-by-Default Android Apps

Even seemingly simple engineering tasks — like updating an API — can become monumental undertakings when you’re dealing with millions of lines of code and thousands of engineers, especially if the changes are security-related. Nowhere is this more apparent than in mobile security, where a single class of vulnerability can be replicated across hundreds of […]

CWChris Wiltz·March 13, 2026Frontend
Security — How Advanced Browsing Protection Works in Messenger

How Advanced Browsing Protection Works in Messenger

We’re sharing the technical details behind how Advanced Browsing Protection (ABP) in Messenger protects the privacy of the links clicked on within chats while still warning people about malicious links. We hope that this post has helped to illuminate some of the engineering challenges and infrastructure components involved for providing this feature for our users. […]

CWChris Wiltz·March 9, 2026Security
Distributed — FFmpeg at Meta: Media Processing at Scale

FFmpeg at Meta: Media Processing at Scale

FFmpeg is truly a multi-tool for media processing. As an industry-standard tool it supports a wide variety of audio and video codecs and container formats. It can also orchestrate complex chains of filters for media editing and manipulation. For the people who use our apps, FFmpeg plays an important role in enabling new video experiences […]

CWChris Wiltz·March 2, 2026Distributed
SRE & Ops — Investing in Infrastructure: Meta’s Renewed Commitment to jemalloc

Investing in Infrastructure: Meta’s Renewed Commitment to jemalloc

Meta recognizes the long-term benefits of jemalloc, a high-performance memory allocator, in its software infrastructure. We are renewing focus on jemalloc, aiming to reduce maintenance needs and modernize the codebase while continuing to evolve the allocator to adapt to the latest hardware and workloads. We are committed to continuing to develop jemalloc development with the […]

CWChris Wiltz·March 2, 2026SRE & Ops
AI & ML — RCCLX: Innovating GPU Communications on AMD Platforms

RCCLX: Innovating GPU Communications on AMD Platforms

We are open-sourcing the initial version of RCCLX – an enhanced version of RCCL that we developed and tested on Meta’s internal workloads. RCCLX is fully integrated with Torchcomms and aims to empower researchers and developers to accelerate innovation, regardless of their chosen backend. Communication patterns for AI models are constantly evolving, as are hardware […]

CWChris Wiltz·February 24, 2026AI & ML
SRE & Ops — Building Prometheus: How Backend Aggregation Enables Gigawatt-Scale AI Clusters

Building Prometheus: How Backend Aggregation Enables Gigawatt-Scale AI Clusters

We’re sharing details of the role backend aggregation (BAG) plays in building Meta’s gigawatt-scale AI clusters like Prometheus. BAG allows us to seamlessly connect thousands of GPUs across multiple data centers and regions. Our BAG implementation is connecting two different network fabrics – Disaggregated Schedule Fabric (DSF) and Non-Scheduled Fabric (NSF). Once it’s complete our AI […] Read Mor

CWChris Wiltz·February 9, 2026SRE & Ops
Security — No Display? No Problem: Cross-Device Passkey Authentication for XR Devices

No Display? No Problem: Cross-Device Passkey Authentication for XR Devices

We’re sharing a novel approach to enabling cross-device passkey authentication for devices with inaccessible displays (like XR devices). Our approach bypasses the use of QR codes and enables cross-device authentication without the need for an on-device display, while still complying with all trust and proximity requirements. This approach builds on work done by the FIDO […]

CWChris Wiltz·February 4, 2026Security
Security — Rust at Scale: An Added Layer of Security for WhatsApp

Rust at Scale: An Added Layer of Security for WhatsApp

WhatsApp has adopted and rolled out a new layer of security for users – built with Rust – as part of its effort to harden defenses against malware threats. WhatsApp’s experience creating and distributing our media consistency library in Rust to billions of devices and browsers proves Rust is production ready at a global scale. […]

CWChris Wiltz·January 27, 2026Security
AI & ML — Adapting the Facebook Reels RecSys AI Model Based on User Feedback

Adapting the Facebook Reels RecSys AI Model Based on User Feedback

We’ve improved personalized video recommendations on Facebook Reels by moving beyond metrics such as likes and watch time and directly leveraging user feedback. Our new User True Interest Survey (UTIS) model, now helps surface more niche, high-quality content and boosts engagement, retention, and satisfaction. We’re doubling down on personalization, tackling challenges like sparse user data […] Re

CWChris Wiltz·January 14, 2026AI & ML
Career & Teams — CSS at Scale With StyleX

CSS at Scale With StyleX

Build a large enough website with a large enough codebase, and you’ll eventually find that CSS presents challenges at scale. It’s no different at Meta, which is why we open-sourced StyleX, a solution for CSS at scale. StyleX combines the ergonomics of CSS-in-JS with the performance of static CSS. It allows atomic styling of components […]

CWChris Wiltz·January 12, 2026Career & Teams
Career & Teams — Python Typing Survey 2025: Code Quality and Flexibility As Top Reasons for Typing Adoption

Python Typing Survey 2025: Code Quality and Flexibility As Top Reasons for Typing Adoption

The 2025 Typed Python Survey, conducted by contributors from JetBrains, Meta, and the broader Python typing community, offers a comprehensive look at the current state of Python’s type system and developer tooling. With 1,241 responses (a 15% increase from last year), the survey captures the evolving sentiment, challenges, and opportunities around Python typing in the […]

CWChris Wiltz·December 22, 2025Career & Teams
SRE & Ops — DrP: Meta’s Root Cause Analysis Platform at Scale

DrP: Meta’s Root Cause Analysis Platform at Scale

Incident investigation can be a daunting task in today’s digital landscape, where large-scale systems comprise numerous interconnected components and dependencies DrP is a root cause analysis (RCA) platform, designed by Meta, to programmatically automate the investigation process, significantly reducing the mean time to resolve (MTTR) for incidents and alleviating on-call toil Today, DrP is used [

CWChris Wiltz·December 19, 2025SRE & Ops
Career & Teams — How We Built Meta Ray-Ban Display: From Zero to Polish

How We Built Meta Ray-Ban Display: From Zero to Polish

We’re going behind the scenes of the Meta Ray-Ban Display, Meta’s most advanced AI glasses yet. In a previous episode we met the team behind the Meta Neural Band, the EMG wristband packaged with the Ray-Ban Display. Now we’re delving into the glasses themselves. Kenan and Emanuel, from Meta’s Wearables org, join Pascal Hartig on […]

CWChris Wiltz·December 17, 2025Career & Teams
Frontend — How AI Is Transforming the Adoption of Secure-by-Default Mobile Frameworks

How AI Is Transforming the Adoption of Secure-by-Default Mobile Frameworks

Meta’s secure-by-default frameworks wrap potentially unsafe OS and third-party functions, making security the default while preserving developer speed and usability. These frameworks are designed to closely mirror existing APIs, rely on public and stable interfaces, and maximize developer adoption by minimizing friction and complexity. Generative AI and automation accelerate the adoption of secure

CWChris Wiltz·December 15, 2025Frontend
SRE & Ops — Zoomer: Powering AI Performance at Meta’s Scale Through Intelligent Debugging and Optimization

Zoomer: Powering AI Performance at Meta’s Scale Through Intelligent Debugging and Optimization

We’re introducing Zoomer, Meta’s comprehensive, automated debugging and optimization platform for AI. Zoomer works across all of our training and inference workloads at Meta and provides deep performance insights that enable energy savings, workflow acceleration, and efficiency gains in our AI infrastructure. Zoomer has delivered training time reductions, and significant QPS improvements, making i

CWChris Wiltz·November 21, 2025SRE & Ops
Security — Key Transparency Comes to Messenger

Key Transparency Comes to Messenger

We’re excited to share another advancement in the security of your conversations on Messenger: the launch of key transparency verification for end-to-end encrypted chats. This new feature enables an additional level of assurance that only you — and the people you’re communicating with — can see or listen to what is sent, and that no […]

CWChris Wiltz·November 20, 2025Security
AI & ML — Efficient Optimization With Ax, an Open Platform for Adaptive Experimentation

Efficient Optimization With Ax, an Open Platform for Adaptive Experimentation

We’ve released Ax 1.0, an open-source platform that uses machine learning to automatically guide complex, resource-intensive experimentation. Ax is used at scale across Meta to improve AI models, tune production infrastructure, and accelerate advances in ML and even hardware design. Our accompanying paper, “Ax: A Platform for Adaptive Experimentation” explains Ax’s architecture, methodology, and h

CWChris Wiltz·November 18, 2025AI & ML
Frontend — Enhancing HDR on Instagram for iOS With Dolby Vision

Enhancing HDR on Instagram for iOS With Dolby Vision

We’re sharing how we’ve enabled Dolby Vision and ambient viewing environment (amve) on the Instagram iOS app to enhance the video viewing experience. HDR videos created on iPhones contain unique Dolby Vision and amve metadata that we needed to support end-to-end Instagram for iOS is now the first Meta app to support Dolby Vision video, […]

CWChris Wiltz·November 17, 2025Frontend
SRE & Ops — Open Source Is Good for the Environment

Open Source Is Good for the Environment

Most people have heard of open-source software. But have you heard about open hardware? And did you know open source can have a positive impact on the environment? On this episode of the Meta Tech Podcast, Pascal Hartig sits down with Dharmesh and Lisa to talk about all things open hardware, and Meta’s biggest announcements […]

CWChris Wiltz·November 14, 2025SRE & Ops
Performance — StyleX: A Styling Library for CSS at Scale

StyleX: A Styling Library for CSS at Scale

StyleX is Meta’s styling system for large-scale applications. It combines the ergonomics of CSS-in-JS with the performance of static CSS, generating collision-free atomic CSS while allowing for expressive, type-safe style authoring. StyleX was open sourced at the end of 2023 and has since become the standard styling system across Meta products like Facebook, Instagram, WhatsApp, […]

CWChris Wiltz·November 11, 2025Performance
AI & ML — Meta’s Generative Ads Model (GEM): The Central Brain Accelerating Ads Recommendation AI Innovation

Meta’s Generative Ads Model (GEM): The Central Brain Accelerating Ads Recommendation AI Innovation

We’re sharing details about Meta’s Generative Ads Recommendation Model (GEM), a new foundation model that delivers increased ad performance and advertiser ROI by enhancing other ads recommendation models’ ability to serve relevant ads. GEM’s novel architecture allows it to scale with an increasing number of parameters while consistently generating more precise predictions efficiently. GEM propagat

CWChris Wiltz·November 10, 2025AI & ML
AI & ML — Video Invisible Watermarking at Scale

Video Invisible Watermarking at Scale

At Meta, we use invisible watermarking for a variety of content provenance use cases on our platforms. Invisible watermarking serves a number of use cases, including detecting AI-generated videos, verifying who posted a video first, and identifying the source and tools used to create a video. We’re sharing how we overcame the challenges of scaling […]

CWChris Wiltz·November 4, 2025AI & ML
Security — Scaling Privacy Infrastructure for GenAI Product Innovation

Scaling Privacy Infrastructure for GenAI Product Innovation

How does Meta empower its product teams to harness GenAI’s power responsibly? In this post, we delve into how Meta addresses the challenges of safeguarding data in the GenAI era by scaling its Privacy Aware Infrastructure (PAI), with a particular focus on Meta’s AI glasses as an example GenAI use case. We’ll describe in detail […]

CWChris Wiltz·October 23, 2025Security
SRE & Ops — Disaggregated Scheduled Fabric: Scaling Meta’s AI Journey

Disaggregated Scheduled Fabric: Scaling Meta’s AI Journey

Disaggregated Schedule Fabric (DSF) is Meta’s next-generation network fabric technology for AI training networks that addresses the challenges of existing Clos-based networks. We’re sharing the challenges and innovations surrounding DSF and discussing future directions, including the creation of mega clusters through DSF and non-DSF region interconnectivity, as well as the exploration of alternati

CWChris Wiltz·October 20, 2025SRE & Ops
SRE & Ops — Branching in a Sapling Monorepo

Branching in a Sapling Monorepo

Sapling is a scalable, user-friendly, and open-source source control system that powers Meta’s monorepo. As discussed at the GitMerge 2024 conference session on branching, designing and implementing branching workflows for large monorepos is a challenging problem with multiple tradeoffs between scalability and the developer experience. After the conference, we designed, implemented, and open sourc

CWChris Wiltz·October 16, 2025SRE & Ops
SRE & Ops — Design for Sustainability: New Design Principles for Reducing IT Hardware Emissions

Design for Sustainability: New Design Principles for Reducing IT Hardware Emissions

We’re presenting Design for Sustainability, a set of technical design principles for new designs of IT hardware to reduce emissions and cost through reuse, extending useful life, and optimizing design. At Meta, we’ve been able to significantly reduce the carbon footprint of our data centers by integrating several design strategies such as modularity, reuse, retrofitting, […]

CWChris Wiltz·October 14, 2025SRE & Ops
SRE & Ops — OCP Summit 2025: The Open Future of Networking Hardware for AI

OCP Summit 2025: The Open Future of Networking Hardware for AI

At Open Compute Project Summit (OCP) 2025, we’re sharing details about the direction of next-generation network fabrics for our AI training clusters. We’ve expanded our network hardware portfolio and are contributing new disaggregated network platforms to OCP. We look forward to continued collaboration with OCP to open designs for racks, servers, storage boxes, and motherboards […]

CWChris Wiltz·October 13, 2025SRE & Ops
Frontend — Introducing the React Foundation: The New Home for React & React Native

Introducing the React Foundation: The New Home for React & React Native

Meta open-sourced React over a decade ago to help developers build better user experiences. Since then, React has grown into one of the world’s most popular open source projects, powering over 50 million websites and products built by companies such as Microsoft, Shopify, Bloomberg, Discord, Coinbase, the NFL, and many others. With React Native, React […]

CWChris Wiltz·October 7, 2025Frontend
SRE & Ops — Introducing OpenZL: An Open Source Format-Aware Compression Framework

Introducing OpenZL: An Open Source Format-Aware Compression Framework

OpenZL is a new open source data compression framework that offers lossless compression for structured data. OpenZL is designed to offer the performance of a format-specific compressor with the easy maintenance of a single executable binary. You can get started with OpenZL today by visiting our Quick Start guide and the OpenZL GitHub repository. Learn more […]

CWChris Wiltz·October 6, 2025SRE & Ops