Performance’s Hardware Horizon: Beyond BPF
Brendan Gregg’s recent plenary at USENIX LISA 2021 turned the spotlight away from the BPF observability work he is best known for, and toward the broader hardware and software trends that will define server performance in the coming years. The talk, “Computing Performance: On the Horizon,” covers developments across processors, memory, storage, and networking, and offers a set of forward-looking predictions from the perspective of a senior performance engineer.
The full session is available on YouTube, with slides published online and as a PDF. The material also expands on chapters from the second edition of Gregg’s Systems Performance book.
What’s Changing at the Silicon Level
The talk leads with the CPU. New processor designs are increasingly leveraging 3D die stacking to overcome the limits of traditional 2D scaling. Cloud vendors are also entering the silicon fray with in-house designs like the AWS Graviton2, which have shifted the performance conversation from raw clock speed to core count, memory bandwidth, and power efficiency.
Close behind the CPU are memory advances. DDR5 is beginning deployment, offering higher bandwidth and capacity per channel. More transformative, however, is the arrival of High Bandwidth Memory (HBM) integrated on-package. HBM provides a generational jump in memory bandwidth, which has become the binding constraint for many modern workloads, particularly in analytics and machine learning inference where data movement outpaces compute.
The Storage and Network Landscape
On the storage side, Gregg highlights new uses for 3D XPoint as an accelerator for 3D NAND, rather than as a permanent replacement for it. In this arrangement, the faster persistent memory acts as a caching or metadata layer in front of the dense but slower NAND. This is a practical evolution for workloads that need persistence without paying a full latency penalty.
Networking also earns a mention, driven by the rise of QUIC as a transport protocol and eXpress Data Path (XDP) as a programmable kernel networking hook. XDP, in particular, moves packet processing earlier in the stack, enabling high-throughput applications to run at near line rate with lower overhead than traditional socket-based paths.
A Prediction, With a Caveat
Gregg’s talk is not shy about looking ahead, and he notes that these predictions are intended to provoke thought rather than to serve as unassailable forecasts. The central theme for the server side is a move away from squeezing single-thread performance and toward orchestrating parallelism, memory tiers, and fast I/O stacks in a way that systems software can reliably exploit.
The full walkthrough of these trends, along with the reasoning and data behind them, is well worth watching for engineers involved in capacity planning, kernel development, or infrastructure design.
Digging Deeper: Reference Material and Further Reading
For those who want to explore the topics covered in Brendan Gregg's LISA 2021 presentation, the talk's reference list is a valuable resource. It spans over a decade of pioneering work across hardware, software, and networking. Below is a selection of the cited sources, which are a mix of blog posts, academic papers, and industry announcements that trace the evolution of the technologies discussed.
The list includes foundational work on analysis methodologies, such as Gregg’s own writing on visualization techniques and flame graphs. On the systems software side, sources cover the introduction of new I/O schedulers in the Linux kernel, the development of advanced congestion control algorithms for TCP, and the tooling built with BPF. Hardware developments are also well documented, covering the roadmap from HBM and DDR4 through to DDR5, the introduction of high-capacity storage technologies, and the rise of custom silicon for the cloud, including FPGA-based infrastructure and area-specific processors.
Systems Software and Analysis
- [Gregg 08] Brendan Gregg, “ZFS L2ARC,” http://www.brendangregg.com/blog/2008-07-22/zfs-l2arc.html, Jul 2008
- [Corbet 17] Jonathan Corbet, “Two new block I/O schedulers for 4.12,” https://lwn.net/Articles/720675, Apr 2017
- [Gregg 13] Brendan Gregg, “Blazing Performance with Flame Graphs,” https://www.usenix.org/conference/lisa13/technical-sessions/plenary/gregg, 2013
- [Gregg 16b] Brendan Gregg, “Linux 4.X Tracing Tools: Using BPF Superpowers,” https://www.usenix.org/conference/lisa16/conference-program/presentation/linux-4x-tracing-tools-using-bpf-superpowers, 2016
- [Spier 20] Martin Spier, Brendan Gregg, et al., “FlameScope,” https://github.com/Netflix/flamescope, 2020
Networking and Protocol Evolution
- [Borkmann 14] Daniel Borkmann, “net: tcp: add DCTCP congestion control algorithm,” https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=e3118e8359bb7c59555aca60c725106e6d78c5ce, 2014
- [Cardwell 16] Neal Cardwell, et al., “BBR: Congestion-Based Congestion Control,” https://queue.acm.org/detail.cfm?id=3022184, 2016
- [Gallatin 19] Drew Gallatin, “Kernel TLS and hardware TLS offload in FreeBSD 13,” https://people.freebsd.org/~gallatin/talks/euro2019-ktls.pdf, 2019
- [Ford 20] A. Ford, et al., “TCP Extensions for Multipath Operation with Multiple Addresses,” https://datatracker.ietf.org/doc/html/rfc8684, Mar 2020
Memory, Processors, and Accelerators
- [Greenberg 11] Marc Greenberg, “DDR4: Double the speed, double the latency? Make sure your system can handle next-generation DRAM,” https://www.chipestimate.com/DDR4-Double-the-speed-double-the-latencyMake-sure-your-system-can-handle-next-generation-DRAM/Cadence/Technical-Article/2011/11/22, Nov 2011
- [Macri 15] Joe Macri, “Introducing HBM,” https://www.amd.com/en/technologies/hbm, Jul 2015
- [Cutress 20] Dr. Ian Cutress, “Insights into DDR5 Sub-timings and Latencies,” https://www.anandtech.com/show/16143/insights-into-ddr5-subtimings-and-latencies, Oct 2020
- [Alcorn 17b] Paul Alcorn, “Hot Chips 2017: Intel Deep Dives Into EMIB,” https://www.tomshardware.com/news/intel-emib-interconnect-fpga-chiplet,35316.html#xenforo-comments-3112212, 2017
- [Shilov 21d] Anton Shilov, “Sapphire Rapids Uncovered: 56 Cores, 64GB HBM2E, Multi-Chip Design,” https://www.tomshardware.com/news/intel-sapphire-rapids-xeon-scalable-specifications-and-features, Apr 2021
- [Vahdat 21] Amin Vahdat, “The past, present and future of custom compute at Google,” https://cloud.google.com/blog/topics/systems/the-past-present-and-future-of-custom-compute-at-google, Mar 2021
Storage Evolution
- [Shimpi 13] Anand Lal Shimpi, “Seagate to Ship 5TB HDD in 2014 using Shingled Magnetic Recording,” https://www.anandtech.com/show/7290/seagate-to-ship-5tb-hdd-in-2014-using-shingled-magnetic-recording, Sep 2013
- [Alcorn 17] Paul Alcorn, “Seagate To Double HDD Speed With Multi-Actuator Technology,” https://www.tomshardware.com/news/hdd-multi-actuator-heads-seagate,36132.html, 2017
- [Shilov 21] Anton Shilov, “Samsung Develops 512GB DDR5 Module with HKMG DDR5 Chips,” https://www.tomshardware.com/news/samsung-512gb-ddr5-memory-module, Mar 2021
- [Shilov 21b] Anton Shilov, “Seagate Ships 20TB HAMR HDDs Commercially, Increases Shipments of Mach.2 Drives,” https://www.tomshardware.com/news/seagate-ships-hamr-hdds-increases-dual-actuator-shipments, 2021
- [ZonedStorage 21] Zoned Storage, “Zoned Namespaces (ZNS) SSDs,” https://zonedstorage.io/introduction/zns, 2021
Gregg notes that he made a point to include author names alongside titles and dates for all sources, rather than simply listing URLs. This practice, which he adopted after feeling it was unfair in his earlier books that some references had author names and others didn't, provides better attribution for the reader. For those interested in his other work from the same conference, he also presented a talk on BPF Internals.



