What “Stack Machine” Really Means for WebAssembly

A recent blog post argues that WebAssembly (WASM) isn’t a genuine stack machine, pointing to its use of locals and the absence of stack manipulation primitives like dup or swap. It’s an interesting position, but the argument hinges on a term that lacks a rigorous, formal definition. Wikipedia’s working description—a processor or process VM whose primary interaction is moving short-lived temporary values to and from a push-down stack—actually fits WASM quite well.

WASM’s core operational model is stack-based; the stack is where the bulk of value passing happens. What sets it apart from purist stack machines like Forth is an additional, infinite set of named locals. In a language like Forth, all data manipulation must occur on the stack, with memory pointers also managed there. WASM provides that same stack and memory access, but augments it with these register-like locals, which changes how you write and read code.

Why Locals Beat Stack Gymnastics

The dup and swap operations mentioned in the critique are staples of stack-oriented programming. They’re necessary in Forth to rearrange data on the stack, but they come at a readability cost. Consider a fundamental Forth routine that adds a value to a memory location. The code often requires intricate stack choreography just to get operands in the right order, making it difficult to follow without a detailed comment trail showing the stack’s state at each step.

Code performing the same operation in WASM is markedly easier to parse. Even in its non-folded, linear form, it remains highly legible. The reason is simple: while calculations and commands use the stack, the key operands—the memory address and the value to add—are stored in named, typed locals. This eliminates the need for tuck-and-swap contortions, letting the reader focus on the logic rather than the data shuffling.

The dup Performance Question

One could argue that a repeated local.get $addr is less efficient than a proper dup instruction. This concern vanishes under scrutiny for two reasons. First, readability is better served by the named local approach, as shown above. Second, and more critically, performance is a non-issue. WASM is an abstraction layer; the underlying hardware that ultimately executes the code is a register machine. The translation from WASM’s stack model to native registers is where the real performance decisions are made.

Compiler toolchains built for modern languages are highly adept at this translation. They are designed to handle arbitrary control flow and register allocation, a legacy of their evolution from compiling C and similar languages. When wasmtime compiles the add_to_byte function to native code (at its default opt-level=2), the output is essentially identical to what you’d expect from hand-written x86-64 assembly for the C statement mem[addr] += addend.

The compiler easily recognizes that two consecutive reads from the same WASM local will yield the same value and performs redundant load elimination. This optimization is straightforward in the WASM model because aliasing locals is impossible; in the absence of an intervening write to that specific local, multiple reads are guaranteed to produce identical results. In contrast, such analysis is considerably harder with C’s unrestricted memory access and pointer aliasing.

Ultimately, the debate over whether WASM is a "pure" stack machine misses the point. The stack is its primary interaction model, but the addition of locals doesn’t detract from that identity—it enhances the model’s usability for humans without imposing any cost on the machine code that’s eventually generated.