Build speed has been a long-running annoyance in Rust projects, and the usual suspects keep showing up in the critical path: syn and serde. Causal profiling in The virtue of unsynn previously showed that speeding up syn would speed up builds, and there is a good chance your own dependency tree contains it. The CMS home depends on syn 1 through 6 different paths, and on syn 2 through 25 different paths. Two versions of thiserror, clap, async-trait, displaydoc, various futures macros, perfect hash maps, tokio macros, tracing, zerovec, zeroize, zerofrom, yoke, and serde all sit on top of that. Replacing some of those is imaginable; serde is not so easy. As of May 2025, syn is the most downloaded crate ever at 900 million downloads, and serde is close behind at 540 million downloads.

Why splitting the crate doesn't help

The reflex response to a slow-building crate is to break it into multiple crates. Under serde's design that makes little difference, and monomorphization is the reason.

Take a bigapi-types crate that houses a heap of types for an API and its JSON payloads, and a catalog that keeps growing:

use chrono::{NaiveDate, NaiveDateTime}; use serde::{Deserialize, Serialize}; use uuid::Uuid; /// The root struct representing the catalog of everything. #[derive(Serialize, Deserialize, Debug, Clone)] pub struct Catalog { pub id: Uuid, pub businesses: Vec<Business>, pub created_at: NaiveDateTime, pub metadata: CatalogMetadata, }
#[derive(Serialize, Deserialize, Debug, Clone)] pub struct CatalogMetadata { pub version: String, pub region: String, }
/// A business represented in the catalog. #[derive(Serialize, Deserialize, Debug, Clone)] pub struct Business { pub id: Uuid, pub name: String, pub address: Address, pub owner: BusinessOwner, pub users: Vec<BusinessUser>, pub branches: Vec<Branch>, pub products: Vec<Product>, pub created_at: NaiveDateTime, }

For narration purposes, a bigapi-indirection crate holds this:

use bigapi_types::generate_mock_catalog; pub fn do_ser_stuff() { // Generate a mock catalog let catalog = generate_mock_catalog(); // Serialize the catalog to JSON let serialized = serde_json::to_string_pretty(&catalog).expect("Failed to serialize catalog!"); println!("Serialized catalog JSON:\n{}", serialized); // Deserialize back to a Catalog struct let deserialized: bigapi_types::Catalog = serde_json::from_str(&serialized).expect("Failed to deserialize catalog"); println!("Deserialized catalog struct!\n{:#?}", deserialized); }

And an application, bigapi-cli, merely calls do_ser_stuff:

fn main() { println!("About to do ser stuff..."); bigapi_indirection::do_ser_stuff(); println!("About to do ser stuff... done!"); }

Going purely by code volume, the CLI should be quick to build, indirection should be quick too since it's a couple of calls, and bigapi-types should be slow with all those struct definitions and a function generating a mock catalog.

On a cold debug build that intuition holds. On a cold release build it very much does not: indirection takes the bulk of the build time.

The cause is that serde_json::to_string_pretty and serde_json::from_str are generic functions, and the instantiations land in the bigapi-indirection crate. Every touch of that crate, even changing a string constant, pays the cost again. Touching bigapi-types is worse still: changing only a string value in generate_mock_catalog triggers a full rebuild.

What monomorphization actually costs

Rust instantiates every generic function: the generic parameters T, K, V are substituted with concrete types. That is monomorphization.

cargo-llvm-lines shows how often it happens:

bigapi on  main [+] via 🦀 v1.87.0 ❯ cargo llvm-lines --release -p bigapi-indirection|head -15 Compiling bigapi-indirection v0.1.0 (/Users/amos/bearcove/bigapi/bigapi-indirection) Finished `release` profile [optimized] target(s) in 0.71s Lines Copies Function name ----- ------ ------------- 80335 1542 (TOTAL) 8760 (10.9%, 10.9%) 20 (1.3%, 1.3%) <&mut serde_json::de::Deserializer<R> as serde::de::Deserializer>::deserialize_struct 3674 (4.6%, 15.5%) 45 (2.9%, 4.2%) <serde_json::de::SeqAccess<R> as serde::de::SeqAccess>::next_element_seed 3009 (3.7%, 19.2%) 11 (0.7%, 4.9%) <&mut serde_json::de::Deserializer<R> as serde::de::Deserializer>::deserialize_seq 2553 (3.2%, 22.4%) 37 (2.4%, 7.3%) <serde_json::ser::Compound<W,F> as serde::ser::SerializeMap>::serialize_value 1771 (2.2%, 24.6%) 38 (2.5%, 9.8%) <serde_json::de::MapAccess<R> as serde::de::MapAccess>::next_value_seed 1680 (2.1%, 26.7%) 20 (1.3%, 11.1%) <serde_json::de::MapAccess<R> as serde::de::MapAccess>::next_key_seed 1679 (2.1%, 28.8%) 1 (0.1%, 11.2%) <bigapi_types::_::<impl serde::de::Deserialize for bigapi_types::Product>::deserialize::__Visitor as serde::de::Visitor>::visit_map 1569 (2.0%, 30.7%) 1 (0.1%, 11.2%) <bigapi_types::_::<impl serde::de::Deserialize for bigapi_types::Business>::deserialize::__Visitor as serde::de::Visitor>::visit_map 1490 (1.9%, 32.6%) 10 (0.6%, 11.9%) serde::ser::Serializer::collect_seq 1316 (1.6%, 34.2%) 1 (0.1%, 11.9%) <bigapi_types::_::<impl serde::de::Deserialize for bigapi_types::User>::deserialize::__Visitor as serde::de::Visitor>::visit_map 1302 (1.6%, 35.9%) 1 (0.1%, 12.0%) <bigapi_types::_::<impl serde::de::Deserialize for bigapi_types::UserProfile>::deserialize::__Visitor as serde::de::Visitor>::visit_map 1300 (1.6%, 37.5%) 20 (1.3%, 13.3%) <serde_json::de::MapKey<R> as serde::de::Deserializer>::deserialize_any
Cool bear
Cool Bear's hot tip: Omitting --release gives slightly different results — LLVM is not the only one doing optimizations!

Roughly 40 copies of a range of generic serde methods appear, each specialized for the given types. That is what makes serde fast, and it is also what makes the build slow — and the resulting binary a little plus-sized:

bigapi on  main [+] via 🦀 v1.87.0 ❯ cargo build --release Finished `release` profile [optimized] target(s) in 0.01s bigapi on  main [+] via 🦀 v1.87.0 ❯ ls -lhA target/release/bigapi-cli Permissions Size User Date Modified Name .rwxr-xr-x 884k amos 30 May 21:16 target/release/bigapi-cli

This is fundamental to how serde works. miniserde, from the same author, works differently, but testing it isn't practical here: neither uuid nor chrono has a miniserde feature, and forking it isn't worth the bother.

Measuring facet against serde

The author of facet did not set out to build a faster serde, reasoning that a second serializer-deserializer framework would need to be dramatically better to justify a migration — serde being both excellent and pervasive. Instead, facet is meant to win on different characteristics entirely.

That premise is now being tested openly. Taking a program that uses serde and forking it to use facet instead means swapping out its dependencies:

/// The root struct representing the catalog of everything. #[derive(Serialize, Deserialize, Debug, Clone)] pub struct Catalog { pub id: Uuid, pub businesses: Vec<Business>, pub created_at: NaiveDateTime, pub metadata: CatalogMetadata, }

for these:

/// The root struct representing the catalog of everything. #[derive(Facet, Clone)] pub struct Catalog { pub id: Uuid, pub businesses: Vec<Business>, pub created_at: NaiveDateTime, pub metadata: CatalogMetadata, }

The indirect crate then relies on facet-json for JSON handling and facet-pretty in place of Debug:

use bigapi_types_facet::generate_mock_catalog; use facet_pretty::FacetPretty; pub fn do_ser_stuff() { // Generate a mock catalog let catalog = generate_mock_catalog(); // Serialize the catalog to JSON let serialized = facet_json::to_string(&catalog); println!("Serialized catalog JSON.\n{}", serialized); // Deserialize back to a Catalog struct let deserialized: bigapi_types_facet::Catalog = facet_json::from_str(&serialized).expect("Failed to deserialize catalog!"); println!("Deserialized catalog struct:\n{}", deserialized.pretty()); }

With a CLI layered on top of that indirect crate, the facade version can be compared directly against the older, serde-powered build. The results are reported as they are, not as one might hope — the author treats the frustration they generate as fuel for further work. There is one bright spot: as of the day before that writing, facet beat serde-json on a single benchmark, serializing a 100 kilobyte string in 351 microseconds on the machine CodSpeed uses.

The serialized long string benchmark takes 351.9 microseconds for facet-json and 460.9 microseconds for serde.

Where the binary goes

The broader picture in the sample program is less flattering: the facet variant produces a larger executable than the serde version. Finding the culprit is harder this time around. cargo-bloat on the serde build shows plainly where the bytes are spent; on the facet build, std is the largest contributor, followed closely by the types crate, facet_deserialize, the indirection crate, facet_json, facet_core and others.

bigapi on  main via 🦀 v1.87.0 ❯ cargo bloat --crates -p bigapi-cli Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.01s Analyzing target/debug/bigapi-cli File .text Size Crate 17.0% 41.6% 351.9KiB bigapi_indirection 13.3% 32.4% 273.9KiB std 3.5% 8.5% 72.2KiB chrono 2.2% 5.3% 44.8KiB serde_json 2.1% 5.2% 44.3KiB bigapi_types ✂️ Note: numbers above are a result of guesswork. They are not 100% correct and never will be.
bigapi on  main via 🦀 v1.87.0 ❯ cargo bloat --crates -p bigapi-cli-facet Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.01s Analyzing target/debug/bigapi-cli-facet File .text Size Crate 6.3% 20.7% 326.3KiB std 5.9% 19.4% 305.5KiB bigapi_types_facet 3.8% 12.7% 200.0KiB facet_deserialize 3.8% 12.6% 198.1KiB bigapi_indirection_facet 2.8% 9.4% 147.9KiB facet_json 2.6% 8.7% 136.5KiB facet_core 2.2% 7.1% 112.3KiB chrono 1.4% 4.8% 75.0KiB facet_reflect 0.4% 1.3% 21.1KiB facet_pretty ✂️ Note: numbers above are a result of guesswork. They are not 100% correct and never will be.

One mitigating factor: the code ends up spread fairly evenly across crates rather than concentrated.

Concurrency in the build graph

Cold and warm builds tell different stories. On a cold debug build, bigapi-types takes longer but does not block the rest from compiling.

(JavaScript is required for this)

On a cold release build, facet-deserialize, pretty, serialize and json all build concurrently, and any crate depending on the indirection crate can build alongside them as well — the purple coloring in the profile makes that visible. Warm release builds are near parity between the serde and facet versions: modifying a bit of bigapi-types-serde and modifying a bit of bigapi-types-facet take approximately the same time. Pinning everything to a single job with -j1 makes matters worse, and that setting is applied to both sides rather than only to the more flattering one.

(JavaScript is required for this)

(JavaScript is required for this)

(JavaScript is required for this)

(JavaScript is required for this)

There is reason for optimism on this front. Facet added a number of marker traits and reimplements standard traits for tuples whose elements implement them, none of which is free.

bigapi on  main [!+?⇡] via 🦀 v1.87.0 ❯ cargo llvm-lines --release -p bigapi-indirection-facet|head -10 Compiling bigapi-indirection-facet v0.1.0 (/Users/amos/bearcove/bigapi/bigapi-indirection-facet) Finished `release` profile [optimized] target(s) in 1.29s Lines Copies Function name ----- ------ ------------- 129037 4066 (TOTAL) 33063 (25.6%, 25.6%) 1509 (37.1%, 37.1%) core::ops::function::FnOnce::call_once 8247 (6.4%, 32.0%) 3 (0.1%, 37.2%) facet_deserialize::StackRunner<C,I>::set_numeric_value 6218 (4.8%, 36.8%) 1 (0.0%, 37.2%) facet_pretty::printer::PrettyPrinter::format_peek_internal 5279 (4.1%, 40.9%) 1 (0.0%, 37.2%) facet_deserialize::StackRunner<C,I>::pop 5010 (3.9%, 44.8%) 50 (1.2%, 38.5%) facet_core::impls_alloc::vec::<impl facet_core::Facet for alloc::vec::Vec<T>>::VTABLE::{{constant}}::{{closure}}::{{closure}} 3395 (2.6%, 47.4%) 1 (0.0%, 38.5%) facet_deserialize::StackRunner<C,I>::object_key_or_object_close 2803 (2.2%, 49.6%) 1 (0.0%, 38.5%) facet_deserialize::StackRunner<C,I>::value

Two numbers stand out in cargo-llvm-lines: call_once accounts for 33 thousand lines of LLVM IR, and set_numeric_value — essentially converting u64s to u16s and back — accounts for over 6% of the total code. No investigation has been done yet; for now these serve as a baseline.

Data instead of code, and what it buys

In exchange for larger binaries and longer builds, facet changes the shape of what the derive macro emits. It generates data rather than code, and along with that a large number of virtual tables enabling runtime interaction with arbitrary values. The data-driven design shows up in practical ways. The Debug implementation is gone from the printing path; facet-pretty renders deserialized values instead:

bigapi on  main via 🦀 v1.87.0 ❯ cargo run -p bigapi-cli-facet Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.01s Running `target/debug/bigapi-cli-facet` About to do ser stuff... Serialized catalog JSON. ✂️ Deserialized catalog struct: /// The root struct representing the catalog of everything. Catalog { id: aa1238fa-8f72-45fa-b5a7-34d99baf4863, businesses: Vec<Business> [ /// A business represented in the catalog. Business { id: 65d08ea7-53c6-42e8-848e-0749d00b7bdd, name: Awesome Business, address: Address { street: 123 Main St., city: Metropolis, state: Stateville, postal_code: 12345, country: Countryland, geo: Option<GeoLocation>::Some(GeoLocation { latitude: 51, longitude: -0.1, }), }, owner: BusinessOwner { user: User { id: 056b3eda-97ca-4c12-883d-ecc043a6f5b4,

That is a one-time cost that yields colored formatting, and it also supports redaction: values marked sensitive, such as street numbers, can be kept out of logs.

#[derive(Facet, Clone)] pub struct Address { // 👇 #[facet(sensitive)] pub street: String, pub city: String, pub state: String, pub postal_code: String, pub country: String, pub geo: Option<GeoLocation>, }
bigapi on  main [!] via 🦀 v1.87.0 ❯ cargo run -p bigapi-cli-facet ✂️ Deserialized catalog struct: /// The root struct representing the catalog of everything. Catalog { id: 61f70016-eca4-45af-8937-42c03f9a5cd8, businesses: Vec<Business> [ /// A business represented in the catalog. Business { id: 9b52c85b-9240-4e73-9553-5d827e36b5f5, name: Awesome Business, address: Address { street: [REDACTED], city: Metropolis, state: Stateville, postal_code: 12345, country: Countryland,

Colors can be turned off, and because facet-pretty works from data rather than code, the depth of what it prints can be limited — something the Debug trait is not flexible enough to allow. Applied together, that data powers both presentation and serialization: the same struct exposure lets facet-json read and write the JSON representation, while a colored, layered view of the struct is printed to the terminal.

Speed, with caveats

Speed remains serde's territory. At the time of writing, facet-json runs anywhere from 3 to 6 times slower than serde-json, with the live benchmarks available for exact figures.

Plotted on a log scale, the gap looks far less dramatic.

A more rigorous or automated comparison was not possible within the deadline, so the microbenchmarks should be taken with a grain of salt. From the standpoint of an end user, both feel instantaneous.

amos in 🌐 trollop in bigapi on  main via 🦀 v1.87.0 took 6s ❯ hyperfine -N target/release/bigapi-cli-serde target/release/bigapi-cli-facet --warmup 500 Benchmark 1: target/release/bigapi-cli-serde Time (mean ± σ): 3.4 ms ± 1.7 ms [User: 2.3 ms, System: 0.9 ms] Range (min … max): 1.8 ms … 10.0 ms 1623 runs Benchmark 2: target/release/bigapi-cli-facet Time (mean ± σ): 4.0 ms ± 1.9 ms [User: 2.5 ms, System: 1.4 ms] Range (min … max): 1.8 ms … 13.7 ms 567 runs Summary target/release/bigapi-cli-serde ran 1.18 ± 0.82 times faster than target/release/bigapi-cli-facet

What the JSON side already buys you

Reflection pays off twice. First, since the shape data is generated once, two different facet-json backends — one multi-threaded, one SIMD-optimized — can parse into the exact same thing and remain interchangeable.

Second, there is the stack. serde-json is a recursive descent parser, but the recursion lives in your code: hand it a struct deep enough with a descent recursion, and it will happily overflow. Given a large struct and deeply nested JSON input, serde-json blows the stack in debug builds, and the same thing happens in release with more padding added.

facet-json instead parses iteratively, gathering tokens and deferring their interpretation — and that design means you must write your own deserializer. Keep in mind the defaults: field defaults are not free with serde, but with facet they are explicitly set through an attribute such as #[facet(default)] (there is no implicit behavior for Option at this time), and there is no initial callback. Because of the iterative shape, if you need a Serialize/Deserialize equivalent, you must write a Facet implementation by hand — given facet's derive, this is usually less work than serde's alternatives. What that yields is worth it:

as an understatement, it works fast.

On a benchmark where the serde build stack-overflows, the facet build currently runs in 118ms vs 112ms (an oversimplification: the serde version never got to the actual parsing), the p99 is 3ms.

The roadmap

At the language level, we have current problems: typeid is not, and possibly never will be, const, and have no plan to be. Rather than allocating a boxed error for every failure, we handle errors via Extensions (not at runtime)?, so we have to use boxed errors.

On the reflection side, facet-reflect uses cursors rather than the closures that the old (unsafe) code had, letting us get rid of a pile of sync. We're working on making the types as transparent as possible, reducing the friction for integrations deep in your code.

Looking at ecosystem work, the gap is not JSON — it's every other format that nobody has benchmarked. Functionality on that front:

  • At the v1 level, our current hypothesis is that we'll finally have something like in reflection:
#[non_exhaustive] #[repr(C)] pub struct StructType<'shape> { pub repr: Repr, pub kind: StructKind, pub fields: &'shape [Field<'shape>], }

We are already used nothing: built-in format flags — xml name, output prefix and nesting spaces in xml — are not available in Rust and there are no plans yet. This means these settings still force format writers to the hand-written API.

As for next steps, most features behind facet aren't realistic as extensibility work until the crate stabilizes. Let's not over-index on building everything at once. if there's value. There's even a talk on that by speed, which many have taken issue with.

Recent integrations: don't reach benchmark coverage or are "private-ish".