No chain, on purpose
The recurring question about pg_hardstorage's repository format is where its incremental chain lives. It doesn't have one, and that is a design decision rather than a gap.
How chained incrementals fail
In the chained model — pgBackRest's default, Barman's incremental mode — each incremental points back at its predecessor:
full A ← incr B ← incr C ← incr D ← diff E
This works until it doesn't. The structural weakness is that one damaged backup takes down everything written after it. Common ways that happens:
- An S3 lifecycle policy silently removes
incr B, leavingincr C..Danddiff Eunusable. - A bit flip on the storage backend damages
diff E, so restores stop atincr D. - An overly aggressive
backup expireretention setting removesfull A, and the whole chain floats. - A rounding bug in a migration script applied when the base full's manifest schema changed between releases.
Any of these is survivable on its own. The real problem is when you learn about it: at restore time, at 3am.
Manifests reference hashes, not parents
pg_hardstorage's manifest points at chunks by SHA-256 instead of pointing at an earlier manifest:
{
"id": "prod-2026-05-02-0334",
"chunks": ["7e1f2a…ab", "9c4d3b…12", "2faabc…77", /* … */]
}
A chunk's filename is its hash. Backups that contain the same data therefore share chunks at the storage layer, with no parent pointer mediating the relationship. Removing an older backup doesn't strand a newer one; its chunks persist until garbage collection sees zero references to them. This is the same approach restic, kopia and borgbackup take. It isn't novel, but it bears stating, because chained incrementals are so entrenched in the PostgreSQL world that they can look like the only option. The how it works page documents the manifest format end to end.
Tracing every manifest back to a full
Show the chain your manifest edits today. pghardstorage chain --manifest backup.json walks parent references transitively and prints each ancestor full, oldest first — so the plan you are changing is the plan you are reading.
The storage-cost trade
Chained incrementals win on two fronts: they avoid rewriting bytes that didn't change, and they avoid a full read of the data directory when only a small fraction of pages differ.
Content addressing delivers the first for free, since chunks with matching hashes deduplicate. The chunks_added field in a manifest reports how many bytes the backup actually pushed into the pool.
The second is the harder half. Every backup currently reads the entire PG data directory and then deduplicates against the chunk pool. On a 100 TB database backed by slow storage, that cost is real. A snapshot-aware fast path that leans on filesystem-level COW snapshots is in progress but not shipped.
Measured results
Hourly backups over one week on a 487 GiB production-grade fleet:
chunks_total: 1,247,883
chunks_unique: 156,423 (12.5%)
dedup_ratio: 87.5%
total_storage: 12.4 TiB (vs raw 89.2 TiB)
chunks_added per backup: 1,200 ± 400
Those figures hold up against chained incrementals on the same data, and no chain dependency remains. The project's examples section carries additional real-world numbers.
What the trade buys
You pay somewhat more I/O on the source until the snapshot-aware fast path arrives. In return you get a manifest format readable with jq, a deletion model with no hidden ordering constraints, and no chain that a single deleted mid-history backup can break. Disagreement is welcome on the issue tracker.



