Content Deduplication
S4 automatically deduplicates stored data using content-addressable storage (CAS). When two or more objects have identical content, only one copy is stored on disk.
How It Works
- When an object is uploaded, S4 computes its SHA-256 hash
- The hash is checked against the deduplication index
- If the hash already exists, the object data is not written again — only a new metadata reference is created
- A reference counter tracks how many objects point to each unique blob
- When all references to a blob are deleted, the space becomes reclaimable by the compactor
Storage Savings
Deduplication provides 30-50% space savings on typical workloads. Savings depend on how much duplicate content exists in your data.
Examples of high-dedup workloads: - Backup storage (incremental backups share most data) - Container image registries (layers are often shared) - Log storage (repeated patterns) - File sync services (multiple users uploading the same files)
Dedup Statistics
Check deduplication effectiveness via the stats API:
curl http://localhost:9000/api/stats
Response includes:
{
"dedup_unique_blobs": 1500,
"dedup_total_references": 3200,
"dedup_ratio": 2.13
}
| Field | Description |
|---|---|
dedup_unique_blobs |
Number of unique data blobs on disk |
dedup_total_references |
Total number of object references |
dedup_ratio |
Ratio of references to unique blobs (higher = more savings) |
A ratio of 2.13 means that on average, each unique blob is referenced by 2.13 objects — roughly 53% storage savings.
Interaction with Other Features
- Versioning: Each version is deduplicated independently. If version 1 and version 3 have identical content, only one copy is stored.
- Object Lock: Deduplication is transparent to Object Lock. Locked objects are protected regardless of dedup status.
- Lifecycle Policies: When lifecycle rules delete objects, dedup reference counts are decremented. The actual data is removed only when no references remain.
Repairs are never deduplicated
Deduplication answers a write with "those bytes are already here". That is the right answer for a client storing a file twice, and the wrong one for the write that exists because the bytes on this disk are damaged — a repair rebuilding an erasure-coded shard after bit rot, for example. Pointing the repaired record at the blob that already holds that content would point it back at the damage.
Such a write therefore carries a reserved metadata key,
_s4_force_physical_write, and the engine treats it as a healing write:
- the bytes go into a volume of their own instead of reusing an existing blob;
- the dedup entry for that content is moved to the copy just written, so the next writer of the same content lands on the healthy one;
- a replica applies it even if it has already applied a write with the same operation id, because a heal is not a retry (decision EC2-164).
The key describes the write, not the object: it is stripped before the record is persisted and never appears in object metadata. Ordinary writes are unaffected and keep the savings above.
A damaged blob leaves the index
Deduplication saves space by having many objects point at one copy of the bytes. That is also the one thing that must never happen to a copy the disk has damaged: a client storing a file whose content happens to match a rotted blob would get an object that was corrupt before it was ever read.
So the moment a read finds that a blob does not hash to what its record claims — bit rot, in one word — that blob's dedup entry loses its location. The entry itself stays, because the objects that point at those bytes still exist and still have to be able to release them; what goes is the engine's willingness to hand those bytes to anybody new. The next write of that content stores a copy of its own and the entry moves onto it, so the pool goes back to sharing one healthy copy.
Nothing about this slows an ordinary write down: a blob nobody has found damaged is shared exactly as before. And the object that was damaged is not forgotten — on a cluster the scrubber fetches a healthy copy from a replica and writes it back (see Bit rot below).
Bit rot
A blob can rot on disk without anything announcing it: the index still names the object, the size and metadata are right, and only the bytes have changed. S4 answers that in three places:
- Reads never serve it. Every read verifies the blob against the SHA-256 the record holds, and a mismatch is an error, not a response body. In a cluster the read then comes from another replica, so the client sees nothing.
- Writes never adopt it. The section above.
- The scrubber repairs it. In cluster mode the background scrubber walks the
volumes, verifies each blob against its CRC32, fetches a healthy copy from a
replica for the ones that fail, and writes it back to this node's disk. Only
the bytes change: the object keeps its ETag, size, metadata, HLC and version.
A cycle that could not heal something comes back in
S4_SCRUBBER_UNHEALED_RETRY_SECSrather than waiting out the full scan period, because the peer it could not reach may be one restart away.
What an operator watches:
# Damage found, and damage actually repaired, per node.
s4_scrubber_corruptions_found_total
s4_scrubber_corruptions_healed_total
s4_scrubber_blobs_scanned_total
The gap between the first two is the number of blobs this node is still carrying damaged — on a single node, where there is no replica to fetch from, that is all of them.
Design Details
- Hashing algorithm: SHA-256 (cryptographically strong, no collisions in practice)
- Dedup granularity: whole-object (not block-level)
- All objects are deduplicated regardless of size
- Dedup index is stored in a dedicated fjall keyspace alongside other metadata