---
title: DNA Data Storage
type: vocabulary
url: "https://www.envisioning.com/vocab/dna-data-storage"
summary: Encoding digital information in synthesized DNA strands for ultra-dense, long-duration archival storage.
year: 2012
generality: 0.42
---

# DNA Data Storage

Encoding digital information in synthesized DNA strands for ultra-dense, long-duration archival storage.
DNA data storage is the practice of encoding digital information as sequences of synthetic nucleotides (A, C, G, T) and recovering that information by reading the synthesized strands back with DNA sequencers. The four-letter nucleotide alphabet is mapped onto the binary alphabet of digital files via an error-correcting code, allowing arbitrary bytes — text, images, video, software — to be translated into a chain of As, Cs, Gs, and Ts that can be chemically synthesized in a test tube and later sequenced to recover the original bitstream.

The mechanism is essentially a four-step pipeline. First, the source data is compressed and chunked into blocks, then each block is encoded with redundancy so that small synthesis or sequencing errors can be corrected on read. Second, oligonucleotides (short, single-stranded DNA sequences) corresponding to the encoded blocks are manufactured — historically by phosphoramidite chemical synthesis, increasingly by template-independent DNA polymerases. Third, the synthesized strands are pooled, dried, or encapsulated for stable storage. Fourth, recovery involves PCR amplification of the relevant blocks (when addressable) and high-throughput sequencing, followed by reassembly and error correction. The information density is roughly a thousand times that of the densest silicon media, and the half-life of dried DNA under cool, dry conditions is on the order of centuries.

The tradeoffs are dominated by cost and write speed. Chemically synthesized DNA currently runs on the order of dollars per megabyte to write and dollars more to read, making the technology economically viable only for cold archival workloads — petabyte-or-larger datasets that must be preserved for decades and read rarely. Tape and hard-disk archives remain orders of magnitude cheaper for actively used data. Enzyme-driven synthesis, as developed at the Wyss Institute and several startups, promises to lower the per-base cost by replacing expensive phosphoramidite reagents with terminal deoxynucleotidyl transferase (TdT) controlled in solution, but the field has not yet crossed the dollar-per-megabyte threshold needed to compete with magnetic media at scale.

Open questions include whether enzymatic synthesis can be multiplexed to throughputs that matter for industrial archival workloads, whether random-access read-out (selecting one file from a petabyte pool without sequencing everything) can be made economical, and whether the technology will move from research to standard infrastructure for exabyte-scale cold storage. The broader research direction — molecular information substrates including synthetic polymers, peptide chains, and other non-nucleotide chemistries — is sometimes grouped under molecular data storage, with DNA as the most mature instance. The plausibility of a long-term cold-storage tier built on DNA is widely accepted in the archival community; the timing of that transition is unresolved.

---
Source: Envisioning — Technology Research Institute (https://www.envisioning.com/vocab/dna-data-storage)
