All known life stores its genetic code with four letters: A, T, C and G. Researchers at the University of California San Diego have now shown that one of biology's most essential enzymes can accurately read and copy an alphabet with eight letters, including four that do not exist in nature.
Beyond the four natural letters
The team studied RNA polymerase from E. coli bacteria, the enzyme that copies DNA into RNA. In experiments described in Nature Communications, the enzyme recognized and transcribed DNA that contained two synthetic base pairs, called P:Z and S:B, alongside the four natural ones. The synthetic letters fit into the double helix using alternative patterns of hydrogen bonds while keeping the same overall shape as natural DNA.
How the enzyme copes with new letters
Using structural imaging, the researchers captured detailed views of the enzyme as it worked. The images showed that RNA polymerase identifies the synthetic letters with the same biochemical and structural signals it uses for natural base pairs, which helps explain why it can copy the expanded alphabet accurately. In a related study published in PNAS, the same group reported that the enzyme can also handle a synthetic pair that lacks hydrogen bonds entirely.
What an eight-letter code could do
The finding supports a long-standing goal in synthetic biology: expanding the language of DNA so cells can carry information that nature never wrote. Earlier work has used expanded alphabets to build synthetic DNA molecules that can recognize liver cancer cells. Scientists say such systems could one day produce new diagnostics, therapies and engineered biological machines, and could make it possible to store far more data in DNA than four letters allow. The team says the next step is learning whether the full eight-letter system can support proteins and normal cell functions, not just copying.