Science is full of sinister-sounding names. First there was dark matter, the invisible substance that makes up most of the Universe. Then there was dark energy, the mysterious force making the Universe expand faster.
Both sound edgy, but they weren't called 'dark' because they're evil. Instead, they acquired the label because they're mysterious. Scientists know they exist – they're trying to figure out what they do, but their true nature remains unclear.
Now, that same idea has entered biology, where scientists are discovering hidden layers of complexity, concealed inside our cells.
Welcome to the shadowy world of dark DNA, where things aren't quite as they seem and molecules don't play by the rulebook.
“There’s just so much about how our DNA works that we don’t fully understand,” says Dr Marie Brunet, a biochemist at the Université de Sherbrooke, Canada.
Dark DNA is the name given to previously overlooked stretches of the genetic code, which are now being shown to be important. Research in this area is forcing scientists to rethink some of the assumptions that underpin modern biology.
Long-held ideas about how cells work are being challenged, revealing insights into how organisms develop and evolution unfolds.
The findings are offering fresh perspectives on diseases, such as cancer and Alzheimer’s, and treatments based on these findings are already being tested.
There’s still much to learn, but it’s looking a lot like dark DNA could be one of the biggest discoveries of the century.
“I’m not someone to oversell things,” says Dr Sebastiaan van Heesch from the Prinses Máxima Center for Paediatric Oncology in Utrecht, the Netherlands, “but I do think it’s a game-changer and it’ll rewrite the textbooks.”
Gene sweepstake
The story begins around 25 years ago, when scientists were decoding the human genome.
The human genome is the blueprint of life. It’s all of the DNA that lives inside your cells, and it contains the instructions needed to build a person and keep them alive.
To keep things interesting, researchers organised a wager. For the price of a dollar or two, they could place a bet on how many genes the genome would be found to contain.
Genes are the fundamental units of heredity; short stretches of DNA that code for specific proteins. These proteins, in turn, do everything.
If genes are the plans for building a house, then proteins are the bricks, mortar, nails and construction workers that make the house happen.

All in all, 228 bets were placed. The winner was geneticist Dr Lee Rowen from the Institute for Systems Biology in Seattle, Washington, in the US, who guessed closest to the correct number and bagged the $1,200 (around £900) prize.
It turns out that the human genome contains around 20,000 protein-coding genes.
This was much lower than most had predicted, but perhaps the biggest surprise was that the vast majority of the human genome – a staggering 99 per cent of its three billion letters of genetic code – didn’t code for proteins at all. They didn’t seem to do much of anything.
Some called it ‘junk DNA’, but you know how sometimes valuable items get thrown out by accident? A sentimental letter ends up in the recycling bin or an earring falls into the trash?
Well, hiding in this ‘junk’ were bits of DNA that weren’t genes, but that were precious all the same. Among them were stretches of DNA that weren’t expected to code for proteins, but that somehow, did.
The sequences were smaller than most regular genes and provided the instructions for making proteins that were tiny too.
In 2001, for example, Prof Ikuo Nishimoto and his team from Keio University in Tokyo, Japan, discovered a teeny gene that makes a teeny protein called humanin.
Most proteins contain hundreds, or even thousands of amino acids. Humanin contains just 24. Six years later, tarsal-less was discovered in fruit flies; another miniature protein, this time less than 33 amino acids long.
Far from being junk, sitting around doing nothing, tiny dark genes were spurring the production of tiny dark proteins.
The genes were found in odd places, too. Some were identified in the vast ‘genetic wasteland’ of the genome, the 99 per cent of the genetic sequence thought not to code for proteins. Others were found hiding inside regular genes.
Take the fused in sarcoma (FUS) gene. It’s a big gene that has been known about for decades. When the gene contains certain spelling mistakes, known as mutations, it can lead to the neurodegenerative disease amyotrophic lateral sclerosis (ALS).
Brunet has shown that this one gene contains a second, much smaller gene, nestled inside it like a Russian doll. This codes for a second, previously unknown dark protein.
In other words, a single stretch of DNA can produce two completely different proteins. It’s a bit like discovering a page in a recipe book that contains two different recipes, depending on where you start reading.
The finding is at odds with the decades-old dogma, still taught at school, which asserts that one gene codes for one protein.
When scientists decoded the human genome and found around 20,000 protein-coding genes, they expected that there would be around 20,000 proteins.
But Brunet’s work, and that of others, suggests the genome may be more information dense than was previously thought. And if that’s the case, there could be far more than 20,000 proteins.
Dark DNA, hiding in plain sight, could be producing an entire anthology of dark proteins with equally dark functions we don’t fully understand.
The question is, how many of these dark proteins are there and what, if anything, do they do?
Read more:
- The simple, surprising habits that can slow your immune ageing by years
- 'A really substantive jump': The most powerful weight-loss drug ever tested has arrived
- 10,000 steps, 8 hours sleep: Here's what science actually says about recommended health targets
A hidden world
Before 2009, dark proteins, such as humanin and tarsal-less, were viewed as a bit of an oddity. Researchers didn’t know if they were quirky one-offs, or the tip of a much larger iceberg.
Then Prof Jonathan Weissman and Prof Nicholas Ingolia from the University of California, San Francisco, developed a new technique called ribosome profiling.
Ribosomes are the cell’s protein-making machines. Ribosome profiling provides a snapshot of where this machinery is active.
The results were a revelation.
The technique revealed that ribosomes were reading genetic instructions from thousands of unexpected regions of the genome, suggesting the existence of a huge hidden world of previously unknown proteins.
Humanin and tarsal-less weren’t one-offs; they were the tip of an enormous iceberg.
Researchers weren’t sure what to call these ‘new’ molecules. Some dubbed them dark proteins. Others named them ghost proteins.
They were called microproteins, alternative proteins and small ORFs (ORF stands for open reading frame, parts of the genome with the potential to code for proteins).
In Brunet’s lab, the plucky, little molecules were likened to the heroes of a story. “We call them our little hobbits,” she says, because hobbits are small.
And just as The Lord of the Rings makes no sense without Frodo and his friends from the Shire, so too human biology makes no sense without them.
In the end, researchers settled on the name ‘peptidein.’ The term is a portmanteau of peptide and protein, and reflects the fact that most of these molecules sit somewhere between the two.

Peptides are short chains of amino acids; proteins are the bigger structures made from them, folded into complex shapes. “We were going to call them pepteins,” says Brunet, “but then we realised it’s the name of a Thai energy drink.”
Peptidein, meanwhile, “is a deliberately vague umbrella term,” according to Dr John Prensner from the University of Michigan, in the US.
Traditionally, a protein can only be called a protein if it’s shown to have a function. Some proteins, for example, help to transport oxygen; others help muscles contract.
Most dark proteins have equally dark roles. Researchers don’t know what, if anything, most of them do. So, the term peptidein is used as a temporary measure.
Then at some point, if the molecule is proven to have a function, it gets an upgrade. A peptidein becomes a microprotein.
In a recent paper, a consortium of more than 60 researchers, including Brunet, van Heesch and Prensner, studied over 7,000 previously ignored genetic regions and found that around a quarter appear to produce peptideins.
This suggests there could be around 1,700 of these tiny molecules and hints that human cells may produce a much greater variety of proteins than was previously thought.
Now, peptideins have become a hot topic in the world of biology, and researchers are rapidly moving from ‘what exists’ to ‘what it does’.
Some peptideins have already been upgraded.
Humanin, for example, appears to protect human cells from stress and death, especially in the context of Alzheimer’s disease.
And tarsal-less plays a pivotal role in embryonic fruit fly development. If it’s not there, the insects can’t develop properly. They have defects in their respiratory systems and outer body structures.
According to the consortium, known as TransCODE, at least 50 of the 1,700 potential peptideins they highlighted could be playing key roles in the everyday workings of cells. As for the rest, the jury is out.
Some could be unintentional byproducts of the cell’s protein-making machinery, products that don’t have a specific function but get made anyway, then are quickly degraded. Others may exist just to gum up the ribosome.
Protein production is tightly controlled. The cell has various methods to do this, but one way could be the production of decoy proteins whose only function is to prevent the ribosome from making anything else.
Some peptideins could be decoy proteins.
Implications for cancer
Peptideins can be found in healthy human cells, but they can also be found in diseased cells. Cancer cells, for example, seem to have lots of them.
Prensner is especially interested in this. As a physician-scientist who studies dark proteins and treats children who have brain cancer, he’s interested in what makes cancer cells different.
“Right now, we’re seeing a bunch of things in cancer cells that we don’t see in non-cancer cells,” he says.

Prensner and colleagues have shown, not just that cancer cells contain peptideins, but that certain cancers depend on them. When the dark DNA that codes for them is disabled – so the peptideins can’t be made – the cells stop growing.
This has been seen in cultured breast, prostate, colorectal and brain cancer cells.
The finding hints that some cancer cells may have a hidden, dark layer of biology driving their expansion. But while this sounds ominous, it could lead to the development of new treatments.
There’s another piece of cellular machinery, called the human leukocyte antigen (HLA) complex, whose job is to display fragments of proteins – big and small – on the outside of the cell.
This gives the immune system a glimpse of what’s happening on the inside, helping it work out which cells to leave alone and which to attack.
This includes many peptideins, which are chopped into pieces and displayed on the surface of cancer cells by HLA molecules.
Now, researchers are designing vaccines that teach the immune system to recognise these fragments and hopefully destroy the cancer cells they’re part of.
Van Heesch has designed a vaccine like this for Ewing sarcoma. Ewing sarcoma is a rare and aggressive cancer that mostly affects the bones. Around 600 young people are diagnosed with it every year in Europe.
Treatment typically includes intensive chemotherapy, surgery and radiation therapy, but it’s brutal. Although the cancer cells take a battering, so do healthy cells, leaving patients with side effects that can seriously affect their quality of life.
Thanks to the HLA complex, Ewing sarcoma cells have lots of different peptidein fragments on their surface. The vaccine has been designed to highlight as many of these as possible to the immune system.
“That’s the beauty of it,” says van Heesch. “The molecules are so small, we could put many of them into a single vaccine.”
Now the vaccine is being tested in an early-stage clinical trial.
Funded by Fight Kids Cancer, in collaboration with the Institut Curie in Paris, France, and the Centre for Paediatric and Adolescent Medicine at Heidelberg University Hospital, Germany, up to 45 patients with Ewing sarcoma will receive the treatment.
Then, if things go well, larger trials will follow.
“It’s looking very promising,” says van Heesch. “We’re really excited to see that our work on dark proteins is starting to have a clinical impact.”
Beyond cancer, the discovery of peptideins could lead to new treatments and diagnostics in other areas too, such as heart disease and neurodegeneration.
And now they have an official label – peptidein – that can be included in official databases, the hope is that researchers worldwide will start to include them in their work.
Slowly but surely, a layer of biology that was once hidden in the shadows is finally coming into view. What was once dismissed as biological dark matter is emerging as an important part of how life works. But there’s one more thing.
Proteins of the future?
Most human proteins have similar counterparts in other, closely related species, suggesting they perform important biological roles that evolution has preserved across millions of years. Peptideins buck this trend.
They don’t seem to have similar versions in other species.

Van Heesch has studied heart tissue from humans, rats and mice and found “a completely different landscape of these things. This suggests that evolution could be more dynamic than we expected,” he says.
Think for a moment about how animals evolve. Scan the fossil record and you’ll see that it contains a mix of fleeting experiments and enduring successes.
Strange animals, such as the five-eyed Opabinia, flourished briefly before their lineage came to an end, while crocodiles and their relatives have persisted for over 200 million years.
The same pattern may hold at the molecular level. Some peptideins may prove to be fleeting evolutionary experiments, here today but gone tomorrow, because they didn’t help the cell do anything useful.
But some potentially useful peptideins could be the proteins of tomorrow, caught in the process of evolving.
We know that evolution works on achingly long timescales, so these ‘tomorrows’ may take a while to come.
But if we could fast forward a million years, there may be new proteins in existence in the far future that owe their origins to the peptideins we see today.
Long before then, however, the terms ‘dark DNA’ and ‘dark proteins’ will become extinct. As scientists discover their secrets, the darkness will disappear, leaving behind a deeper understanding of life itself.
Read more:

