Molecular evolution · illustrated

A gene from nothing

Most new genes are copies of old genes. A minority are not — they are read out of DNA that, in the ancestor, coded for nothing at all. This site walks through de novo gene birth: the mechanism step by step, the competing routes to the same-looking result, how the dating software genEra decides a gene is young, and what all of that looks like in a real hybrid plant genome.

Every figure is animated and steppable. There is also a game.

The problem, in one sentence

Duplication explains where most new genes come from: copy an existing gene, let one copy drift, and you have a novel function built out of an old part. But every genome sequenced turns up a tail of genes with no detectable relatives outside a narrow clade — in the genome examined here, several hundred per haplotype. Either they are genuinely new, or our ability to detect their ancient relatives has failed. Deciding which is the entire game, and it is harder than it sounds: the two possibilities produce identical output from a homology search.

The claim being tested

A de novo gene is one whose coding sequence derives from ancestrally non-coding DNA — not from a duplicate, not from a transposable element's own genes, not from a horizontal transfer. The positive evidence is a syntenic, verifiably non-coding ortholog in an outgroup. Absence of BLAST hits is not evidence; it is the thing that needs explaining.

How it happens

The canonical account has five moving parts. The animation below builds them one at a time — each step adds a layer to a single figure, so the last frame is the whole story.

Outgroup lineage — locus is non-coding gene A gene B Focal lineage — same position, same neighbours gene A gene B syntenic intergenic DNA low-level transcript (lncRNA) reading frame M * ORF, ~59 codons stop lost 80S peptide — now visible to selection proto-genes most are lost within a few My — one is retained allele freq. fixed — every individual carries it genEra call: PS 19 — youngest phylostratum correct here · ambiguous in general

Two dead ends that look like progress

Steps 2 and 3 are independent. A locus that only ever completes one of them never becomes a gene, and both intermediates are abundant in real genomes — which is exactly why de novo birth is thought to be a continuous process rather than a rare accident.

Transcript, no ORF

A locus with a promoter and no long open reading frame is a lncRNA. Plant genomes carry tens of thousands. Most are transcriptional noise; a few have RNA-level functions of their own. Neither outcome is gene birth.

ORF, no transcript

Random 1 kb of AT-rich intergenic DNA contains ORFs of 50+ codons by chance alone. Silent ones are invisible to selection and drift freely, accumulating stops. They are raw material, not genes.

The corollary for annotation: an ORF prediction is not a gene call. A large share of "lineage-specific genes" in any annotation are spurious ORF predictions in non-transcribed DNA, and they inflate the youngest phylostratum systematically.

What it takes to actually believe one

Ranked roughly by how much each piece of evidence buys you:

  1. A syntenic, non-coding ortholog in an outgroup The only positive evidence there is. Align the region, show the flanking genes match, and show the outgroup sequence is disabled — no ATG, a stop where the focal lineage lost one.
  2. The enabling mutation itself Best case: you can name the substitution or indel, and see the ancestral state in two independent outgroups. Now it is a described event, not an inference from absence.
  3. Expression RNA-seq in some tissue or condition. Weak, narrow, stress-inducible expression is the expectation, not a red flag.
  4. Translation Ribosome profiling or peptide-level mass spec. Distinguishes a real young protein from a transcribed ORF nobody reads.
  5. Population genetics Is it fixed or segregating? Is there a signal of purifying selection on non-synonymous sites? A segregating ORF is a proto-gene caught in the act.
  6. Structure, used carefully Predicted structure is weak evidence of age. Young de novo proteins are often intrinsically disordered and predict at low pLDDT — but so do ancient disordered proteins. Low pLDDT is consistent with youth; it does not demonstrate it.
The asymmetry that governs everything

Every step above is positive evidence. The homology search that started the whole process only ever supplies negative evidence — no hits — and the absence of a hit has two causes: the relative is not there, or the search could not see it. The genEra page works through a real case where the second cause won, and the concepts page covers how to put a probability on it.

Or: play it

The mechanism has a shape that is easier to feel than to read. De Novo is a QWOP-style game: Q and W alternate to drive RNA polymerase along the locus; O and P extend the open reading frame. Both decay while you attend to the other, and you only make protein when both are up. Which pair you reach for first decides which published scenario your run turns out to be.

→ Play De Novo