Skip to main content

Campbell Biology · Chapter 21

Genomes and Their Evolution

pp. 426–449 · 6 sections

This chapter zooms out from single genes to whole genomes. It covers how scientists read all of an organism's DNA, how computers find the genes in it, what fills the huge noncoding part of eukaryotic genomes, and how duplications, rearrangements and jumping DNA make new genes. Most of it is background for the AP course, but DNA sequencing (Topic 6.8), mutations and polyploidy (Topics 6.7 and 7.10), and DNA and protein evidence for common ancestry (Topics 7.6, 7.7 and 7.9) are all tested.

Independent review — not affiliated with or endorsed by the publisher. You'll need your own copy of the book.

21.1 Sequencing whole genomes

pp. 427–429

On the AP exam? Background

Topic 6.8 expects you to know that DNA sequencing reads the order of bases and is used in medicine, forensics and evolution. How the Human Genome Project mapped and assembled the genome is background, and linkage maps tie back to recombination frequency in Topic 5.4.

In the course: Topic 6.8 Biotechnology, Topic 5.4 Non-Mendelian Genetics (notes, videos and more questions)

Key points

  • A genome is the full set of DNA an organism carries. Genomics studies whole genomes at once, and bioinformatics is the computer work that stores, searches and compares all that sequence data.
  • The public Human Genome Project, which began in 1990, worked step by step. It first ordered genetic markers on a linkage map using how often crossing over separates them, then built a physical map measured in base pairs from overlapping cloned pieces, and only then read the pieces base by base.
  • Shotgun sequencing skips the maps. You break many copies of the genome into random fragments, read them all, and let software rebuild the original sequence by matching up the overlapping ends. Long repeated stretches are the hard part, because a read from a repeat could belong in many places.
  • A public team and a private company announced rough drafts of the human sequence in 2000–2001, and the public project called it essentially complete in 2003. Roughly 8% of it, mostly highly repetitive DNA around centromeres, stayed unread until a gap-free 'telomere-to-telomere' sequence was published in 2022.
  • Sequencing has become vastly faster and cheaper. The first human genome took more than a decade and cost hundreds of millions of dollars; today a genome can be read in a day or two for a few hundred dollars.
  • Modern machines read millions of fragments at the same time without cloning them first. Short-read machines give pieces of a few hundred bases, while long-read machines can read tens of thousands of bases in one go, which helps them get through repeats.
  • Metagenomics means sequencing all the DNA in an environmental sample, like seawater, soil or the gut, at once. It reveals microbes that nobody has managed to grow in the lab.
Key terms (14)
genome
The complete set of DNA in an organism, including its genes and all the DNA between them. One human copy is about 3.1 billion base pairs.
genomics
Studying whole genomes at once, including all the genes, how they're arranged and how they work together, instead of one gene at a time.
bioinformatics
Using computers, databases and statistics to store, search and compare huge amounts of biological data such as DNA and protein sequences.
Human Genome Project
The international public effort, from 1990 to 2003, that produced the first reference sequence of the human genome.
genetic marker
Any stretch of DNA with a known location that varies between individuals, so you can follow it through families or populations.
linkage map
A map that orders markers by how often crossing over separates them. Its distances are in map units, which reflect recombination, not actual DNA length.
physical map
A map that places markers by actual DNA distance, counted in base pairs.
whole-genome shotgun sequencing
Reading a genome by breaking many copies into random fragments, sequencing every fragment and letting software reassemble the genome from overlaps.
read
The sequence a machine reports for one DNA fragment. Short reads are a few hundred bases; long reads can be tens of thousands.
sequence assembly
Joining reads together by their overlapping ends to rebuild longer continuous sequences, ideally whole chromosomes.
coverage
How many times, on average, each base of a genome appears among the reads. Higher coverage helps catch errors and fill gaps.
high-throughput
Describes lab methods that handle huge numbers of samples or molecules at once and produce data very quickly.
metagenomics
Sequencing the mixed DNA of a whole community straight from an environmental sample, so you can study microbes without growing them.
telomere-to-telomere (T2T) assembly
A genome sequence with no gaps, running from one end of each chromosome to the other. A complete human one was first published in 2022.

Check yourself: 21.1 Sequencing whole genomes

4 questions on 21.1 Sequencing whole genomes. Pick an answer to see if you got it, and why.

Question 1 of 4

Three reads from the same strand of one region of a genome are listed 5′ → 3′: Read 1: TTAGCCATGG Read 2: CATGGTCAAG Read 3: GACTTAGCCA Which sequence results from assembling the reads by their overlapping ends?

Question 2 of 4

A plant genome contains thousands of nearly identical copies of a 5,000-base-pair transposable element. A team sequences the genome using only 150-base reads, and the assembly breaks into many separate pieces wherever a copy of the element sits. Which change would most likely close these gaps?

Question 3 of 4

On a linkage map, markers P and Q are 5 map units apart, and so are markers R and S. Sequencing later shows that P and Q are 0.8 million base pairs apart, while R and S are 6 million base pairs apart. Which explanation best accounts for this?

Question 4 of 4

Microbiologists want to learn which bacteria live in a mat of microbes around a deep-sea hydrothermal vent. Fewer than 1% of the cells grow on any lab medium they have tried. Which approach would give the most complete picture of the community?

0 of 4 answered

21.2 Bioinformatics: finding genes and what they do

pp. 429–432

On the AP exam? Background

The exam won't ask about databases or software by name. It does expect you to know that sequence data are used in medicine and research (Topic 6.8) and that comparing sequences shows how species are related (Topic 7.9).

In the course: Topic 6.8 Biotechnology, Topic 7.9 Phylogeny, Topic 6.3 Transcription and RNA Processing (notes, videos and more questions)

Key points

  • Sequences from labs everywhere go into free public databases. GenBank is run by the US National Center for Biotechnology Information and swaps data with partner databases in Europe and Japan, and the Protein Data Bank stores 3-D protein structures.
  • Search tools such as BLAST compare a new DNA or protein sequence with everything already stored and list the closest matches, often from other species.
  • Gene annotation means finding the genes in raw sequence. Software looks for promoters, start and stop codons, open reading frames and splice sites, and checks them against RNA that cells actually make.
  • Similar sequence usually means similar job. If part of a new protein looks like a known enzyme's active region, it probably does something related. Protein sequences are often better to compare than DNA, because the redundant genetic code lets many DNA changes leave the amino acid the same.
  • When a sequence matches nothing known, researchers turn it off with CRISPR or RNA interference and look at what changes. Working from gene to trait this way is called reverse genetics.
  • Systems biology studies how many genes and proteins work together as networks. Researchers measure all of a cell's proteins (the proteome) or all of its RNA at once, now mostly by RNA sequencing, which has largely replaced microarray chips.
  • Large projects map the genome's working parts. ENCODE catalogued promoters, enhancers and other elements and found that much of the genome is copied into RNA at some point, though whether all of that RNA has a job is still debated. The Cancer Genome Atlas compared tumors with normal tissue in 33 cancer types to find drug targets and support personalized medicine.
Key terms (13)
GenBank
A huge free public database of DNA sequences, run by the US National Center for Biotechnology Information (NCBI).
BLAST
A search tool that compares a DNA or protein sequence with a database and lists the most similar sequences it finds.
gene annotation
Finding where the genes are in a raw genome sequence and working out what they probably do.
open reading frame (ORF)
A stretch of DNA that runs from a start codon to a stop codon in the same frame, with no stop codon in between. It's a candidate protein-coding gene.
protein domain
A part of a protein that folds on its own and has its own job, such as binding ATP. The same kind of domain can show up in many different proteins.
reverse genetics
Starting with a gene's sequence and working out its role, usually by switching the gene off and seeing what changes.
gene knockout
An organism or cell in which one gene has been deliberately disabled, often with CRISPR, so you can see what the gene does.
RNA interference (RNAi)
Using small RNAs that match a gene's mRNA to get that mRNA destroyed or blocked, which silences the gene.
proteome
The full set of proteins a cell, tissue or organism makes at a given time.
proteomics
Studying whole proteomes, including which proteins are present, how much of each, and how they interact.
RNA sequencing (RNA-seq)
Reading all the RNA in a sample to measure how strongly each gene is being expressed.
systems biology
Studying how many parts, such as genes, proteins and metabolites, work together as a network rather than one at a time.
personalized medicine
Choosing how to prevent or treat a disease based on a person's own genes, or on the genes of their tumor.

Check yourself: 21.2 Bioinformatics: finding genes and what they do

4 questions on 21.2 Bioinformatics: finding genes and what they do. Pick an answer to see if you got it, and why.

Question 1 of 4

Gene-finding software flags open reading frames: a start codon (ATG on the coding strand) followed, in the same frame, by codons with no stop codon (TAA, TAG or TGA) until a stop codon is reached. Each stretch below is part of a coding strand, written 5′ → 3′ and read in frame from its first base. Which one runs from ATG to a stop codon at its end with no earlier stop?

Question 2 of 4

The same gene is compared between a shark and a human. Their coding DNA sequences are 64% identical, but the proteins they code for are 81% identical in amino acid sequence. What best explains why the proteins are more alike than the DNA?

Question 3 of 4

A gene of unknown function is sequenced from a deep-sea worm. A database search finds that one region of its predicted protein closely matches the zinc-finger region that many known transcription factors use to grip DNA. The rest of the protein matches nothing in the database. Which hypothesis is best supported?

Question 4 of 4

Researchers find a zebrafish gene, zx1, whose sequence doesn't resemble any known gene. Which experiment would most directly test what zx1 does?

0 of 4 answered

21.3 Genome size, gene count and gene density

pp. 432–434

On the AP exam? Background

You won't be asked to compare genome sizes, but alternative splicing (Topic 6.3), how bacterial and eukaryotic DNA differ (Topic 6.1) and introns as a shared eukaryote feature (Topic 7.7) are all tested.

In the course: Topic 6.3 Transcription and RNA Processing, Topic 6.1 DNA and RNA Structure, Topic 7.7 Common Ancestry (notes, videos and more questions)

Key points

  • Genome size is counted in base pairs, often in megabases (Mb, millions of base pairs). Most bacteria and archaea have 1–6 Mb, though some bacteria that live inside other cells have far less. Eukaryotes usually have much more: baker's yeast has about 12 Mb and humans about 3,100 Mb.
  • Among eukaryotes, genome size doesn't track how complex an organism is. An onion has roughly five times as much DNA as you, and some lungfish, salamanders and ferns have tens of times more. The extra is mostly repetitive, noncoding DNA, not extra genes.
  • Free-living bacteria and archaea typically have a few thousand genes, while eukaryotes range from a few thousand in some single-celled fungi to tens of thousands in many plants. Humans have about 20,000 protein-coding genes, about the same as a tiny roundworm and fewer than rice or wheat.
  • Humans get a lot out of relatively few genes. Most genes with several exons are spliced in more than one way, proteins are modified after they're made, and small RNAs and transcription factors fine-tune when and where each gene works.
  • Gene density is the number of genes per stretch of DNA. Bacteria pack roughly 900–1,000 genes into each Mb, while humans average only about 6–7, because so much of our DNA lies in introns and between genes.
  • Bacterial genes run without interruption and most of a bacterial genome codes for proteins or RNAs. Eukaryotic genes are broken up by introns, so a typical human gene spans tens of thousands of base pairs, and long stretches of DNA lie between genes.
Key terms (11)
base pair (bp)
One matched pair of nucleotides across the two strands, like A with T. Genome sizes are counted in base pairs.
megabase (Mb)
One million base pairs. A typical bacterial genome is a few Mb, and the human genome is about 3,100 Mb.
genome size
The total amount of DNA in one copy of an organism's genome.
protein-coding gene
A gene whose final product is a polypeptide, as opposed to a gene for a working RNA such as rRNA or tRNA.
gene density
The number of genes packed into a stretch of DNA, often counted per million base pairs.
C-value paradox
The puzzle that genome size doesn't match an organism's complexity. It's explained mostly by differing amounts of repetitive DNA.
noncoding DNA
DNA that doesn't code for a protein. Some of it, like promoters and enhancers, is regulatory, and much of it is repeats.
exon
A part of a gene that stays in the mature mRNA after splicing. Exons hold the coding sequence, plus the untranslated ends.
intron
A part of a gene that is copied into pre-mRNA but cut out before the mRNA leaves the nucleus. Bacterial genes don't have them.
alternative splicing
Keeping different combinations of exons from the same pre-mRNA, so one gene can code for several related proteins.
post-translational modification
A change made to a protein after it's built, such as cutting it or adding a sugar or phosphate group, that alters how it works.

Check yourself: 21.3 Genome size, gene count and gene density

4 questions on 21.3 Genome size, gene count and gene density. Pick an answer to see if you got it, and why.

Question 1 of 4Calculator allowed

Species X has a 5.2-Mb genome with about 4,900 protein-coding genes. Species Y has a 1,450-Mb genome with about 23,000 protein-coding genes. Which statement best compares them?

Question 2 of 4

Two closely related salamander species have about the same number of protein-coding genes, but one has a genome about 10 times larger than the other. What most likely makes up most of the extra DNA?

Question 3 of 4

One mouse gene has seven exons. The table shows which exons end up in its mature mRNA in three tissues (invented data). Tissue | Exons in mature mRNA Brain | 1, 2, 3, 5, 7 Liver | 1, 2, 4, 5, 7 Muscle | 1, 3, 4, 6, 7 Exon 2 codes for the region that anchors the protein in the plasma membrane. In which tissues would the protein most likely be held in the membrane?

Question 4 of 4

Which feature would you most expect to find in the genome of a newly sequenced free-living bacterium?

0 of 4 answered

21.4 Noncoding DNA, jumping genes and gene families

pp. 434–438

On the AP exam? Background

Topic 6.7 includes transposition as one way DNA moves and creates variation, and short tandem repeats are what DNA profiling compares (Topic 6.8). The percentages, the names Alu and L1, and the details of gene families are background.

In the course: Topic 6.7 Mutations, Topic 6.3 Transcription and RNA Processing, Topic 7.7 Common Ancestry, Topic 6.8 Biotechnology (notes, videos and more questions)

Key points

  • Just 1.5% or so of human DNA codes for proteins, rRNA or tRNA. Introns and gene-control regions such as promoters and enhancers add roughly another quarter, and most of the rest is repetitive DNA.
  • Some noncoding DNA clearly matters. Certain noncoding stretches are more alike across mammals than many genes are, and studies comparing many mammal genomes suggest that roughly a tenth of the human genome is kept in place by natural selection. Much of the rest has no known job, and how much of it does anything is still debated.
  • Transposable elements are pieces of DNA that can move to new spots in the genome. Barbara McClintock first found them in maize in the 1940s, long before most scientists accepted the idea. Elements and their broken-down remains make up about 45% of human DNA and up to 85% of the maize genome.
  • There are two main kinds. DNA transposons move with an enzyme called transposase, by cut-and-paste or copy-and-paste. Retrotransposons, the more common kind in eukaryotes, are copied into RNA and then turned back into DNA by reverse transcriptase, so the original always stays put and copies pile up.
  • In humans, Alu elements (about 300 base pairs, over a million copies, about 10% of the genome) and LINE-1 elements (about 6,000 base pairs long, about 17%) are the main families. Only a small number of LINE-1 copies can still move.
  • Simple sequence DNA is a short unit repeated over and over in a row. Large blocks sit at centromeres and telomeres, where they help chromosomes separate and protect their ends. Short tandem repeats (STRs) vary in length from person to person, which is why they're used in DNA profiling.
  • Many genes belong to multigene families of similar copies. The rRNA genes are repeated hundreds of times in tandem so cells can make ribosomes quickly. The α-globin and β-globin families hold related genes that are switched on at different stages of life, plus pseudogenes, which are copies that no longer work.
Key terms (14)
repetitive DNA
DNA sequences that appear many times in a genome. In humans, they make up about half of all the DNA.
transposable element
A segment of DNA that can move, or be copied, to a new location in the genome. It's sometimes nicknamed a 'jumping gene'.
transposition
When a transposable element shifts, or is copied, from one spot in the genome to a new one.
transposon
A transposable element that moves as DNA, either cut out and pasted elsewhere or copied and pasted.
transposase
The enzyme, usually coded by the transposon itself, that cuts a transposon out or copies it and inserts it at a new site.
retrotransposon
A transposable element that moves through an RNA copy, which reverse transcriptase turns back into DNA and inserts at a new site.
reverse transcriptase
An enzyme that builds DNA using RNA as its template. Retrotransposons and retroviruses such as HIV both use it.
Alu element
A short (about 300-base-pair) repeated element found mostly in primates. It has over a million copies in the human genome.
LINE-1 (L1)
A family of long retrotransposons, about 6,000 base pairs each when complete, that makes up about 17% of human DNA.
simple sequence DNA
DNA made of a short unit, often under 15 bases, repeated many times in a row.
short tandem repeat (STR)
A run of a 2–5 base unit repeated in a row. The number of repeats varies between people, so STRs are used in DNA profiling.
multigene family
A set of genes in one genome that are identical or closely similar, usually because they all descend from one ancestral gene by duplication.
pseudogene
A copy of a gene that has built up mutations, such as an early stop codon, and no longer makes a working product.
globin
A family of oxygen-carrying proteins. Hemoglobin is built from α-globin and β-globin chains.

Check yourself: 21.4 Noncoding DNA, jumping genes and gene families

4 questions on 21.4 Noncoding DNA, jumping genes and gene families. Pick an answer to see if you got it, and why.

Question 1 of 4

Cultured cells carry two active mobile elements, M and N. When the cells are treated with a drug that blocks reverse transcriptase, new insertions of M stop, but N keeps moving. Each time N moves, it disappears from its old site. Which conclusion is best supported?

Question 2 of 4

Transposable elements and sequences derived from them account for about 45% of human DNA. Which statement best describes most of this DNA?

Question 3 of 4

A family is tested at a short tandem repeat (STR) locus where the unit GATA is repeated. The mother's alleles have 9 and 12 repeats, and the father's have 10 and 14. Their child's alleles have 12 and 11 repeats. What is the most likely origin of the 11-repeat allele?

Question 4 of 4

In fruit flies, mutants that have lost part of their cluster of tandemly repeated rRNA genes develop more slowly than normal flies and grow short, thin bristles. Which explanation best fits these traits?

0 of 4 answered

21.5 How genomes change over time

pp. 438–442

On the AP exam? Yes

Topic 6.7 tests mutations and polyploidy as sources of new variation, and Topic 7.10 tests polyploidy and chromosome changes as routes to new species. Gene duplication, exon shuffling and the history of particular gene families are background, but they explain where new genes come from.

In the course: Topic 6.7 Mutations, Topic 7.10 Speciation, Topic 7.2 Natural Selection, Topic 7.6 Evidence of Evolution (notes, videos and more questions)

Key points

  • Every change in a genome starts with a mutation. Most changes are harmful or have no effect, but a rare useful one gives natural selection something to work with. To be passed on, a change has to happen in cells that make eggs or sperm.
  • A mistake in meiosis can give an organism extra full sets of chromosomes (polyploidy). Then every gene has spare copies, and one copy can keep the old job while the others are free to change. Polyploidy is common in plants, such as wheat, and rarer in animals, though whole-genome duplications happened early in vertebrate history and later in some fish and frogs.
  • Chromosomes also break and rejoin in new ways: pieces get duplicated, flipped (inverted), moved to another chromosome (translocated) or fused. Human chromosome 2, for example, formed when two ancestral ape chromosomes fused end to end. Hybrids between populations with different arrangements may have trouble with meiosis, which can help split one species into two.
  • Smaller duplications come from unequal crossing over, when homologous chromosomes pair slightly out of line, often at matching repeats, and swap uneven pieces, leaving one with an extra copy and the other with a gap. Slippage during DNA replication can also add or drop repeated units.
  • After a gene is duplicated, the copies can go three ways. They can stay similar and do related jobs, like the globin genes switched on at different life stages. One copy can evolve a new job. Or one copy can break down into a pseudogene.
  • Exons often code for separate protein domains. Because introns are long and get spliced out, crossing over or insertions inside them can copy or move whole exons, building proteins with new mixes of domains (exon shuffling).
  • Transposable elements speed things up. They break genes or change how genes are controlled when they land nearby, give misaligned chromosomes matching sequences to cross over at, and sometimes carry exons or whole genes along to new places.
Key terms (11)
polyploidy
Having more than two complete sets of chromosomes, such as 3n or 4n. It's common in plants and can create a new species in one step.
gene duplication
A mutation that makes an extra copy of a gene, giving evolution a spare copy to modify.
unequal crossing over
Crossing over between homologous chromosomes that are paired slightly out of line, leaving one with a duplicated segment and the other with a deletion.
replication slippage
A copying error in which the new strand slips along the template in a repeated region, adding or deleting repeat units.
inversion
A chromosome change in which a segment breaks out and is reinserted backward.
translocation
A chromosome change in which a segment moves to a different, nonhomologous chromosome.
chromosome fusion
Two chromosomes joining end to end to form one, which lowers the chromosome number.
paralogs
Genes in the same genome that are related because one arose from a duplication of the other, like the α- and β-globin genes.
exon shuffling
The mixing of exons within or between genes by recombination or transposition, producing proteins with new combinations of domains.
gene family
A set of related genes descended from one ancestral gene by repeated duplication and divergence.
divergence
The build-up of differences between two sequences, or two lineages, after they separate.

Check yourself: 21.5 How genomes change over time

4 questions on 21.5 How genomes change over time. Pick an answer to see if you got it, and why.

Question 1 of 4

In a family of animal proteins, each protein is built from several domains: regions that fold on their own and have their own jobs. When researchers compare the proteins with their genes, they find that the borders between domains almost always fall where the genes' introns are. Which hypothesis does this pattern best support?

Question 2 of 4

After a gene is duplicated, why is one copy often free to collect mutations that would have been harmful if the gene existed in only one copy?

Question 3 of 4

Four genes in one species' genome belong to the same family. The table shows how identical their amino acid sequences are (invented data). Pair | Identity (%) A–B | 94 A–C | 68 A–D | 67 B–C | 69 B–D | 68 C–D | 83 Assuming copies become steadily less alike over time after each duplication, which history fits best?

Question 4 of 4

Antarctic notothenioid fish make antifreeze glycoproteins that keep their blood from freezing. The antifreeze gene contains stretches nearly identical to parts of the gene for trypsinogen, a digestive enzyme precursor, and the fish still have a working trypsinogen gene. Which process best explains how the antifreeze gene arose?

0 of 4 answered

21.6 Comparing genomes to trace evolution and development

pp. 442–447

On the AP exam? Yes

Topics 7.6, 7.7 and 7.9 test using DNA and protein similarities as evidence of common ancestry and to build trees, and Hox genes come up in Topic 4.3. Human–chimp percentages, the plant MADS-box genes and specific evo-devo cases are background.

In the course: Topic 7.6 Evidence of Evolution, Topic 7.7 Common Ancestry, Topic 7.9 Phylogeny, Topic 6.6 Gene Expression and Cell Specialization, Topic 4.3 Signal Transduction Pathways, Topic 7.11 Variations in Populations (notes, videos and more questions)

Key points

  • Species whose genomes are more alike split from each other more recently. Close relatives are best for spotting recent changes, while very distant ones reveal deep history, such as the ancient split of bacteria, archaea and eukaryotes.
  • Some genes have changed remarkably little over a billion years or more. Yeast and humans share many genes for basic cell machinery, and some human genes can even stand in for their missing yeast versions. That's strong evidence of common ancestry, and it's why model organisms tell us about human biology.
  • Humans and chimpanzees differ at only a little over 1% of single bases, with more difference from insertions, deletions and duplications. Many of the differences that matter are thought to change when, where and how much genes are expressed, not just the proteins themselves.
  • Any two people match at about 99.9% of single bases. Most of the differences are SNPs, single bases that vary, but there are also CNVs, stretches of DNA that some people carry in more or fewer copies than others. Modern humans arose in Africa about 300,000 years ago, African populations hold the most genetic diversity, and most people outside Africa carry about 1–2% Neanderthal DNA.
  • Evolutionary developmental biology (evo-devo) compares how embryos are built. Animal Hox genes share a 180-base-pair homeobox that codes for a 60-amino-acid DNA-binding homeodomain, and in flies and mice the genes even sit in the same order along the chromosome as the body regions they control.
  • Since so many animals share the same developmental toolkit, body differences come mostly from changes in where, when and how strongly those genes are switched on, and in which target genes they control. Shifting the region where a Hox gene is active can change, for example, how many vertebrae carry ribs.
  • Plants and animals became multicellular separately. Both build their bodies with cascades of transcription factors, but plants use a different family, the MADS-box genes, as master switches for things like which flower part grows where.
Key terms (11)
comparative genomics
Comparing whole genomes of different species, or of different individuals, to learn about evolution and gene function.
highly conserved sequence
DNA or protein sequence that has barely changed over long evolutionary times, usually because changes to it are harmful.
model organism
A species that is easy to study in the lab, such as yeast, fruit flies or mice, used to learn about biology shared with other organisms.
single nucleotide polymorphism (SNP)
A spot in the genome where individuals in a population commonly have different single bases.
copy-number variant (CNV)
A stretch of DNA, often including genes, that different individuals carry in different numbers of copies.
evo-devo
Evolutionary developmental biology: comparing how embryos develop to learn how changes in development produce new body forms.
homeotic gene
A master regulatory gene that decides what a body region becomes, such as whether a segment grows legs or antennae.
Hox genes
A group of homeotic genes in animals, lined up on chromosomes in the same order as the body regions they help pattern from head to tail.
homeobox
A 180-base-pair DNA sequence found in Hox and many other developmental genes. It codes for the protein's DNA-binding region.
homeodomain
The 60-amino-acid part of a protein, coded by the homeobox, that binds DNA so the protein can switch other genes on or off.
MADS-box genes
A family of transcription factor genes that act as many of the master switches in plant development, such as deciding flower organ identity.

Check yourself: 21.6 Comparing genomes to trace evolution and development

4 questions on 21.6 Comparing genomes to trace evolution and development. Pick an answer to see if you got it, and why.

Question 1 of 4

The table gives the average number of single-base differences per 1,000 bases between the two copies of the genome that people carry, in several human populations (invented data that follow the real pattern). Population | Differences per 1,000 bases Eastern and southern Africa | 1.00 Middle East | 0.80 Europe | 0.77 East Asia | 0.72 Indigenous South America | 0.60 Which explanation best fits the pattern?

Question 2 of 4

Suppose researchers find that a Hox gene that gives segments a rib-bearing trunk identity is switched on in only a limited stretch of a lizard embryo, but along nearly the whole body of a snake embryo. Which explanation of the snake's long, rib-covered body does this best support?

Question 3 of 4

When the mouse gene Pax6, which contains a homeobox, is switched on in the leg tissue of a developing fruit fly, the fly grows extra eye structures on its legs. These are compound fly eyes, not mouse-style eyes. Which conclusion is best supported?

Question 4 of 4

Salivary amylase breaks down starch. People carry anywhere from 2 to more than 15 copies of the salivary amylase gene, and populations with traditionally starch-rich diets tend to have more copies on average. What kind of variation is this, and what is the most likely link to diet?

0 of 4 answered