A joint guide
+ Guide · Omics

One code, many states.

A first-principles guide to omics: what each molecular layer records that the genome cannot, how each one is measured, and how a measurement becomes a result someone can act on.
+ The central question

Every cell in your body carries the same genome. So what does each omics layer measure that the genome alone cannot?

Starts at:the central dogma at school-biology level. No omics, statistics or instrument background assumed.Ends at:reading an omics report critically, and knowing what evidence a molecular signature needs before it becomes a test. Built around:one pair of identical twins, one of whom spent 340 days in orbit: the same genome, followed by ten research teams for 25 months.
+ Before we start

Before we start.

Most introductions to omics start with a list. Genomics studies genes, transcriptomics studies RNA, proteomics studies proteins, metabolomics studies metabolites, and each has its instruments. The list is correct and it explains almost nothing. It does not say why there are so many layers, why some are cheap to measure and others stubbornly hard, or why a result from one layer so often disagrees with a result from the next.

This guide is built around one question instead: every cell in your body carries the same genome, so what does each omics layer measure that the genome alone cannot? A neuron and a liver cell hold the same three billion letters. So do you at twenty and at seventy, and so did the astronaut in this guide's case study before and after a year in orbit. Whatever makes those states different is not written in the sequence. Each omics layer is a different way of reading that difference.

Here is the short answer, which the rest of the guide exists to justify. The genome is the instruction set, inherited and nearly fixed. The epigenome records which instructions a cell has made accessible. The transcriptome records which ones it is reading right now, and how often. The proteome records which molecular machines actually exist and in what form. The metabolome records the chemistry those machines are carrying out. Around these sit layers that add context rather than chemistry: which cells are present and in what state, where they sit in a tissue, which other organisms' genomes live alongside ours, and what cells shed into the blood. Each step away from the genome is closer to what the body is doing, changes faster, and is harder to measure.

That last point has a physical cause, and it runs through every chapter. DNA and RNA can be copied by enzymes and recognized by pairing with a partner strand written from their own sequence. Those two tricks let us amplify one molecule into billions and design a probe for any sequence on a keyboard. Proteins and metabolites allow neither. Every measurement of them has to work with the molecules already in the tube, which is why the layers closest to function are also the hardest to see.

How to read this guide

Chapters are numbered straight through, and each one opens with the question a careful reader would ask after the previous chapter. Plates are numbered separately so that any one can be cited on its own. Four devices recur:

  • Insight boxes (blue edge) carry the structural point of a section, or a worked calculation.
  • Caution boxes (orange edge) name a common misreading, a trap, or the limit of a claim.
  • In practice boxes (green edge) map the idea onto the reader's own work: at the bench, in instrument and assay design, or in a regulatory file.
  • Each chapter closes with what it established, in three or four lines.

The mathematics stays at ratios, percentages, logarithms and powers of ten. Where a number is derived, the arithmetic is shown once with real values. Where a plate is schematic rather than data, it says so.

Inside the plates, color is a legend and never decoration. DNA, RNA, protein, metabolite and chemical modification each have one color throughout the guide, the measured readout is always blue, and one orange mark in each plate points at the detail that matters most. A key strip under every plate lists only the colors that plate uses.

The limits of this guide

This is a reference for understanding, not a laboratory protocol, a regulatory opinion or clinical advice. Technology status, prices and regulatory positions are stated as of September 2026 and will date; where a figure is a manufacturer's claim rather than an independent measurement, the text says so.

+ Part I · Foundations

What the genome cannot tell.

The first three chapters build the model the rest of the guide leans on: why a single genome can produce very different cells, why some molecules are easy to count and others are not, and what physical tricks every omics instrument uses. Nothing here assumes more than school biology.

01 — One genome, many cell types

One genome, many cell types.

+ The questionIf a neuron and a liver cell carry identical DNA, what makes them different?

The instruction set and the state

Start with what is fixed. A human cell carries its genome as two copies of about 3.1 billion base pairs, one inherited from each parent. Somewhere around 20,000 stretches of that sequence are genes that encode proteins; their protein-coding portions occupy only one or two percent of the total, and the rest includes the switches, spacers and structural sequence that control when and where those genes are used. With a handful of exceptions, every nucleated cell in the body carries essentially the same sequence. The exceptions are instructive rather than fatal to the rule: red blood cells discard their nucleus, immune cells cut and rejoin their receptor genes, eggs and sperm carry one copy instead of two, and every dividing cell slowly accumulates a few new mutations. None of these explains why a neuron is a neuron.

Now look at what differs. A neuron is long, electrically excitable and packed with ion channels; a liver cell is compact, stuffed with enzymes that detoxify drugs and make blood proteins, and has no use for an axon. They are built from the same instruction set but they are running different parts of it. Biologists count a few hundred distinct cell types in the human body, and every one of them is a different selection from the same genome.

That is the distinction this guide turns on. The genome is the instruction set: it constrains what a cell can become. What the cell actually is at a given moment is its state: which instructions it is using, how intensely, and with what result. Two cells with identical genomes can be in radically different states, and one cell can change state within minutes.

01 — Same code, different cells
DNA RNA MADE PROTEIN MADE NEURON G1 G2 G3 G4 G5 G6 G7 G8 G9 G10 G11 G12 LIVER CELL G1 G2 G3 G4 G5 G6 G7 G8 G9 G10 G11 G12 identical in both cells
DNARNAProteinFocal detail
Plate 01 — A neuron and a liver cell carry the same genome letter for letter. What separates them is which genes are transcribed and translated — the part of the cell's state the genome does not record.

Reading is regulated at every step

The route from instruction to function has several steps, and the cell controls each one. DNA is packed around protein spools called nucleosomes, and chemical marks on the DNA and on the spools decide whether a gene is accessible at all. An accessible gene can be copied into RNA, a process called transcription, at a rate the cell sets. The RNA is edited, exported and eventually destroyed, and its lifetime is regulated too. A messenger RNA is read by a ribosome to build a protein, a process called translation, again at a controllable rate. The finished protein may be cut, folded, decorated with phosphate groups or sugars, sent out of the cell, or degraded. Many proteins are enzymes, and their activity determines which small molecules, the metabolites, are made and consumed.

Every one of those control points is a place where information enters that was not present in the step before. Knowing that a gene is accessible does not tell you how much RNA is being made from it. Knowing the RNA level does not tell you how much protein exists, because translation rates and protein lifetimes vary by orders of magnitude between genes. Knowing how much of an enzyme exists does not tell you how much chemistry it is doing, because enzymes are switched on and off by modification and by the metabolites around them.

This is the first answer to the central question. Each layer downstream of the genome records a decision the cell has made, and that decision is not written in the sequence.

The genome is a constraint, not a state

A genome tells you what a cell can do and what it is predisposed to. It does not tell you what the cell is doing. Every other omics layer is a way of reading some part of the state.

What the suffix adds

The suffix -ome means the complete set: the genome is all of an organism's DNA sequence, the proteome all of its proteins. Omics is the practice of measuring an entire set at once instead of one member at a time. A classical assay asks a narrow question: how much of this one protein is in this sample? An omics experiment asks a wide one: which of the thousands of proteins in this sample differ between these two groups?

The width is the point and also the price. Measuring everything at once lets you find what you did not know to look for, and it is where many biomarker candidates are now first found. But an instrument that measures twenty thousand things in one run cannot optimize itself for any one of them. Each individual measurement is typically less precise than a dedicated assay would be, and with twenty thousand measurements some will look interesting by chance alone. Both costs return later in this guide, the first in the chapters on individual layers and the second in Chapter 17.

02 — The omics stack
INHERITED CODE LIVING STATE GENOME EPIGENOME TRANSCRIPTOME PROTEOME METABOLOME RECORDS the inherited sequence which genes are accessible which genes are being read now which machines exist and act which chemistry is running HOW MANY ~20,000 protein- coding genes ~28 million CpG sites ~250,000 annotated transcripts proteoforms far outnumber genes ~218,000 listed in HMDB 5.0 CHANGES OVER a lifetime (near-fixed) minutes to years minutes to hours hours to days seconds to minutes READ BY sequencing sequencing after chemical conversion sequencing (via cDNA) mass spectrometry, affinity binding mass spectrometry, NMR copying and base-pairing are not available beyond this line
DNARNAProteinMetaboliteModificationFocal detail
Plate 02 — Each layer sits one step further from the inherited code and one step closer to what the cell is doing now. The boundary after RNA is the one that matters for measurement: to its left every molecule can be copied and read by pairing; to its right, neither trick works.

Each layer has its own shutter speed

The layers also differ in how fast they change, and that matters as much as what they record. The inherited sequence is fixed for life. Epigenetic marks range from histone modifications that can change within minutes to DNA methylation patterns that persist for years, and many are copied through cell division. Messenger RNAs in mammalian cells have lifetimes measured in hours, with a median of roughly nine hours in one widely cited study of mouse cells, while proteins in the same study lasted about two days on median. Many metabolites turn over in seconds to minutes.

So each layer is a photograph taken with a different shutter speed. The transcriptome is a snapshot of the last few hours of decisions; the metabolome is a snapshot of the last few minutes of chemistry. This has an immediate practical consequence. The faster a layer changes, the more it can change after the sample is taken: a blood tube left on a bench keeps metabolizing, and cells under stress during tissue collection start changing their RNA before anyone measures it. The further a layer sits from the genome, the more the handling of the sample becomes part of the result.

In practice

Before choosing a technology, name the layer that carries the signal you need. A question about susceptibility or inherited risk is a genome question. A question about what a tissue is doing now is a transcriptome, proteome or metabolome question, and it brings a sample-handling requirement with it: time to freeze, temperature, anticoagulant, storage. Those conditions belong in the protocol and the specification from the first draft, not as a later fix.

The case this guide is built around

Between March 2015 and March 2016 one of a pair of identical twins spent 340 days aboard the International Space Station while his brother stayed on Earth. Ten research teams sampled both men before, during and after the mission, over 25 months in total, and measured nearly every layer this guide covers: DNA methylation, gene expression, telomeres, proteins, metabolites, gut microbes and cognition. The combined results were published in Science in April 2019 as the NASA Twins Study.

Identical twins arise from a single fertilized egg, so the two men started with the same inherited genome. The spaceflight twin then lived through microgravity, radiation, confinement and a disrupted daily rhythm, while his brother served as a genetically matched control. It is close to an ideal experiment for this guide's question. With the genome held constant, anything that differs between the twins, or between one twin before and after flight, must have been recorded in another layer. Chapter 15 reads the results layer by layer, once the guide has built the tools to interpret them.

The misreading this case is famous for

In March 2018, news coverage of preliminary results reported that a year in space had changed 7% of the astronaut's DNA. It had not. The inherited sequence was unchanged; what had shifted, and in part stayed shifted, was the expression of a set of genes. The error is the one this chapter exists to prevent: treating a change in state as a change in the code.

+ What this chapter established
  • Nearly every cell carries the same genome; cell types differ in which parts of it they use.
  • Each step from DNA to function is regulated, so each layer carries information the previous one does not.
  • Omics measures a whole layer at once, trading per-molecule precision for the chance to find the unexpected.
  • Layers change at very different speeds, so the faster ones make sample handling part of the measurement.
02 — Why some molecules are easy to count

Why some molecules are easy to count.

+ The questionIf a whole genome can now be read for a few hundred dollars, why can't we measure every protein in a drop of blood?

Signal per molecule is the whole problem

Every detector, whether a camera, a photomultiplier or an ion counter, needs a minimum amount of signal before it can report anything above its own noise. One molecule, on its own, produces very little: a few photons if it carries a fluorescent dye, one ion if it survives the trip through a mass spectrometer. So the central engineering question of every omics measurement is the same. How do you get enough signal per molecule of interest, and how do you know that the signal came from that molecule and not from something else?

There are only a few ways to answer it. You can make more copies of the target until the signal is large. You can attach a label that recognizes the target and nothing else, and make the label bright. Or you can build a detector sensitive enough to see individual molecules and count them one at a time. Nucleic acids happen to allow the first two answers cheaply and almost universally. That single fact explains most of the difference in cost and depth between the layers of the omics stack.

Trick one: copying

DNA is a template for itself. A DNA polymerase walks along one strand and builds its partner, adding the base that pairs with each one it reads. Supply two short primers that flank the target, the four building blocks and heat cycling to separate the strands, and each cycle doubles the number of copies. This is the polymerase chain reaction. After 30 cycles one starting molecule has become roughly a billion, because 2 multiplied by itself 30 times is about 1.07 billion. A signal that was invisible is now large enough to measure with ordinary optics.

RNA can join in with one extra step. An enzyme called reverse transcriptase copies RNA into DNA, called complementary DNA or cDNA, and from there the DNA machinery takes over. That is why transcriptomics rides on the same instruments as genomics.

Proteins cannot be copied. Information in the cell flows from nucleic acid to protein and never back, so there is no enzyme that reads a protein and builds another one like it. Metabolites cannot be copied either; they are made by chains of enzymes, not from templates. Whatever proteins and metabolites are in the tube are all you will ever have to measure.

03 — Two tricks only nucleic acids allow
COPYING PAIRING 1 2 4 8 ~1 billion copies 30 doubling cycles: 1 molecule → ~1 billion copies (230 ≈ 1.07 × 109) 1 molecule stays 1 molecule — no enzyme copies a protein A T T A G C C G C G G C T A A T 5′ 3′ 3′ 5′ TARGET PROBE the probe is designed from the sequence alone each protein needs its own binder, raised and validated separately
DNAProteinFocal detail
Plate 03 — Nucleic acids can be copied by enzymes and recognized by a partner written from their own sequence. Proteins and metabolites allow neither, so every measurement of them must work with the molecules that are already there.

Trick two: pairing

The second property is just as important and less often noticed. Because A pairs with T and G pairs with C, a probe for any nucleic acid sequence can be written from the sequence alone. If you know that a target reads ATGCCGTA, its partner is known before anything is made, it can be synthesized chemically within days, and its binding strength can be predicted from its composition. The same chemistry works for every target. A panel of ten thousand DNA probes is ten thousand entries in a design file.

Proteins are recognized by shape, not by a code. A protein-binding reagent such as an antibody has to be raised or selected for each target, one at a time, and then tested to show that it binds the right protein and not a similar one. Nothing about a protein's sequence lets you write down its binder. A panel of five thousand proteins needs five thousand validated reagents, and ten thousand if the method uses two binders per protein. Every one of them can fail in its own way, and its specificity has to be proved, not predicted.

The asymmetry in one line

Nucleic acids can be multiplied and can be recognized by a partner designed on a keyboard. Proteins and metabolites allow neither, so every measurement of them must work with the molecules already present and must recognize them by mass or by bespoke binding.

The dynamic-range problem

A second difference compounds the first. Nearly every nuclear gene is present at two copies per cell, one from each parent, so nearly every genomic target starts at the same abundance. The genome is flat. Blood plasma is the opposite. Albumin circulates at around 40 milligrams per milliliter, while signaling proteins such as interleukin-6 sit at a few picograms per milliliter in a healthy person. That is a span of about ten orders of magnitude, and the rare proteins are often the informative ones.

Worked example: how outnumbered is a signaling protein?

Albumin: 40 mg/mL, molecular weight about 66,500. Moles per mL = 0.04 g ÷ 66,500 g/mol ≈ 6.0 × 10⁻⁷.

Interleukin-6: 2 pg/mL, molecular weight about 21,000. Moles per mL = 2 × 10⁻¹² g ÷ 21,000 g/mol ≈ 9.5 × 10⁻¹⁷.

Ratio ≈ 6.0 × 10⁻⁷ ÷ 9.5 × 10⁻¹⁷ ≈ 6 × 10⁹. For every molecule of interleukin-6 there are about six billion molecules of albumin.

Now picture an instrument that measures everything at once. A mass spectrometer receives ions in proportion to what is in the sample, and it has a limited capacity per second and a limited range between the largest and smallest signal it can record in one scan. If albumin and a few dozen other abundant proteins take almost all of that capacity, the rare proteins are never counted at all. This is why plasma proteomics spends so much effort removing or damping the abundant proteins before measurement, and why it still reaches far fewer proteins per sample than a sequencer reaches genes.

04 — Ten orders of magnitude in one tube
10−13 10−12 1 pg/mL 10−11 10−10 10−9 1 ng/mL 10−8 10−7 10−6 1 µg/mL 10−5 10−4 10−3 1 mg/mL 10−2 10−1 g/mL, log scale albumin ~40 mg/mL immunoglobulin G ~10 mg/mL fibrinogen ~3 mg/mL transferrin ~2.5 mg/mL C-reactive protein (healthy) ~1 µg/mL insulin (fasting) ~0.5 ng/mL cardiac troponin I (healthy) ~a few pg/mL interleukin-6 (healthy) ~2 pg/mL the signal you want ≈ 10 orders of magnitude molar ratio albumin : IL-6 ≈ 6 billion : 1
ProteinFocal detail
Plate 04 — Plasma proteins span about ten orders of magnitude in concentration. Any method that reads all molecules at once spends almost all of its capacity on the most abundant few, which is why the informative, rare proteins are the hardest to see.

What the two tricks predict

Put copying, pairing and dynamic range together and a ranking falls out that matches practice. The genome is the easiest layer: it can be copied, it can be paired, and every target starts at the same abundance. The transcriptome is next: it can be copied after reverse transcription and paired, but its molecules span a wider range of abundance and change quickly. The proteome is harder: nothing copies or pairs, abundances span ten orders, and each protein comes in modified forms. The metabolome is hardest to identify: nothing copies or pairs, the molecules are chemically diverse, and there is no template from which to predict what should be there.

The ranking also explains why so many clever methods for proteins and other molecules work by converting them into DNA. If two antibodies that recognize a protein each carry a short DNA tag, the tags can be joined and then copied and counted like any other DNA. Chapter 9 shows one such method in detail. It is a way of borrowing the two tricks for a molecule that does not have them.

What would break this account

The ranking is a claim about signal per molecule, not a law of nature. Any method that detects single molecules directly, with a specificity that does not depend on copying or pairing, removes the advantage nucleic acids enjoy. Single-molecule fluorescence counting, digital assays and early single-molecule protein sequencing all move in that direction. If one of them reached whole-proteome depth in plasma at sequencing-like cost, the ordering in this chapter would change. As of 2026 none has.

In practice

When an assay claims to measure thousands of analytes, ask two questions before anything else. Where does the signal per molecule come from: amplification, a bright label, or single-molecule detection? And how does the method know which molecule produced the signal: sequence, mass, or a binder whose specificity was tested for that analyte? The answers predict the depth, the failure modes and the validation burden more reliably than the headline number of analytes.

+ What this chapter established
  • Measurement depth is set by the signal you can get per molecule and how surely you can attribute it.
  • Nucleic acids can be copied and recognized by designed partners; proteins and metabolites cannot.
  • Plasma proteins span about ten orders of magnitude, so capacity is consumed by the abundant few.
  • Together these predict the ranking: genome easiest, then transcriptome, proteome, metabolome.
03 — The measuring toolkit

Five ways to measure a molecule.

+ The questionIf a protein cannot be copied or paired, how is it detected at all?

Every instrument measures a proxy

No omics instrument observes biology directly. Each one converts some physical property of a molecule into a signal a detector can count: photons, ions, or an electrical current. The biology is inferred from that signal through a chain of conversions, and Chapter 16 follows that chain in detail. For now the useful observation is that there are only five physical properties the whole field leans on, and every method in this guide is one of them or a combination.

The five are the order of the bases in a nucleic acid, the mass of a molecule, the shape that lets a reagent bind it, the light it emits or scatters, and the speed at which it travels through a medium. Each has a characteristic strength and a characteristic blind spot. Knowing which property a method reads tells you most of what you need to judge it.

Sequence: reading order by pairing

Sequencing reads the order of bases in DNA, or in RNA converted to DNA. Most instruments do it by making a partner strand one base at a time and detecting which base was added, usually as a flash of colored light. Others pull a single strand through a nanometer-scale pore and read the change in electrical current as each group of bases passes. Either way, sequencing depends on the pairing rule that makes nucleic acids special, which is why it is the dominant readout for the genome, the transcriptome and, after chemical conversion, the epigenome. Chapters 4 to 7 cover it.

Weigh: mass spectrometry

A mass spectrometer gives molecules an electric charge, moves them through electric or magnetic fields, and measures how they respond. The response depends on the ratio of mass to charge, written m/z, which modern instruments measure to within a few thousandths of a unit. Mass is universal: any molecule that can be charged and brought into a vacuum can be weighed, and the instrument does not need to know in advance what to look for.

That universality has two costs. The instrument sees whatever is abundant first, so rare molecules compete with common ones for its capacity, as Chapter 2 showed. And mass alone rarely identifies a molecule; many different molecules share a mass, so identity usually comes from breaking the molecule and weighing the pieces as well. Mass spectrometry is the workhorse for proteins and metabolites, and Chapters 8 and 10 cover it.

Bind: affinity reagents

A binder, usually an antibody or a short folded nucleic acid called an aptamer, grabs its target by shape. The binder itself carries a reporter: an enzyme that produces color, a fluorescent dye, a metal atom, or a DNA tag. Because the reporter can be amplified after binding, affinity methods can reach very low concentrations. An enzyme reporter turns over many substrate molecules per bound target, and a DNA tag can be copied a billion-fold.

The limit is that a binder finds only what it was made for. An affinity method measures a predefined list of targets, and each measurement is only as specific as its reagent. A binder that also sticks to a related protein reports both, and nothing in the signal reveals the mistake. Chapter 9 covers affinity proteomics; Chapters 11 and 14 rely on binders to label cells and vesicles.

See and count: one object at a time

Light is the reporter behind most of the other methods, but it becomes a measurement principle of its own when an instrument looks at individual objects. A flow cytometer measures each cell as it passes a laser; a microscope images each molecule in a tissue section; the most sensitive instruments detect single fluorescent molecules. When objects can be seen individually, the measurement changes character: instead of estimating how much signal a mixture gives, the instrument counts events. Counting is more robust than estimating intensity, and it is the basis of the digital methods that recur throughout this guide.

Separate: sorting by travel time

Separation methods do not identify anything on their own. Liquid chromatography pushes a mixture through a packed column, and molecules emerge at different times depending on how strongly they stick to the packing. Gas chromatography does the same for volatile molecules; electrophoresis sorts by size and charge in an electric field. Their job is to reduce complexity: a detector presented with a few molecules at a time performs far better than one presented with thousands at once. The time at which a molecule emerges, its retention time, also becomes an extra coordinate for identifying it.

05 — Five ways to measure a molecule
SEQUENCE base order, read by pairing WEIGH mass-to-charge, from motion in a field BIND shape-complementary antibody, aptamer SEE / COUNT light: fluorescence, scatter, one by one SEPARATE travel time by size, charge, hydrophobicity GENOME EPIGENOME TRANSCRIPTOME PROTEOME METABOLOME CELLS VESICLES + NMR reads nuclear spin no direct sequencing route main route minor/emerging
DNARNAProteinMetaboliteModificationFocal detail
Plate 05 — Every omics method reads a physical proxy — base-pairing, mass, shape-complementarity, light, or travel time — never the biology directly. Layers that cannot be sequenced lean on mass and binding, and inherit their limits.

Real methods combine the principles

Almost every method in practical use chains two or more of these principles. The pairing matters because each principle covers the blind spot of another.

MethodPrinciples combinedWhat the combination buys
LC-MS proteomics and metabolomicsseparate + weighfewer molecules reach the detector at once; retention time adds identity
Flow cytometrybind + seea specific label, read one cell at a time
Proximity extension assaybind + sequenceprotein recognition, with DNA's copying and counting
Single-cell sequencingsequence, after barcoding each cellevery molecule carries the address of its cell
Spatial imagingbind (by pairing) + seemolecule identity with its position in the tissue
CITE-seqbind + sequencesurface proteins and RNA from the same cell

A pattern runs down that table. Whenever a molecule lacks the two nucleic-acid tricks, methods try to borrow them, most often by attaching a DNA tag to a binder so that the measurement ends as a count of DNA. Much of the progress in protein and cell measurement over the last decade has come from that one move.

A proxy is not the thing

Each principle measures a property that correlates with the biology, not the biology itself. A fluorescence signal reports bound dye, not protein; an ion count reports molecules that ionized well, not molecules that were abundant; a sequencing read reports a molecule that survived library preparation. Every quantitative claim in omics rests on the assumption that the proxy tracks the quantity of interest, and much of good experimental design is checking that assumption.

In practice

Reading a methods section, reduce each technique to its principles before judging it. A protein assay that ends in sequencing inherits the dynamic range of counting DNA and the specificity of its antibodies. A particle count from a scatter-based instrument inherits that instrument's size limit. Stating the principles in a specification or a technical file also makes the verification plan obvious: each principle has a known failure mode, and each failure mode needs a test.

+ What this chapter established
  • Every omics method reads one of five physical properties: sequence, mass, binding, light, or travel time.
  • Mass is universal but favors the abundant; binding reaches deep but only for predefined targets.
  • Counting single objects is more robust than estimating the intensity of a mixture.
  • Practical methods chain principles, and many convert non-nucleic-acid targets into DNA to borrow its advantages.

+ Part II · The nucleic-acid layers

Reading by copying and pairing.

Genome, epigenome and transcriptome share one advantage: their molecules can be copied and read by pairing. These four chapters show how sequencing exploits that, how three generations of instruments answer three different questions, and how chemistry turns marks and activity into letters a sequencer can read.

04 — Genomics

Reading a genome in fragments.

+ The questionHow do you read three billion letters when no instrument can read a chromosome from end to end?

Shatter, read, reassemble

The largest human chromosome is about 250 million base pairs long. No sequencer reads anything close to that in one pass: the most common instruments read a few hundred bases at a time, and even the long-read machines described in Chapter 5 typically read tens of thousands. So every genome is read the same way, by breaking it into pieces, reading each piece, and putting the pieces back in order with software.

The laboratory half of that process is called library preparation. DNA is extracted from cells, fragmented, and fitted with short synthetic adapter sequences at both ends. The adapters are what the instrument grips and where the reading starts, so every fragment, whatever its origin, presents the same handle. The result is a library: hundreds of millions of fragments, each a random sample of the genome, each ready to be read.

The computational half turns reads back into a genome. Most human sequencing uses alignment: each read is matched to the position it best fits in a reference genome, the agreed standard sequence that acts as a coordinate system. Where no suitable reference exists, reads are assembled from scratch by finding overlaps between them, the way a torn page can be reconstructed from overlapping scraps. Both approaches rely on the reads being long enough, and unique enough, to find their place.

Every base is read many times

A single read is not trusted, for two reasons. The first is error. Sequencers report a quality score for every base, on a logarithmic scale: Q30 means an estimated one chance in a thousand that the base is wrong, Q20 one in a hundred. Even at Q30, a genome read once would contain millions of errors. The second is sampling. Fragments land on the genome at random, so some positions are covered many times and some not at all, the same way raindrops leave some paving stones soaked and others dry.

The answer to both is coverage, the average number of reads covering each position. A standard clinical or research genome is sequenced to about 30-fold coverage, written 30×.

Worked example: what 30× coverage means in reads

Haploid human genome: about 3.1 billion bases. At 30× coverage: 3.1 × 10⁹ × 30 ≈ 9.3 × 10¹⁰ bases to read.

With reads of 150 bases: 9.3 × 10¹⁰ ÷ 150 ≈ 620 million reads, usually produced as about 310 million pairs read from both ends of each fragment.

At 30×, a position read by, say, 28 reads with 14 showing one letter and 14 another is a confident call. A position read by 4 reads is not, which is why coverage is reported as a distribution, not just an average.

What a variant looks like in the data

Once reads are stacked against the reference, differences stand out, and their pattern says what they are. A heterozygous variant, one present on only one of the two parental copies, appears in roughly half the reads at that position. A homozygous variant appears in nearly all of them. A random sequencing error usually appears in one read and nowhere else, because such errors rarely repeat at the same position across independent fragments. Systematic errors can recur, for example in runs of a single base, in duplicated fragments or in reads misplaced in repeats, which is why variant callers also weigh strand balance, mapping quality and duplicates.

That is the logic of variant calling: a variant is a difference that recurs in many independent reads at the same position, at a proportion consistent with one or two copies. Software weighs the number of supporting reads, their base qualities and how confidently each read was placed, and reports each call with a confidence of its own. The output of genome sequencing is therefore not a text read from start to finish. It is a list of differences from the reference, each with a measure of how sure the software is.

06 — Reading a genome in fragments
REFERENCE G A T T A C A G C T T A G C C T A G G C A T C A READS T C T C T C G error: seen once 7× 0 COVERAGE coverage 30× across a whole genome ≈ 30 reads over every base variant: seen in ~half the reads at the same position
DNAReadoutFocal detail
Plate 06 — No instrument reads a chromosome end to end. The genome is shattered, each fragment is read, and the reads are stacked against a reference. A true variant recurs in many reads at the same position; a random error appears once.
A genome is a list of differences, not a transcript

It is tempting to picture a sequenced genome as three billion letters read in order. What is actually delivered is a set of variant calls relative to a reference, each with a quality, plus regions where no call could be made because coverage was too thin or the reads could not be placed. Two laboratories using different references, software or thresholds can report different variant lists from the same DNA. The list is only interpretable alongside those choices.

The reference problem

A reference genome is a coordinate system, and its gaps become blind spots. The reference used through most of the last two decades was missing about 8% of the genome, mostly long repetitive regions that short reads could not assemble. The Telomere-to-Telomere consortium published the first gapless human sequence in 2022, 3.055 billion base pairs, adding roughly 200 million bases and nearly 2,000 predicted genes; the Y chromosome followed in 2023.

A single reference also comes from very few people, so reads carrying DNA that the reference lacks either fail to align or align to the wrong place. The Human Pangenome Reference Consortium is building a reference from many people instead: 47 individuals in its first release in 2023 and 232 in a second data release in 2025. Diversity in reference data has practical consequences for Indian readers in particular. The GenomeIndia project released whole genomes of 10,074 healthy, unrelated people from 83 Indian population groups in January 2025, because variant frequencies measured mainly in European-ancestry cohorts do not transfer reliably to South Asian populations.

Variant size decides what you can see

Variants come in sizes spanning more than eight orders of magnitude. A single-nucleotide variant changes one base. Small insertions and deletions change up to about 50. Structural variants, meaning deletions, duplications, inversions and insertions of more than 50 bases, can span thousands to millions. Copy-number changes and whole-chromosome gains or losses are larger still.

Short reads handle the smallest and the largest variants well. A single-base change sits comfortably inside a 150-base read, and a whole extra chromosome shows up as a proportional rise in coverage. The middle is hard. A structural variant several thousand bases long cannot be seen inside any one short read; it has to be inferred from indirect clues such as reads that align oddly or a change in coverage, and in repetitive regions the clues are often missing. Much of the reason long-read sequencing exists is to read across that middle range directly.

07 — Variant size against read length
VARIANTS 1 bp SNV 2–50 bp small indel deletion, duplication, inversion, insertion structural variants · 50 bp – 1 Mb >1 Mb to whole chromosome aneuploidy large copy-number change / 1 bp 10 bp 100 bp 1 kb 10 kb 100 kb 1 Mb 10 Mb READS short read · 100–300 bp long read (HiFi) · 15–20 kb typical ultra-long nanopore reads >100 kb (record ≈ 4 Mb) longer than a short read: detected indirectly, often missed
DNAReadoutFocal detail
Plate 07 — A variant is easiest to see when a single read spans it. Structural variants — thousands of bases long — sit beyond the reach of short reads, which is much of the reason long-read sequencing exists.

Genome, exome or panel

Not every question needs the whole genome. Only about 1–2% of it encodes protein, and sequencing just those regions, the exome, costs less and concentrates reads where most known disease-causing variants sit. Targeted panels go further, reading tens to hundreds of chosen genes at very high coverage, often 500× to more than 1,000×. High coverage is what lets a tumor panel find a mutation present in only a few percent of the cells in a biopsy. The trade is the usual one in omics: breadth against depth. A panel sees deeply into what it was designed for and nothing else; a genome sees everything, less deeply.

In practice

A sequencing request or specification should state, at minimum: the question and the variant classes that matter; genome, exome or panel; coverage, as a minimum across target regions rather than an average; read length; the reference build; and acceptance criteria for quality such as the fraction of bases at Q30 or above, duplicate rate and mapping rate. A result delivered without those parameters cannot be compared with any other result, including a repeat of itself.

+ What this chapter established
  • Genomes are read by shattering DNA into fragments, reading each, and aligning the reads to a reference.
  • Every position is read many times; coverage defeats both random errors and uneven sampling.
  • A variant recurs across independent reads at one position; an error appears once.
  • Structural variants longer than a short read, from a few hundred to tens of thousands of bases, are hardest for short reads, which motivates long reads.
05 — Three generations of sequencing

Three generations, three questions.

+ The questionIf newer sequencers are faster and cheaper, why do laboratories still run three different kinds?

What stayed constant

Every sequencer in routine use reads DNA through the same underlying fact: a base is identified by what it pairs with, or by the distinctive signal it produces as it passes a detector. Most machines copy the template strand one base at a time and detect each addition; nanopore machines detect each group of bases as it moves through a pore. What changed across three generations of instruments is not that principle. It is how many molecules are read at once, whether they are copied before reading, and how long a stretch of one molecule can be read in one go. Each change answered a sharper question, and that is the clearest way to understand them.

Generation one: what is this sequence?

Frederick Sanger's chain-termination method, published in 1977, copies a single purified template in the presence of a small amount of modified bases that stop the copying wherever they are incorporated. The result is a set of copies ending at every possible position, each ending in a base carrying its own dye. Separated by length in a thin capillary, the fragments pass a detector in order, and the colors read out the sequence directly.

Sanger sequencing reads up to about a thousand bases with very high accuracy, but one template per capillary. It answers one question well: what is the sequence of this particular piece of DNA? The Human Genome Project was built largely on it, at a cost of billions of dollars over more than a decade. It is still the standard way to confirm a single variant found by other methods.

Generation two: what are all the sequences, and how many of each?

The step change came from parallelism. Short-read instruments attach millions to billions of library fragments to a glass surface, copy each fragment in place into a small cluster of identical molecules, and then read every cluster simultaneously. Each chemical cycle adds one dye-labeled base to every growing strand and blocks further addition; a camera images the whole surface; the block is removed; the next cycle begins. After 100 to 300 cycles, every cluster has produced a read.

The individual read is short and the chemistry has an error rate per base, but the output is enormous: current high-throughput instruments produce several terabases per run, with most bases at Q30 or better. Parallelism did two things. It collapsed the cost, and it turned sequencing into a counting instrument. When you read hundreds of millions of fragments, the number of reads that come from each gene, each allele or each organism is itself a measurement. RNA sequencing, single-cell methods and liquid biopsy all depend on that counting property more than on reading any one sequence.

Generation three: what is the whole molecule?

Single-molecule instruments skip the copying step. One approach watches a single polymerase at the bottom of a tiny well as it builds a partner strand, recording each incorporated base as a flash of light; reading the same circular molecule many times over gives consensus reads, typically 15,000 to 20,000 bases long, with accuracy that the manufacturer states at Q30 and above. The other approach pulls a single strand through a protein pore in a membrane and reads the changing electrical current as bases pass; its reads can exceed 100,000 bases, the longest reported run to millions, and the manufacturer reports a modal raw single-read accuracy of about 99.75%.

Reading a native molecule has a second benefit that is easy to miss. Because nothing is copied, chemical marks on the original DNA are still there to be detected. Both approaches report methylated cytosines directly from the signal, so a long read can carry sequence and epigenetic state together. The question this generation answers is the one short reads struggle with: what does this whole stretch of molecule look like, in one piece, across repeats and structural variants, marks included?

08 — Three generations, three questions
WHAT IS THIS SEQUENCE? Sanger (chain termination) WHAT ARE ALL THE SEQUENCES? Short-read (sequencing by synthesis) WHAT IS THE WHOLE MOLECULE? Long-read (single molecule) template one template per capillary · up to ~1,000 bases millions of clusters, each a copy-cloud of one fragment parallelism is what collapsed the cost cycle 1 cycle 2 cycle 3 one base per cycle, every cluster at once · 100–300 bases methyl marks current one native molecule, no copying · 15 kb to >100 kb · reads methylation directly
DNAModificationReadoutFocal detail
Plate 08 — The chemistry of reading DNA did not change as much as the question did. Sanger reads one sequence well; short-read machines read millions at once; long-read machines read one native molecule end to end, marks included.

The evolution at a glance

GenerationWorking principleQuestion it answersTypical readStrengthBlind spot
Sangerchain termination, one template per capillaryWhat is this sequence?up to ~1,000 basesaccuracy on a single targetthroughput, cost per base
Short readcluster amplification, one base per cycle, all clusters at onceWhat are all the sequences, and how many of each?100–300 basescost, counting, scalerepeats, structural variants, phasing
Long readsingle native molecule, real-time optics or nanopore currentWhat is the whole molecule?15 kb to >100 kbstructure, phasing, native methylationcost per base, historically raw accuracy

What the cost curve shows

The practical consequence of parallelism is visible in the cost of a genome. The US National Human Genome Research Institute tracked the full cost of sequencing a human genome, including labor, equipment and overheads, from about $95 million in 2001 to about $525 in 2022. The fall accelerated abruptly in 2008, when sequencing centers switched from capillary instruments to massively parallel ones. Since then, manufacturers have announced reagent costs of $80 to $350 per genome on their newest systems. Those claims measure reagents on a fully loaded instrument and are not comparable with the institute's all-in figure, but the direction is not in doubt.

09 — Cost of sequencing one human genome
$10 $100 $1k $10k $100k $1M $10M $100M 2001 2005 2010 2015 2020 2025 centers switch to massively parallel sequencing (NHGRI) Moore's-law pace $95.3M · Sep 2001 $525 $200 (Illumina) $80 (Ultima) $150 (Roche) $345 (PacBio HiFi) NHGRI all-in cost per genome vendor reagent-cost claims vendor claims are reagents at scale — not comparable with NHGRI's all-in figure ≈181,000-fold fall vs ≈1,450-fold at Moore's-law pace (2001→2022)
ReadoutFocal detail
Plate 09 — Between 2001 and 2022 the all-in cost of a genome fell about 181,000-fold, some 125 times more than Moore's-law pace would have produced. The break came in 2008, when sequencing went massively parallel.
Worked example: faster than Moore's law

All-in cost: $95.3 million in 2001, $525 in 2022. Fold reduction = 95,300,000 ÷ 525 ≈ 181,000.

Moore's-law pace, halving every two years over 21 years: 2^(21 ÷ 2) = 2^10.5 ≈ 1,450-fold. And 181,000 ÷ 1,450 ≈ 125.

The total fall was about 125 times larger than Moore's-law pace would have produced: a halving roughly every 14 months instead of every 24.

Newer did not replace older

Each generation is a tool matched to a different question, not a strict successor. Sanger still confirms single variants. Short reads remain the cheapest way to count molecules and to find single-base variants across many samples. Long reads resolve structural variants, repeats and phasing, and read methylation without conversion. New chemistries keep arriving, including a sequencing-by-expansion method that reached the market in 2026, and hybrid approaches that add long-range information to short reads. The right choice follows from the question.

In practice

Choose the generation from the question, not from the newest brochure. Confirming a known variant: Sanger. Counting molecules across many samples, such as expression, allele fraction or pathogen load: short read. Resolving structure, phasing or repetitive regions, or reading methylation directly: long read. When a supplier quotes a cost per genome, ask what it includes: reagents only, or instrument time, labor, analysis and storage.

+ What this chapter established
  • All three generations read base order; they differ in parallelism, copying and read length.
  • Parallel short reads collapsed the cost and made sequencing a counting instrument.
  • Long single-molecule reads resolve structure and read methylation without conversion.
  • The all-in cost of a genome fell about 181,000-fold from 2001 to 2022.
06 — Epigenomics

Memory without changing the letters.

+ The questionIf the DNA sequence never changes, how does a cell remember what it is?

Identity has to survive division

When a liver cell divides, both daughters are liver cells. The sequence was copied, but the sequence is the same in a neuron, so something else must have been copied too: the record of which genes this cell keeps available and which it keeps silent. That record is the epigenome, from the Greek for "on top of" the genome. It consists of chemical marks and physical packaging that sit on the DNA without changing its letters, and that a dividing cell passes to its daughters.

Three kinds of record matter most for measurement. The first is DNA methylation, a small methyl group attached to cytosine, in most cell types almost always where a C is followed by a G, a site written CpG. Neurons are a notable exception, carrying substantial methylation at other sites too. The human genome has about 28 million of them. Methylation is copied through division by an enzyme that reads the mark on the old strand and adds it to the new one, which is what makes it a memory. The second is chromatin accessibility: DNA is wound around protein spools called nucleosomes, and a gene wound tightly is hard for the transcription machinery to reach. The third is histone modification, chemical marks on the spool proteins themselves that recruit or repel the machinery. Together they decide which parts of the genome a cell can read.

Methylation: turning a mark into a letter

A sequencer that copies DNA cannot see methylation, because the copying enzyme puts an ordinary C opposite a methylated one. The mark is lost at the first copy. So the classic methods convert the question into one a sequencer can answer. Bisulfite treatment chemically converts every unmethylated cytosine into uracil, which is copied as T, while methylated cytosines are protected and stay C. After sequencing, each CpG in each read shows either C, meaning it was methylated, or T, meaning it was not. One limit is built in: standard bisulfite and enzymatic conversion both protect 5-hydroxymethylcytosine as well, so they cannot tell it from 5-methylcytosine without an extra chemical step.

Bisulfite chemistry is harsh and fragments much of the DNA, which matters when the starting material is scarce, as with the cell-free DNA in blood described in Chapter 14. Enzymatic conversion methods do the same job with enzymes instead, with less damage. Newer approaches read methylation alongside the ordinary sequence in one library, and long-read instruments detect methylated bases directly from the native molecule without any conversion.

10 — Turning a chemical mark into a letter
REFERENCE 5′ 3′ A C G T T C G A C G T C A AFTER CONVERSION 5′ 3′ A U G T T C G A U G T U A READ AS 5′ 3′ A T G T T C G A T G T T A unmethylated C → U U reads as T methylated C resists conversion 10 reads at the marked CpG 7 read C, 3 read T → 70% methylated bisulfite (chemical) or enzymatic conversion — same logic
DNAModificationFocal detail
Plate 10 — Sequencers cannot see a methyl group, so the chemistry converts the question into one they can answer. Every unmethylated C is turned into a letter that reads as T; a C that survives was methylated, and the share of reads keeping it gives the methylation level.

The output of any of these methods is a methylation level: at each CpG, the fraction of reads that carried the mark. That fraction is a sampling estimate, and its precision depends on how many reads covered the site.

Worked example: why depth matters for a methylation level

At one CpG, 7 of 10 reads read C. Methylation level = 7 ÷ 10 = 70%.

The standard error of a proportion is √(p × (1 − p) ÷ n) = √(0.7 × 0.3 ÷ 10) ≈ 0.145, so the estimate is roughly 70% ± 28 points at 95% confidence, and at so few reads even that simple interval is only a rough guide.

With 100 reads the same calculation gives about ± 9 points. A difference between two samples of 60% and 70% methylation is invisible at 10 reads and detectable at a few hundred.

A methylation level is a population average

A CpG measured at 50% in a tissue can mean every cell has one of its two copies methylated, or that half the cells are fully methylated and half are not, or any mixture in between. Bulk measurements cannot tell these apart, and the difference is often the biology of interest, for example when a small population of tumor cells carries a pattern the surrounding tissue does not. Reading single molecules, as long reads do, or single cells, as Chapter 12 describes, is how the ambiguity is resolved.

Accessibility: an enzyme does the measuring

Accessibility is measured by asking an enzyme to find it. ATAC-seq, introduced in 2013, uses an engineered transposase, an enzyme that inserts DNA, loaded with sequencing adapters. Added to cell nuclei, it can reach only DNA that is not wrapped around nucleosomes or covered by other proteins. Wherever it can reach, it cuts the DNA and inserts adapters in the same step. Sequencing the resulting fragments and stacking them on the genome produces peaks where the chromatin was open: active promoters, enhancers and other regulatory switches.

This is a good example of the measuring-by-proxy principle from Chapter 3. Nobody observes chromatin opening. What is counted is where an enzyme was able to act, and the inference is that those regions were available to the cell's own machinery too.

11 — Reading which DNA is open
nucleosomes: DNA wrapped on histones Tn5 cuts exposed DNA and inserts adapters Tn5 cannot reach wrapped DNA open region (e.g. an active promoter) READ COVERAGE reads pile up where DNA was accessible position along the genome accessible = available to be read by the cell
DNAReadoutFocal detail
Plate 11 — An enzyme that can only cut exposed DNA does the measuring. Where chromatin is open it inserts sequencing adapters; wrapped DNA is protected. Reading the fragments back maps which regions of the genome the cell has made available.

Marks on the spools, and the fold

Histone modifications are read with antibodies. In ChIP-seq, chromatin is cross-linked, fragmented, and pulled down with an antibody against a particular mark, and the recovered DNA is sequenced. CUT&Tag, published in 2019, instead guides the adapter-inserting enzyme to the mark with the antibody, working from far fewer cells. Both inherit the strengths and weaknesses of binding: they find exactly the mark the antibody recognizes, and they are only as specific as the antibody.

The genome also has a three-dimensional shape, and regulatory switches often act on genes far away along the sequence but close in space. Chromosome conformation methods such as Hi-C, introduced in 2009, cross-link DNA that is touching in the nucleus, join the touching ends, and sequence the junctions. The frequency of each junction reports how often two distant regions sit together.

Why methylation reaches beyond the cell

Methylation patterns differ between tissues and change in disease, and they survive in DNA fragments after cells die. That makes them useful far from the cell they came from. A fragment of DNA in the blood carries the methylation pattern of the tissue that released it, which is the basis of methods that estimate the tissue of origin of cell-free DNA and of blood tests that look for the methylation signatures of cancer, described in Chapter 14. Methylation at certain sites also changes steadily with age, and statistical models built on those sites, often called epigenetic clocks, estimate age from a blood sample. They estimate; they do not measure a biological age directly, and different clocks give different answers for the same person.

In practice

Conversion-based methods need proof that conversion worked. Specify a control, typically unmethylated DNA spiked into every sample, and an acceptance criterion such as conversion above 99%, because incomplete conversion reads as false methylation. For low-input material such as cell-free DNA, weigh enzymatic conversion or direct long-read detection against bisulfite, and state the minimum read depth per CpG that the decision needs.

+ What this chapter established
  • The epigenome records which parts of the genome a cell keeps available, and is copied through division.
  • Methylation is invisible to copying, so methods convert it into a letter or read it from the native molecule.
  • Accessibility is measured by where an enzyme can reach; histone marks by where an antibody binds.
  • A methylation level is a population fraction whose precision depends on read depth.
07 — Transcriptomics

Counting what the cell is reading.

+ The questionKnowing a gene is present, how do we know the cell is using it, and how much?

A measurement that is really a count

A gene in use is transcribed into RNA, and the more intensely it is used, the more copies of its messenger RNA the cell holds at any moment. Measuring gene activity therefore means measuring how many copies of each RNA are present, a steady-state level that reflects both how fast each RNA is made and how fast it is degraded. A typical mammalian cell holds on the order of a few hundred thousand messenger RNA molecules, spread across roughly ten thousand genes that are active in that cell type, with some genes represented by thousands of copies and others by one or two.

RNA sequencing, or RNA-seq, measures this by counting. RNA is extracted from a sample. Most of it is ribosomal RNA, the structural RNA of the protein-making machinery, which carries no information about gene activity, so it is removed, or messenger RNA is selectively captured by the string of A bases on its tail. The remaining RNA is copied into cDNA by reverse transcriptase, fragmented, fitted with adapters and sequenced exactly as a genome library would be. The reads are aligned to the genome or to a catalog of known transcripts, and the number of reads landing on each gene is tallied.

12 — From tissue to a count table
tissue RNA AAAA AAAA AAAA cDNA reverse transcription: RNA copied into DNA so it can be amplified and sequenced fragments + reads map to genes exon intron one read is split across an exon junction count table S1 S2 S3 GENE A 1240 1310 1185 B 12 8 15 C 530 610 488 D 0 3 1 E 88 42 97 reads per gene ≈ how much of that gene was being expressed schematic illustrative counts
DNARNAReadoutFocal detail
Plate 12 — RNA cannot be sequenced directly on most machines, so it is copied into DNA first. What comes out is not a measurement of any one molecule but a count: how many reads landed on each gene, which rises with the gene's steady-state RNA level and with its length.

The result is a count table: genes down one side, samples across the top, and in each cell the number of reads. It is worth being precise about what that number is. It is not a count of molecules. It is a count of sequenced fragments, which is proportional to the number of molecules, multiplied by the length of the transcript, multiplied by how efficiently that transcript survived capture and copying, and divided among all the other transcripts competing for the same fixed number of reads.

Why raw counts mislead

Two of those factors distort comparisons unless they are corrected. The first is sequencing depth. If one sample receives 40 million reads and another 20 million, every gene in the first will have roughly twice the count for reasons unrelated to biology. The second is transcript length. A longer transcript is cut into more fragments, so at the same number of molecules it collects more reads.

Worked example: normalizing for length and depth

Two genes in one sample. Gene A is 1,000 bases long and has 500 reads. Gene B is 4,000 bases long and has 1,000 reads. On raw counts B looks twice as active.

Divide by length in kilobases: A = 500 ÷ 1 = 500 reads per kb; B = 1,000 ÷ 4 = 250 reads per kb. Per molecule, A is present at twice B's level.

Then scale so that each sample's length-adjusted values sum to one million. The result, transcripts per million or TPM, can be compared between samples of different depth. For comparisons of the same gene across samples, statistical packages use related methods that also guard against a few very highly expressed genes absorbing a large share of the reads.

The third factor, composition, is subtler. Because the total number of reads is fixed, a sample in which one gene becomes enormously abundant leaves fewer reads for everything else. Every other gene then appears to fall, although nothing changed for it. The normalization methods used for differential expression estimate a scaling factor from the genes that did not change, precisely to avoid this. Chapter 13 meets the same problem in a different guise.

More reads is not more activity

A higher count for one gene than another says little on its own: the longer gene collects more reads, and a deeper sample collects more of everything. Counts become comparable only after normalization, and even then the comparison that holds up best is the same gene across samples under the same protocol, not different genes within one sample.

How many reads is enough

Depth decides which genes can be measured at all. The ENCODE consortium's standard for bulk RNA-seq is about 30 million aligned reads per replicate. At that depth abundant transcripts are counted precisely and rare ones are sampled sparsely: a gene that yields 5 reads in one sample and 10 in another has not necessarily doubled, because counts that small fluctuate by chance by about the square root of their value. Detecting modest changes in low-abundance genes needs more reads, more biological replicates, or both. Replicates matter more than depth in most designs, because the variation between individuals or cultures is usually larger than the counting noise.

One gene, many transcripts

Most human genes can be spliced in more than one way, joining different subsets of their coding segments, or exons, into different messenger RNAs that may encode different proteins. Short reads see only fragments of each transcript, so the relative amounts of each version are inferred statistically from reads that happen to span a particular junction. Long-read sequencing can read a whole transcript in one pass, from start to tail. Library methods that join several cDNA molecules end to end before long-read sequencing now make full-length transcript sequencing practical at scale.

Before sequencing: arrays and targeted assays

Before RNA-seq, gene expression was measured with microarrays: glass slides carrying thousands of fixed DNA probes, each designed to pair with one gene's RNA. Labeled cDNA from the sample hybridizes to its probes, and the brightness of each spot reports abundance. Arrays exploit pairing, just as sequencing does, but they read intensity rather than counting, and intensity has a floor set by background and a ceiling set by saturation. Arrays also measure only what their probes were designed for. RNA-seq counts, which extends the range, and it reads sequence, which finds transcripts nobody knew to look for.

Targeted assays remain important, especially in the clinic. Quantitative PCR measures a handful of chosen RNAs with high precision, and several established multigene tests that guide cancer treatment decisions are built on it. The general pattern recurs across omics: discovery uses the broad method, and a decision that will be made repeatedly uses a narrow, validated one.

What RNA can and cannot tell you

A transcriptome is a snapshot of the last few hours of the cell's decisions, averaged over every cell in the sample. Both limits matter. A tissue is a mixture, and a change in the bulk signal can mean that every cell changed or that the mixture of cell types changed; Chapter 12 describes how single-cell methods separate those. And RNA is not protein: the relationship between the two is imperfect, for reasons Chapter 15 sets out.

In practice

Put RNA quality and design into the specification before the first sample is collected. Set a minimum RNA integrity value as an acceptance criterion, because degraded RNA biases counts toward the ends of transcripts. Plan at least three biological replicates per condition, and more when the effect sought is small. Balance conditions across processing batches, a point Chapter 17 returns to. If the result will drive a repeated decision, plan a targeted confirmation assay from the start.

+ What this chapter established
  • RNA-seq measures gene activity by counting sequenced fragments per gene.
  • Counts must be normalized for sequencing depth, transcript length and composition before comparison.
  • Replicates and depth together decide which changes can be detected.
  • The transcriptome is an average over cells and hours, and it does not stand in for the proteome.

+ Part III · Beyond nucleic acids

Measuring what cannot be copied.

Proteins and metabolites are where the cell's decisions become action, and neither can be copied or read by pairing. These three chapters show the two routes around that limit, weighing molecules and binding them, and why identifying a small molecule is harder than detecting it.

08 — Proteomics by mass

Weighing the workforce.

+ The questionIf proteins cannot be copied or paired, how does a mass spectrometer identify thousands of them in one run?

The layer that acts

Proteins do most of the work that makes a cell what it is: they catalyze its chemistry, carry its signals, build its structure and defend it. The same gene can give rise to several different protein molecules, because RNA can be spliced in different ways and because finished proteins are cut, folded and decorated with chemical groups such as phosphates and sugars. Each distinct molecular form is called a proteoform. The proteome is therefore larger and more varied than the list of about 20,000 protein-coding genes suggests, and many of the differences that matter, such as whether an enzyme has been switched on by a phosphate, exist only at this layer.

Proteomics has made steady progress on coverage. The Human Proteome Project's 2025 report found credible protein-level evidence for about 93.6% of the 19,435 proteins predicted from the genome. The challenge is no longer whether a protein can be detected somewhere, but how many can be measured in one sample, how precisely, and how cheaply.

How a molecule is weighed

A mass spectrometer measures the ratio of a molecule's mass to its electric charge. The molecule first has to become a charged particle in a vacuum. For proteins and peptides this is almost always done by electrospray: the liquid emerging from a chromatography column is sprayed from a fine needle held at high voltage, the charged droplets shrink as the solvent evaporates, and the dissolved molecules are left carrying one or more extra protons.

The ions are then sorted in electric or magnetic fields. Different instruments do it differently. A quadrupole passes only ions of a selected m/z through an oscillating field, like a tunable filter. A time-of-flight analyzer gives every ion the same push and times its arrival, since lighter ions fly faster. An orbital trap lets ions circle a central electrode and measures how fast they oscillate, which depends on m/z. The best instruments measure m/z to within a few parts per million, enough to distinguish molecules whose masses differ in the third decimal place.

Why proteins are cut before they are weighed

Most proteomics is done bottom-up: proteins are first digested into peptides with trypsin, an enzyme that cuts after the amino acids lysine (K) and arginine (R). There are good reasons. Whole proteins are large, carry many charges, and break unpredictably. Peptides of roughly 7 to 30 amino acids ionize well, separate well by chromatography, and break in predictable places, so their sequences can be read from their fragments.

The peptides are separated by liquid chromatography over a gradient lasting tens of minutes to hours, so that at any moment only a fraction of them enter the instrument. For each slice of time the instrument records a first spectrum, MS1, of all the peptide ions present. It then selects ions, breaks them by collision with gas, and records the masses of the fragments in a second spectrum, MS2. Tens of thousands of MS2 spectra are recorded per run.

13 — Bottom-up proteomics
protein K R K trypsin cuts after K and R peptides time LC column: peptides separate over a gradient + + + electrospray: charged droplets ions enter the mass spectrometer m/z MS1: measure m/z of every peptide ion select one, fragment it m/z MS2: fragment spectrum match against predicted spectra from a protein database LVNELTEFAK · BSA next candidate · lower score decoy sequence · lower score identification is a statistical match, not a direct reading schematic spectra
ProteinReadoutFocal detail
Plate 13 — A mass spectrometer never sees a protein whole in this workflow. Proteins are cut into peptides, the peptides are weighed and broken, and the fragment masses are matched against what every known protein would produce. The protein list is an inference from those matches.

Reading a peptide from its fragments

A peptide breaks mostly along its backbone, between one amino acid and the next. Each break produces two pieces, one containing the start of the peptide and one containing the end, and across many copies of the peptide the breaks occur at every position. The pieces containing the end, called y ions, form a ladder: y1 is the last amino acid, y2 the last two, and so on. The difference in mass between consecutive rungs is the mass of one amino acid.

Worked example: reading an albumin peptide

The tryptic peptide LVNELTEFAK, from bovine serum albumin, a protein widely used as an instrument standard, carries two charges and appears in MS1 at m/z 582.32.

In its MS2 spectrum the y ions fall at 147.11 (y1), 218.15 (y2) and 365.22 (y3). The gap from y1 to y2 is 71.04, the mass of alanine (A). The gap from y2 to y3 is 147.07, the mass of phenylalanine (F). Continuing up the ladder reads E, T, L, E, N, V: the sequence, from the end backwards.

One limit is built in. Leucine and isoleucine have the same mass, 113.08, so this ladder cannot tell them apart without additional evidence.

14 — Reading a peptide from its fragments
0 200 400 600 800 1000 m/z y1 y2 y3 y4 y5 y6 y7 y8 y9 A 71.04 F 147.07 E 129.04 T 101.05 L 113.08 E 129.04 N 114.04 V 99.07 each gap is one amino acid's mass precursor [M+2H]2+ m/z 582.3 peak heights schematic read right to left: …V N E L T E F A K BSA standard peptide LVNELTEFAK
ProteinReadoutFocal detail
Plate 14 — When a peptide breaks along its backbone, the fragments form a ladder. The spacing between rungs is the mass of one amino acid, so the sequence can be read off the gaps — here LVNELTEFAK, a peptide of bovine serum albumin used as an instrument standard.

From spectra to a protein list

In practice the ladder is not read by eye. Search software takes every protein sequence in a database, digests it in silico, predicts the fragment spectrum of every resulting peptide, and scores each observed spectrum against the candidates. The best-scoring match is reported if its score clears a threshold.

That threshold is set statistically. The search is run against the real database and against a decoy database of reversed or scrambled sequences that cannot be present; the number of decoy matches above a score estimates how many of the real matches are false. Results are typically reported at a 1% false discovery rate for peptides and again for proteins.

A further step, protein inference, turns peptides into proteins, and it is the weakest link. Many peptides are shared between related proteins or between proteoforms of one gene, so a set of peptides often cannot say which member of a family was present. This is why results list protein groups rather than proteins.

A protein list is an inference, not an inventory

A proteomics result lists protein groups inferred from peptide matches above a statistical threshold. A protein missing from the list was not detected in that run; it was not necessarily absent. Shared peptides make some identities ambiguous, and a protein identified from a single peptide deserves less confidence than one supported by many.

Choosing what to fragment

The instrument cannot fragment every ion at once, so it has to choose. In data-dependent acquisition it selects the most intense ions in each MS1 scan. That choice is partly random, so a peptide measured in one run may be skipped in the next, leaving gaps in the data. In data-independent acquisition it fragments everything within successive m/z windows and relies on software to untangle the mixed fragment spectra. The result is more complete and more reproducible across samples. With current instruments and narrow-window acquisition, one 2024 study measured about 10,000 human protein groups from a cell digest in a 30-minute run.

Plasma is harder, for the dynamic-range reason Chapter 2 set out. Neat plasma yields roughly 700 to 1,100 protein groups per sample on current instruments; removing the most abundant proteins, or enriching lower-abundance ones on nanoparticle surfaces, roughly doubles that in independent comparisons, and developers of the newest enrichment workflows report several thousand. The next chapter describes the other route to the faint end.

How much, and in what form

Mass spectrometry measures ion signal, and different peptides ionize with different efficiency, so signal is not directly a quantity. Relative quantity, the same peptide compared across samples, is measured from signal intensity or by labeling each sample with an isotope tag and measuring the samples together. Some tags differ in mass and are told apart in the first spectrum; isobaric tags weigh the same and are told apart by reporter ions released on fragmentation. Absolute quantity requires a known amount of an isotope-labeled copy of the peptide added to the sample as an internal standard. That is how clinical mass spectrometry assays are built: narrow, targeted, and calibrated.

Modifications can be studied directly, because a phosphate or a sugar adds a known mass. Phosphoproteomics enriches phosphorylated peptides before measurement to map which signaling switches are on. Top-down proteomics, which weighs and fragments intact proteins, measures proteoforms whole, at lower throughput.

In practice

A proteomics deliverable should state the acquisition mode, the false discovery rate at peptide and protein level, how protein groups were formed, and how missing values were handled, since the choice changes downstream statistics. In biopharmaceutical work, mass spectrometry of host-cell proteins complements immunoassays because it identifies each contaminant rather than reporting a total. Where a number will support a release decision, it needs a targeted, calibrated method with defined limits of quantitation.

+ What this chapter established
  • A mass spectrometer measures mass-to-charge; identity comes from weighing fragments as well as the whole.
  • Proteins are digested into peptides, whose fragment ladders spell out their sequence.
  • Identification is a statistical match controlled by decoys; protein inference is the weakest step.
  • Signal is not quantity: absolute amounts need isotope-labeled internal standards.
09 — Proteomics by binding

Reaching the faint end.

+ The questionWhere mass spectrometry cannot see the rare proteins in blood, how do affinity methods reach them, and what do they give up?

One binder, one protein

The oldest way to measure a single protein at low concentration is the sandwich immunoassay, often called an ELISA. One antibody fixed to a surface captures the protein; a second antibody, recognizing a different part of it, binds on top; the second antibody carries an enzyme that converts a colorless substrate into a colored or light-emitting product. Two features make it work. Requiring two antibodies to bind the same molecule makes the signal specific, because a molecule that happens to stick to one antibody rarely also binds the other. And the enzyme amplifies: each bound enzyme turns over thousands of substrate molecules. Routine assays reach picograms per milliliter this way.

The obvious next step, measuring many proteins in one well by mixing many antibody pairs, runs into a combinatorial problem. With 10 pairs in a mixture, each detection antibody can meet any of the 10 capture antibodies. With 100 pairs there are 10,000 possible combinations, and every one that forms by accident produces false signal. Multiplexed immunoassays that read the pairs on beads or spots have therefore tended to stop at a few dozen proteins per well.

Turning a protein into DNA

The methods that broke through that ceiling did it by converting the protein measurement into a DNA measurement, borrowing the copying and pairing that proteins lack. In a proximity extension assay each protein is targeted by two antibodies, each carrying a short single strand of DNA. When both antibodies bind the same protein molecule, their DNA strands are held close enough to pair with each other at their tips. A polymerase then extends them into a new double-stranded DNA barcode unique to that protein. The barcodes are copied and counted, by quantitative PCR or by sequencing.

15 — Turning a protein into DNA
target protein antibody 1 (one site) antibody 2 (another site) DNA tag signal only when two binders meet on one molecule polymerase extended into one new DNA barcode amplify and count (qPCR or sequencing) surface no partner one antibody stuck non-specifically: no barcode, no signal
DNAProteinReadoutFocal detail
Plate 15 — Proximity extension converts a protein measurement into a DNA measurement. Two antibodies must land on the same molecule for their DNA tags to meet and be extended into a countable barcode — the second binder is what suppresses false signal.

The design solves the mixing problem at its root. Antibodies that stick where they should not do no harm unless their partner also lands on the same molecule, close enough for the DNA tips to meet. And because the readout is DNA, thousands of assays can share one tube and be read together on a sequencer. Current panels measure about 5,400 proteins from 2 microliters of plasma. A related approach that adds purification steps to suppress background reports detection limits in the attomolar range, several orders below conventional immunoassays, in its developers' own comparison.

Worked example: why a second binder helps so much

Suppose each antibody, on its own, produces a false signal from 1 in 100 of the off-target molecules it encounters. If a signal requires two antibodies binding the same molecule, and their mistakes are independent, the false-signal rate is 1/100 × 1/100 = 1 in 10,000.

The condition matters. The mistakes are independent only if the two antibodies recognize unrelated parts of the target. Two antibodies that share a weakness, for example both binding a feature common to a protein family, multiply much less than this.

Aptamers: binders made of nucleic acid

A second large-scale route uses aptamers instead of antibodies. An aptamer is a short single strand of chemically modified DNA that folds into a shape that binds a particular protein. It is found by selection: a vast random library of strands is exposed to the target, the binders are kept and copied, and the cycle is repeated until strong binders dominate. Because the binder is itself a nucleic acid, it can be read directly by hybridization or sequencing once the unbound strands are washed away. Current aptamer panels report about 11,000 measurements from 55 microliters of plasma.

An aptamer platform typically uses one binder per protein rather than a pair, which raises throughput and lowers sample volume but leans harder on the specificity of each binder. Each approach makes a different trade.

What each route sees

Mass spectrometry and affinity methods answer different questions, and the difference is structural rather than a matter of maturity.

Mass spectrometry identifies proteins by their sequence, needs no prior list, and sees modifications and variants, but it is shallow in plasma because abundant proteins consume its capacity. Affinity methods reach deep into the low-abundance range and run thousands of samples a week, but they measure only the proteins they were designed against. Their identity claim is only as good as the evidence that each reagent binds its intended target and nothing else. They report whatever the reagent binds, which can include complexes, fragments, or a related protein.

16 — Where each proteomics route can see
LC-MS, neat plasma — ~700–1,100 proteins per sample LC-MS after depletion or nanoparticle enrichment — ~1,400–2,400 Affinity panels (antibody pairs, aptamers) — thousands of preselected targets deep, but only for the targets it was built to find Single-molecule protein sequencing — emerging not yet at plasma scale 10−13 10−12 10−11 10−10 10−9 10−8 10−7 10−6 10−5 10−4 10−3 10−2 10−1 1 pg/mL 1 ng/mL 1 µg/mL 1 mg/mL plasma concentration, g/mL (log scale) bar ends are approximate albumin IL-6
ProteinReadoutFocal detail
Plate 16 — Mass spectrometry reads whatever is abundant without needing to know what to look for; affinity panels reach far deeper but only for the proteins they were designed against. The two routes answer different questions, which is why large studies increasingly run both.

A subtle consequence matters for genetics. A genetic variant that changes the part of a protein an antibody or aptamer recognizes can change how well the reagent binds without changing how much protein is present. In a large study, that shows up as a strong genetic effect on the measured level, and it can be mistaken for real regulation.

Two platforms, one protein name, different numbers

When two affinity platforms each report a protein called IL-6, they report what their own reagents bind, which may differ in which forms, fragments and complexes are captured. Published comparisons of the large platforms have found good agreement for some proteins and poor agreement for many others. Values from different platforms are not interchangeable, and a change in platform during a study is a change in the measurement.

Proteomics at population scale

Affinity methods are what made population-scale proteomics possible. The UK Biobank Pharma Proteomics Project measured 2,923 proteins in 54,219 participants, and its 2023 analysis mapped over 14,000 associations between genetic variants and protein levels. An expansion announced in January 2025 aims to measure up to 5,400 proteins in 600,000 samples, including repeat samples from the same people years apart, with full data expected in 2027. Studies like these link genetic risk to the proteins that carry it, and are a major source of new drug targets.

The third route: reading proteins one molecule at a time

Chapter 2 noted that single-molecule detection is the one route that could remove the advantage nucleic acids enjoy. Several groups are pursuing it for proteins. One commercial instrument, shipping since 2023 and in an upgraded version since 2025, immobilizes peptides on a chip and watches fluorescent recognizer proteins bind the amino acid at the exposed end while enzymes trim the peptide one residue at a time. In September 2026 its developer reported detecting 18 of the 20 amino acids with a developmental kit. Academic groups have shown that a protein strand can be pulled through a nanopore by a molecular motor and read with single-residue sensitivity, and that re-reading the same molecule improves accuracy substantially. A third company has opened early access to a platform that maps billions of intact protein molecules by repeated probing.

These are real demonstrations, and none yet measures a plasma proteome at a depth or cost that competes with the two established routes. The distance between those two statements is where Chapter 20 places them.

In practice

For an affinity measurement that will support a decision, treat each analyte as its own assay. Ask for evidence of specificity for that target, lot-to-lot consistency of the reagents, and the platform version. In a regulated product each claimed analyte needs its own analytical validation, whatever the multiplex level. In a study design, fix the platform and version for the whole study, and keep aliquots for orthogonal confirmation by a targeted mass spectrometry or immunoassay method.

+ What this chapter established
  • Sandwich immunoassays reach low concentrations through two-binder specificity and enzyme amplification.
  • Proximity extension converts proteins into DNA barcodes, so thousands of assays can share one tube.
  • Affinity methods reach deep but only for predefined targets; mass spectrometry is unbiased but shallow in plasma.
  • Single-molecule protein sequencing is demonstrated and advancing, but not yet at plasma scale.
10 — Metabolomics, lipidomics, glycomics

The chemistry in motion.

+ The questionWhy is the layer closest to the phenotype also the hardest to identify?

Small molecules, enormous variety

Metabolites are the small molecules of life, typically under about 1,500 daltons: sugars, amino acids, fatty acids and other lipids, nucleotides, vitamins, hormones, and the intermediates of every pathway that makes or breaks them. A blood or urine sample also carries molecules the body did not make: drug residues, food components, pollutants, and the products of gut microbes. The Human Metabolome Database, in its 2022 release, listed 217,920 annotated metabolites.

Two features set this layer apart. Metabolites are the direct substrates and products of enzymes, so they reflect what the cell's machinery is actually doing rather than what it could do. They are the closest molecular layer to the phenotype, and a change in diet, drug or disease shows up here within minutes. And they have no template. Nothing in the genome spells out a metabolite the way it spells out a protein's sequence, so the list of what should be present cannot be fully predicted from the genome. Every identification has to be earned from the chemistry of the molecule itself.

Three instruments, three trade-offs

Liquid chromatography coupled to mass spectrometry, LC-MS, covers the widest range of metabolites. Different columns retain polar or fatty molecules, and running the ion source in positive and negative mode captures molecules that prefer to gain or to lose a proton. Gas chromatography with mass spectrometry, GC-MS, handles molecules that are volatile or can be made volatile by chemical derivatization; its fragmentation is so reproducible that large reference libraries of spectra exist for matching.

Nuclear magnetic resonance, NMR, works on a different principle entirely. Placed in a strong magnetic field, certain atomic nuclei absorb radio waves at frequencies that depend on their chemical surroundings, so an NMR spectrum reports the atomic environments in a molecule. NMR is less sensitive than mass spectrometry, typically reaching micromolar rather than nanomolar concentrations, but it is quantitative without standards for each compound, it does not destroy the sample, and it is highly reproducible between laboratories. Those properties suit it to very large studies of the most abundant metabolites and lipoproteins.

One mass, several molecules

Chapter 3 noted that mass alone rarely identifies a molecule. Metabolomics is where that bites hardest. Glucose, galactose and fructose all have the formula C₆H₁₂O₆ and the same exact mass, 180.0634 daltons. A mass spectrometer that weighs them to four decimal places still cannot tell them apart, because they are isomers: the same atoms, arranged differently. The same problem runs through the whole metabolome. Many lipids differ only in where a double bond sits; many drug metabolites differ only in which position carries a hydroxyl group.

17 — One mass, several molecules
100 150 200 250 m/z [M−H]− 179.06 · C6H12O6 one peak, three identities glucose galactose fructose three candidates one OH flipped same formula, same exact mass (180.063 Da) retention time chromatography separates what mass alone cannot NMR distinguishes by atomic environment schematic
MetaboliteReadoutFocal detail
Plate 17 — A mass spectrometer can weigh a metabolite to a thousandth of a dalton and still not know what it is: isomers share a formula and a mass. Identification needs a second property — when the molecule elutes, how it fragments, or what an authentic standard does.

Identification therefore needs a second property. Chromatography separates many isomers, so the time at which a molecule emerges from the column helps. The way it fragments helps. The decisive evidence comes from running an authentic standard of the suspected compound under the same conditions and showing that it matches on at least two independent properties.

The identification bottleneck

An untargeted LC-MS run detects features: signals defined by an m/z and a retention time. A typical dataset contains tens of thousands of them. Most are not distinct compounds. The same molecule appears as several features because it picks up different ions such as sodium or potassium, because of natural carbon-13 isotopes, and because some molecules break apart in the ion source. A 2025 analysis that used isotope patterns to separate real compounds from these artifacts estimated that a typical dataset holds about 1,000 to 2,000 genuine compounds, and that more than half of them have no match in any database.

The Metabolomics Standards Initiative set out confidence levels for reporting identifications in 2007, and they remain the working vocabulary:

LevelMeaningEvidence
1Identifiedmatches an authentic standard, run in the same laboratory, on two or more independent properties such as retention time and fragment spectrum
2Putatively annotatedmatches a spectral library or literature, without a standard
3Putatively characterized compound classproperties consistent with a chemical class, such as a phosphatidylcholine
4Unknowndetected and quantified, not identified
18 — From feature to identified metabolite
Features detected (m/z × retention time): tens of thousands Real compounds after removing noise, adducts, isotopes: ~1,000–2,000 per dataset Annotated by database or library match: fewer than half Identified against an authentic standard (MSI level 1) the only level that is a confirmed identity MSI CONFIDENCE LEVELS Level 1 identified standard, two orthogonal properties Level 2 putatively annotated library / spectral match Level 3 compound class Level 4 unknown funnel tiers not to scale
MetaboliteFocal detail
Plate 18 — Untargeted metabolomics detects far more signals than it can name. Most features are not compounds at all, and of the real compounds more than half have no database match. A reported 'identification' should always carry its confidence level.
A name in a metabolomics table is a claim of a particular strength

Untargeted results often list compound names without saying how each was assigned. A level 2 annotation can be wrong about which isomer is present, and a level 3 annotation names only a class. A metabolite reported as a biomarker should carry its identification level, and one that will support a decision should be confirmed at level 1.

Targeted metabolomics: the version already in routine use

The opposite approach measures a predefined set of metabolites, each with an isotope-labeled internal standard, and reports absolute concentrations. It is less exploratory and far more reliable, and it is the form in which metabolomics entered routine medicine long before the word existed. Newborn screening measures amino acids and acylcarnitines in a dried blood spot by tandem mass spectrometry to detect inherited metabolic disorders. Developed in the early 1990s and adopted by public screening programs from the late 1990s, it now covers a large share of the conditions on screening panels such as the US recommended panel.

Handling is part of the measurement

The metabolome changes in seconds, and it keeps changing after the sample is taken. Blood cells in an unprocessed tube continue to consume glucose and release lactate; red cells that break release their contents; repeated freezing and thawing degrades labile molecules. Biological context adds further variation: fasting state, time of day, recent meals, exercise and medication all move metabolite levels. A metabolomics study that does not standardize collection, processing time, temperature and storage is measuring its own logistics as much as its subjects.

Lipids and glycans

Two sub-fields deserve separate mention because they matter in both biology and manufacturing. Lipidomics measures the thousands of lipid species in membranes, stores and signaling pathways. It uses mass spectrometry and faces an extreme version of the isomer problem. Glycomics measures glycans, the branched sugar chains attached to proteins and lipids. Glycans are not encoded by any template; they are built by enzymes, which makes them heterogeneous. Their branching produces isomers that share a mass and often co-elute. A standard workflow releases glycans from proteins enzymatically before chromatography and mass spectrometry. In therapeutic antibodies, the glycan pattern affects how the drug behaves in the body, which is why it is monitored during manufacturing as a critical quality attribute.

In practice

Write pre-analytical handling into the protocol as acceptance criteria: time from collection to processing, temperature, anticoagulant, freeze–thaw limits, and fasting status. When commissioning untargeted work, ask for identification levels per compound and for the fraction of features annotated at each level. For any metabolite that will support a clinical or release decision, move to a targeted, internally standardized method.

+ What this chapter established
  • Metabolites are the layer closest to phenotype, change in seconds, and have no template to predict them from.
  • Isomers share a mass, so identification needs retention time, fragmentation and ultimately an authentic standard.
  • Most untargeted features are not distinct compounds, and most real compounds lack a database match.
  • Targeted, internally standardized metabolomics has been routine in newborn screening since the late 1990s.

+ Part IV · Cells, space and shed material

From averages to individuals.

Everything so far measured a sample as a whole. These chapters measure its parts: individual cells, where they sit in a tissue, the other organisms living alongside them, and the vesicles and DNA fragments cells release into the blood. The recurring problem changes from identification to detection limits.

11 — Cytomics

Measuring cells one at a time.

+ The questionIf a tissue is a mixture, how do we measure each cell instead of the average?

The averaging problem

A blood sample holds billions of cells of dozens of types. Grind it up and measure a protein, and you get one number: the average across every cell, weighted by how common each type is. That average hides the most important distinction in cell biology. Suppose a marker rises by 10%. Either every cell now carries 10% more of it, or 10% more of the cells are of a type that carries it, or a rare population has expanded dramatically while everything else stayed still. The bulk number is the same in all three cases. A bulk measurement is a smoothie; to know what fruit went in, you need to look at the pieces.

Cytomics is the measurement of many properties of individual cells across very many cells. Its oldest and most widely used instrument is the flow cytometer, and the ideas it established, labeling cells with specific binders and reading them one at a time, carry straight through to the single-cell sequencing of the next chapter and to the vesicle measurements of Chapter 14.

Inside a flow cytometer

A flow cytometer has three subsystems. The fluidics inject the cell suspension as a thin core stream inside a faster-moving sheath of fluid. Because the sheath accelerates and narrows the core, a process called hydrodynamic focusing, the cells are forced into single file and pass the measurement point one at a time.

The optics shine one or more lasers across the stream. Each passing cell scatters light. Light scattered at small forward angles grows with the cell's size, and light scattered to the side grows with its internal complexity, such as granules and nuclear shape. Cells that carry fluorescent labels also emit light at longer wavelengths, which a series of mirrors and filters routes to separate detectors.

The electronics convert each flash into a pulse, a signal that rises as the cell enters the beam and falls as it leaves, and digitize its height, area and width. Each cell becomes one row in a table: forward scatter, side scatter and one value per fluorescence detector. Conventional instruments record thousands to tens of thousands of cells per second. Sorting instruments add a final step, breaking the stream into droplets and deflecting those containing wanted cells into separate tubes.

19 — Inside a flow cytometer
FLUIDICS · SIDE VIEW sample sheath hydrodynamic focusing cells in single file same point, from above OPTICS · FROM ABOVE laser FSC forward: size SSC · 90° internal complexity FL1 FL2 fluorescence: bound antibody-dye dichroic mirrors one cell at a time, ~microseconds SIGNAL · DATA height width area one event row FSC SSC FL1 FL2 FSC SSC
ReadoutFocal detail
Plate 19 — A flow cytometer turns a suspension into a table with one row per cell. Hydrodynamic focusing lines cells up single file, each crosses the laser in microseconds, and the scatter and fluorescence pulses become that cell's measurements.

Labels, panels and controls

The fluorescence comes from antibodies against cell-surface or internal proteins, each linked to a fluorescent dye. A set of such antibodies is a panel, and panel design is where most of the skill lies. Bright dyes are paired with proteins present at few copies per cell and dim dyes with abundant ones; markers found on the same cells are assigned dyes whose emissions overlap least. Every panel needs controls that show where background ends: unstained cells, cells stained with all antibodies but one, and a dye that marks dead cells, which bind antibodies nonspecifically.

More colors, more overlap

Each dye emits across a broad band of wavelengths, so a detector meant for one dye also catches some light from its neighbors. This spillover is corrected mathematically, a step called compensation, but the correction adds noise, and the problem grows quickly as colors are added. Conventional instruments with one filter per dye become impractical at a few dozen colors.

Two designs push further. Spectral cytometers record each cell's full emission spectrum across many narrow detectors and then unmix it, using each dye's reference signature to work out how much of each contributed. A 50-color panel was published in 2024 on spectral instruments, and a manufacturer placed a 60-color instrument in early access in 2026. Mass cytometry abandons light entirely. Antibodies are tagged with isotopes of heavy metals, each cell is vaporized and its atoms ionized, and a time-of-flight mass spectrometer counts the metal tags. Metal masses barely overlap and cells have no natural metal background, so about 50 parameters can be read cleanly. The costs are speed, hundreds rather than thousands of cells per second, the loss of every cell measured, and a substantial fraction of cells lost before measurement.

20 — Separating colors: three designs
CONVENTIONAL bandpass filter per dye wavelength → SPILLOVER one detector per dye; overlap spills into neighbors → compensation SPECTRAL array of narrow detectors wavelength → full spectrum per cell; dyes unmixed by their signatures ~40–50-color panels demonstrated MASS CYTOMETRY metal-tagged antibodies, read by time of flight 141 150 160 170 176 mass (Da) → metal tags, near-zero overlap ~50 parameters slower, and the cell is vaporized spectra and peak heights schematic
ReadoutFocal detail
Plate 20 — More colors means more overlap. Conventional cytometers fight spillover with one filter per dye; spectral instruments record the whole emission and unmix it; mass cytometry avoids light altogether by reading metal tags by mass, at the cost of speed and the cell.

From events to populations

The data are analyzed by gating: drawing boundaries on two-dimensional plots of one parameter against another to select populations step by step, for example living cells, then single cells, then T cells, then a subset of T cells. With dozens of parameters, manual gating gives way to clustering algorithms and two-dimensional maps of the high-dimensional data. Either way the result is a set of populations, each reported as a fraction of a parent population or as an absolute count per microliter.

Rare populations bring a statistical limit that no instrument removes. Counts of independent events fluctuate by about the square root of the count, so the precision of a population estimate depends on how many cells were counted in it.

Worked example: how many cells must be acquired?

A population makes up 0.01% of cells. To see 100 of them, acquire 100 ÷ 0.0001 = 1,000,000 cells.

Counting noise on 100 events is √100 = 10, a coefficient of variation of 10%. To halve it to 5%, count 400 events, which means acquiring 4,000,000 cells. The number of events in the gate, not the total acquired, sets the precision.

Fluorescence intensity is not a molecule count

Fluorescence is reported in arbitrary units that depend on laser power, detector gain and optical path, so the same cells read differently on two instruments or on one instrument after a service visit. Converting intensity into an equivalent number of reference fluorophores requires calibration beads of known brightness. Comparisons across instruments, sites or time are meaningful only after that calibration, a point that becomes decisive for vesicles in Chapter 14.

Why cytometry is an omics method

With 30 to 50 proteins measured on each of millions of cells, a cytometry run describes the composition and state of a cell population in more detail than most bulk proteomics can. Methods that bridge to sequencing go further. Antibodies carrying DNA barcodes instead of dyes can be read by the sequencer together with each cell's RNA, as Chapter 12 describes. The lineage is direct: label with a specific binder, read one object at a time, count.

In practice

Panel work should be written down as it is done: antibody clone, dye, titration result, and the controls that set each gate. Calibration beads should be run with every experiment so that intensities can be compared over time. Studies should be reported to the community's minimum-information standard for flow cytometry. For instrument design, the pulse chain from photons to digitized height, area and width is where coincidence, dead time and threshold settings silently decide which events exist.

+ What this chapter established
  • Bulk measurements average over cell types; cytometry measures each cell and counts populations.
  • A flow cytometer focuses cells into single file, reads scatter and fluorescence, and turns each cell into a row.
  • Spectral unmixing and metal tags push past the spillover that limits conventional color counts.
  • Precision for rare populations is set by events counted, and intensities need calibration to compare.
12 — Single-cell and spatial omics

Every cell, and where it sat.

+ The questionCan we read the whole transcriptome of each cell, and still know where in the tissue it came from?

Labeling instead of separating

Cytometry measures dozens of proteins per cell. Measuring the whole transcriptome of each cell is a different scale of problem. A single cell holds only about 10 picograms of RNA, and early methods handled one cell per tube, which limited studies to hundreds of cells. The methods that scaled to millions did so by separating cells only long enough to label them. If every RNA molecule from one cell carries the same DNA tag, all the cells can be pooled and sequenced together, and the reads sorted back into cells afterwards.

The most common implementation uses droplets. A microfluidic chip combines a stream of cells, a stream of gel beads and oil, producing thousands of droplets per second. Cells are loaded dilute, so most droplets hold no cell and an occupied droplet ideally holds one cell and one bead. Every bead carries millions of copies of a DNA strand with three working parts: a cell barcode, identical on every strand of that bead and different from every other bead; a unique molecular identifier, or UMI, a short random sequence that differs from strand to strand; and a string of T bases that pairs with the A-rich tail of messenger RNA. Inside the droplet the cell is broken open, its RNA is captured by the strands, and reverse transcription writes the barcode and UMI into each cDNA molecule. Then the droplets are broken and everything is pooled, amplified and sequenced.

21 — Barcoding every cell
IN EACH DROPLET cell bead ONE CAPTURE OLIGO, ENLARGED CELL BARCODE UMI POLY-T which cell this molecule came from AAAAAAAA captured mRNA cDNA copied by reverse transcription now carries barcode + UMI AFTER POOLING droplets broken, everything pooled and sequenced together BC3 BC1 BC5 BC2 BC1 BC4 BC3 sort reads by barcode CELLS (BARCODES) BC1 BC2 BC3 BC4 BC5 gene A 12 0 3 25 1 gene B 0 8 0 2 14 gene C 5 5 19 0 0 gene D 31 2 0 7 4 gene E 0 0 6 1 9 schematic counts UMI: unique tag per captured molecule — PCR copies share it and are counted once
DNARNAReadoutFocal detail
Plate 21 — Single-cell sequencing separates cells only long enough to label them. Each occupied droplet ideally holds one cell and one bead whose DNA barcode is stamped onto every RNA it captures, so after pooling, each read still carries the address of its cell.

Two labels, two jobs

The cell barcode answers "which cell did this molecule come from?" The UMI answers a subtler question: "is this read a new molecule, or a copy of one already counted?" Amplification makes many copies of each captured molecule, and some are copied more than others. Counting reads would reward the lucky ones. Counting distinct UMIs counts the original molecules.

Worked example: counting molecules, not copies

For one gene in one cell, 12 reads carry the same cell barcode. Their UMIs are: AGT (5 reads), CCA (4 reads), TTG (3 reads).

Reads = 12. Distinct UMIs = 3. The cell yielded 3 captured molecules of this gene's RNA; the other 9 reads were amplification copies. (Real UMIs are 10 to 12 bases long, so two molecules rarely share one by chance.)

Only a fraction of each cell's RNA molecules is ever captured and converted, so the data are sparse: most genes read zero in any given cell even when they are expressed at low levels. Single-cell data are powerful for telling cell types and states apart, and weak for measuring small differences in a single low-abundance gene.

Scale, and what it has produced

Throughput has grown by orders of magnitude. Droplet systems now process up to a few million cells per run, at a reagent cost the manufacturer states at around one cent per cell. Methods that barcode cells by repeatedly splitting and pooling them in plates, adding one barcode segment per round, reach a million cells per kit without droplets. The Human Cell Atlas consortium reported data on about 62 million cells in November 2024. An openly available single-cell data resource held about 125 million unique cells in its November 2025 release. A single 2025 dataset profiled 100 million cells from 50 cancer cell lines under about 1,100 drug-and-dose treatments.

From a matrix to cell types

The output is a matrix of cells by genes holding UMI counts. Analysis groups cells with similar expression profiles into clusters, draws them on two-dimensional maps, and labels the clusters by the marker genes they express. Cells caught at different stages of a continuous process, such as differentiation, can be ordered into trajectories.

Two artifacts recur. A droplet sometimes captures two cells, a doublet, which looks like a hybrid cell type that does not exist. And RNA released by broken cells into the suspension is captured in every droplet, adding a faint shared background. Both are handled computationally and both need to be checked.

A cluster is not a cell type

Clusters are statistical groupings whose number and boundaries depend on the analysis settings. A cluster may be a genuine cell type, a transient state, a set of doublets, or cells stressed by the dissociation that freed them from the tissue. Naming a cluster is an interpretation that needs independent evidence, such as known markers, protein measurements or spatial location.

Adding proteins and chromatin

The barcoding logic extends to other layers. Antibodies carrying DNA barcodes, instead of fluorescent dyes, bind the cell surface before the cell enters the droplet; their barcodes are captured and sequenced alongside the RNA. The method, CITE-seq, gives surface protein levels and transcriptome from the same cell, joining cytometry's view to sequencing's. Accessibility can be read per cell by running the adapter-inserting enzyme of Chapter 6 on nuclei before barcoding, and some kits read RNA and accessibility from the same nucleus.

Putting cells back in place

Dissociation throws away location, and location matters: a cell's neighbors decide much of what it does, and tissue structure is how a pathologist reads disease. Spatial methods keep the map, by two routes that are converging.

Capture arrays place a tissue section on a slide covered with barcoded spots whose positions are known. RNA from the tissue is captured by the spot beneath it, and the spot barcode records where it came from. The whole transcriptome is read, and resolution is set by spot size: some current arrays use features 2 micrometers wide or smaller, smaller than a cell, whereas earlier arrays used spots 55 micrometers across that pooled several cells.

Imaging methods leave the RNA in place and identify each molecule under a microscope. Panels of probes hybridize to their target RNAs over several rounds, and each gene's pattern of on and off signals across rounds forms a code that identifies it. Resolution is subcellular and each molecule is individually placed, but only genes in the designed panel are read. Panels have grown from hundreds to 5,000 or 6,000 genes, and one whole-transcriptome imaging panel covering more than 18,000 RNAs began shipping in 2025. Protein imaging works the same way with barcoded antibodies, reaching a hundred or more markers per section.

22 — Keeping the map: two approaches
SEQUENCING-BASED CAPTURE (1,1) (12,8) top view side view tissue section RNA barcoded spots on the slide whole transcriptome · resolution set by spot size (2 µm squares in current arrays) IMAGING-BASED IN SITU imaging rounds 1 2 3 gene B each molecule decoded to a gene at its exact position subcellular resolution · predefined panel: hundreds to ~18,000 genes SPATIAL RESOLUTION GENES MEASURED subcellular spot-sized a chosen panel the trade-off: how many genes vs how precisely placed imaging in situ capture arrays both are converging schematic
DNARNAReadoutFocal detail
Plate 22 — Spatial methods keep the address that single-cell dissociation throws away. Capture arrays read the whole transcriptome at the resolution of their spots; imaging reads a chosen panel molecule by molecule. The two are converging from opposite ends.

Every spatial method must also decide which molecules belong to which cell, a step called segmentation, usually from a nuclear or membrane stain. Segmentation errors assign molecules to the wrong neighbor, and in dense tissue they are a leading source of false cell-to-cell signals.

In practice

Most single-cell failures are decided before the instrument runs. Minimize the time and temperature of dissociation, and record both. Measure viability and set an acceptance threshold. Load cells at a density that keeps the doublet rate acceptable. For spatial work, agree the resolution, gene panel and tissue area with the biological question first; each costs the others. Keep a matched section for conventional histology so that clusters can be checked against morphology.

+ What this chapter established
  • Single-cell sequencing labels each cell's molecules with a barcode, then pools and sorts reads back into cells.
  • UMIs count original molecules rather than amplification copies; data are sparse because capture is incomplete.
  • Clusters are statistical groupings; calling them cell types needs independent evidence.
  • Spatial methods keep location, either by barcoded capture arrays or by imaging molecules in place.
13 — Metagenomics and the microbiome

The other genomes.

+ The questionWhen a sample holds a thousand species, how do we tell whose DNA is whose?

A second genome, carried everywhere

A human body carries roughly as many bacterial cells as human ones: one widely cited 2016 estimate put the bacteria at about 38 trillion against about 30 trillion human cells, most of them in the colon. Collectively those microbes carry far more genes than the human genome, and they do chemistry the body cannot do alone. They digest fibers, make vitamins and short-chain fatty acids, modify drugs, and train the immune system. The metabolites of Chapter 10 are partly theirs. So the question "what is this person's molecular state?" includes a question about other organisms' genomes, and this chapter is about how they are read from a mixed sample without growing each organism in culture, which most of them resist.

Marker-gene sequencing: who is there, cheaply

Every bacterium carries a gene for 16S ribosomal RNA, about 1,500 bases long. Parts of it are nearly identical across all bacteria and other parts, nine hypervariable regions, differ between lineages. That structure is ideal for a census. Primers that pair with the conserved stretches amplify a hypervariable region from every bacterium in the sample at once, and sequencing the products tells you which lineages were present and roughly in what proportions. Fungi are surveyed the same way with a different marker.

The method is cheap and sensitive, and it has three known limits. It usually resolves organisms only to genus, because closely related species share the same marker sequence. Amplification favors some sequences over others. And bacteria carry between one and about fifteen copies of the 16S gene depending on species, so a species with many copies looks more abundant than it is unless the analysis corrects for it.

Shotgun metagenomics: who is there, and what they can do

Shotgun metagenomics skips amplification of a marker and sequences all the DNA in the sample, exactly like the genome sequencing of Chapter 4 but without a single reference. Reads are assigned to species and often strains by matching them against genome databases, and genes are cataloged directly. The catalog shows what the community is equipped to do, including which antibiotic-resistance genes it carries.

The costs are money and interference. Deep shotgun sequencing costs several times more per sample than a marker census, though shallow shotgun sequencing narrows the gap. And in clinical samples most of the DNA is often human, so most reads are spent on the host unless human DNA is depleted first. Sequencing RNA instead of DNA, metatranscriptomics, adds which genes the community is actively using, the same step from genome to transcriptome made earlier in this guide.

23 — Two ways to read a community
SAMPLE microbes + host cells 16S AMPLICON one marker gene region, amplified many times who is there, usually to genus Genus A Genus B Genus C Genus D other 0% 100% relative abundance (schematic) SHOTGUN microbial DNA host DNA — often the majority in clinical samples all DNA, fragmented who is there (species/strain) + which genes they carry (function) relative abundances sum to 100%: when one taxon grows, the others appear to shrink even if they did not change
DNAFocal detail
Plate 23 — Amplifying one marker gene tells you who is present cheaply; sequencing everything tells you who and what they can do, at higher cost and with host DNA in the way. Either way the result is a share, not a count, and shares can mislead.

Shares, not counts

Both methods produce relative abundances: each organism's share of the reads. Because shares must sum to 100%, they behave in ways that mislead the unwary.

Worked example: how a constant population appears to fall

Before: species A and species B each make up 50% of the reads, and each is present at 1 billion cells per gram.

After: A doubles to 2 billion cells per gram; B is unchanged at 1 billion. Shares are now 2 ÷ 3 = 67% for A and 1 ÷ 3 = 33% for B.

B's share fell by a third although not one B cell was lost. From shares alone, "A increased" and "B decreased" are equally consistent with the data.

This is the same composition effect Chapter 7 met in RNA-seq, where one abundant gene absorbs reads from all the others. The remedies are also similar: add a known quantity of a reference organism or DNA to each sample before extraction, so that shares can be converted into absolute amounts; measure total microbial load independently, for example by quantitative PCR or by counting cells; or use statistical methods designed for proportions, which compare ratios between organisms rather than shares.

Low biomass and contamination

Laboratory reagents are not sterile at the level of DNA. Extraction kits and water carry traces of bacterial DNA, and in a stool sample, rich in microbes, those traces vanish. In a sample with little microbial DNA, such as blood, spinal fluid or a tissue biopsy, they can dominate the result, and a few published claims of microbes in supposedly sterile sites have been traced to contamination. The defense is procedural: process blank samples alongside real ones through every step, include a mock community of known composition, and treat any organism that also appears in the blanks with suspicion.

Metagenomics in the clinic

Unbiased sequencing is valuable where the cause of an infection is unknown, because it does not require a hypothesis about which pathogen to test for. A University of California San Francisco study published in 2024 reviewed seven years of metagenomic testing of cerebrospinal fluid, 4,828 samples. The test detected 63% of infections, and about one in five infections was identified by metagenomic sequencing alone, missed by all other tests run. A blood test that sequences microbial cell-free DNA to identify more than a thousand pathogens is offered as a laboratory-developed test and received FDA breakthrough-device designation in 2024. Nanopore sequencing has identified the pathogens in respiratory samples within about six hours in research settings.

Detecting DNA is not detecting an infection

A read from an organism shows that its DNA was in the tube. That DNA may come from dead organisms, from harmless colonization, from a neighboring body site picked up during sampling, or from reagents. Turning a metagenomic detection into a diagnosis needs thresholds calibrated against controls, knowledge of what is normally present at that site, and clinical judgment.

In practice

Every metagenomic run needs negative controls carried from extraction onward and a positive control of known composition. Report the fraction of reads that were host, and the depth remaining for microbes. If a decision depends on how much of an organism is present rather than whether it is present, add a spike-in or an independent load measurement, because relative abundance cannot answer that question.

+ What this chapter established
  • Microbial communities are read from mixed DNA without culture, by marker-gene census or by shotgun sequencing.
  • Marker genes say who is present, cheaply; shotgun sequencing adds species, strains and gene function.
  • Results are shares that sum to 100%, so a stable organism can appear to fall; spike-ins restore absolute amounts.
  • In low-biomass samples contamination can dominate, so blanks and mock communities are essential.
14 — EV and liquid-biopsy omics

What cells leave behind.

+ The questionWhat can the vesicles and DNA fragments that cells shed into the blood tell us about tissue we cannot biopsy?

A tube of blood as a sample of the whole body

Cells are constantly releasing material into the fluids around them. When a cell dies, its DNA is cut into fragments and some of them reach the bloodstream. Living cells release membrane-bound vesicles, extracellular vesicles or EVs, that carry proteins, RNA and lipids from the cell that made them. Both end up in blood, and they come from almost every tissue: bone marrow, liver, vessel walls, a growing fetus, and a tumor. That makes a blood draw a sample, however diluted, of tissues that could otherwise be reached only by biopsy. This is the idea behind liquid biopsy.

The same fact defines the difficulty. Material from the tissue of interest is mixed with far more material from everywhere else, and much of it sits at sizes and concentrations near the limits of detection.

24 — What a tube of blood holds, by size
proteins, antibodies 3–15 nm white blood cells 7–20 µm cfDNA on a nucleosome ~10 nm · fragment ~167 bp small EVs < 200 nm large EVs 200 nm – ~1 µm red blood cells ~7–8 µm lipoproteins HDL/LDL/VLDL 7–80 nm chylomicrons to ~1 µm platelets 2–3 µm 1 nm 10 nm 100 nm 1 µm 10 µm conventional flow cytometry (scatter) many instruments cannot resolve particles below ~300–600 nm nanoparticle tracking (scatter): from ~50 nm single-molecule fluorescence: counts labeled particles regardless of size EVs overlap in size with lipoproteins, which outnumber them in plasma most EVs are smaller than a conventional cytometer can see
DNAProteinMetaboliteReadoutFocal detail
Plate 24 — Extracellular vesicles occupy the same size range as lipoproteins and sit below what a conventional cytometer's scatter can resolve. Measuring them is a detection-limit problem before it is a biology problem.

Cell-free DNA: fragments with a fingerprint

Cell-free DNA in plasma is mostly short fragments of about 166 to 167 base pairs. The number is not arbitrary. In the nucleus, DNA is wrapped around nucleosomes, 147 base pairs per spool plus a short linker, and when a cell dies, enzymes cut the exposed linkers first, releasing the protected spool-length pieces. Fragments are cleared from the blood within hours, so what is in a tube reflects recent cell death.

The fragments carry three kinds of information. Their sequence carries any mutations of the cell that released them. Their methylation, the tissue-specific pattern of Chapter 6, reveals which tissue they came from: in healthy people most cell-free DNA comes from blood-forming cells, and about 1% from the liver. And their fragment length and end positions differ by origin: fetal DNA fragments are shorter, peaking around 143 base pairs, and tumor-derived fragments are shorter than healthy ones too.

The first large success was prenatal. Placental DNA enters the mother's blood early in pregnancy, and by about ten weeks it is typically around a tenth of the cell-free DNA, enough to test reliably. Sequencing millions of fragments and counting how many map to each chromosome reveals an extra copy of chromosome 21: the counts rise slightly but measurably above expectation. Launched in 2011, this non-invasive prenatal screening is now offered to all pregnant patients under US professional guidance.

Tumor DNA: a detection problem set by the tube

Tumor-derived DNA, circulating tumor DNA, is the basis of blood tests that choose cancer treatments, monitor response, and increasingly screen for cancer. In advanced disease it can be a large share of the cell-free DNA. In early disease it is tiny, often below 1% of fragments, and in about half of stage I cancers in one study it was undetectable. The reason is arithmetic, not instrument sensitivity.

Worked example: how many tumor fragments are in a tube?

Healthy plasma holds about 6.6 ng of cell-free DNA per mL. One copy of the haploid genome weighs about 3.3 pg, so 6.6 ng ÷ 3.3 pg ≈ 2,000 genome copies per mL.

A 10 mL blood tube yields about 4 to 5 mL of plasma: roughly 8,000 to 10,000 copies of any given position in the genome.

If the mutant allele makes up 0.1% of the copies at that position, roughly a 0.2% tumor fraction for a mutation on one of two copies, the whole tube holds about 8 to 10 mutant fragments. After extraction and library losses, a few may remain. No sequencer, however accurate, can find fragments that were never in the sample.

So the methods that work in early disease pool evidence across many positions rather than relying on one: many mutations tracked at once, methylation patterns across thousands of regions, and fragment-length profiles across the genome. Error correction with molecular identifiers, like the UMIs of Chapter 12, separates true rare variants from copying errors. A blood test for colorectal cancer screening that combines mutations, methylation and fragment patterns was approved by the FDA in July 2024. Tests aiming to detect many cancers from one draw are further behind, and Chapter 19 examines why the evidence bar for them is so high.

Extracellular vesicles: packages with a return address

EVs are particles enclosed by a lipid bilayer, released by essentially every cell type, ranging from about 30 nanometers to a micrometer or more. They carry surface proteins from the membrane of the cell that made them and cargo from its interior, including proteins, RNA fragments and metabolites. In principle that makes each vesicle a small, protected sample of a specific cell, which is why EVs attract so much interest as biomarkers and as drug-delivery vehicles.

The field's reference guideline, MISEV2023, recommends careful language because the biology is easy to overclaim. "Small EVs" means particles under 200 nanometers, as an operational size range; terms that imply a particular origin inside the cell should be used only when that origin has been shown. An EV preparation should also report its purity, because EVs share their size range with other particles in blood. Lipoprotein particles, which carry fats and cholesterol, overlap EVs in size and outnumber them by a wide margin in plasma; protein aggregates are also co-isolated.

25 — An EV, and what comes with it
CD9 CD81 other surface proteins miRNA, mRNA fragments proteins metabolites CD63 markers identify EVs — no single marker is universal ONE EV PREPARATION lipoprotein free protein protein aggregate non-vesicular particle MISEV2023: report purity and co-isolates; 'small EV' (<200 nm) is an operational size term
RNAProteinMetaboliteFocal detail
Plate 25 — An EV preparation is a mixture: vesicles carrying protein, RNA and lipid cargo, alongside lipoproteins and protein aggregates of similar size. Marker proteins such as the tetraspanins help identify the vesicles, but no single marker marks them all.

Every isolation method trades purity against yield. Ultracentrifugation pellets particles by density and size; size-exclusion chromatography separates them by size in a column; precipitation reagents recover a great deal of material and a great deal of contamination; capture with antibodies against surface markers gives high purity for one subpopulation and excludes the rest. The choice defines which EVs the study is actually about.

Measuring vesicles: the detection limit decides the answer

A typical EV is smaller than the wavelength of light, scatters very little of it, and carries only a limited number of any one surface protein. So EV measurement is first a detection-limit problem. Many conventional flow cytometers cannot distinguish a small EV's scatter from background: in a multicenter study of 46 flow cytometers, fewer than half detected the scatter of vesicles about 600 nanometers across, sizes estimated from bead calibration. Other methods reach smaller sizes by different routes. Nanoparticle tracking follows the scattered light of each particle as it drifts and sizes it from how fast it moves, down to about 50 nanometers, with fluorescence modes for labeled subsets. High-sensitivity flow instruments and single-molecule fluorescence cytometers count particles by their fluorescent labels rather than their scatter, so detection depends on labeling rather than size. Surface-capture chips and super-resolution microscopy image individual captured vesicles.

The consequence is that a concentration of EVs is not a property of the sample alone. It is a property of the sample, the instrument's detection limit and the analysis settings together.

26 — One sample, three instruments
10 6 10 7 10 8 10 9 particles / mL (log scale) 5.5 × 10 8 /mL SINGLE-MOLECULE FLOW CYTOMETER median diameter 57 nm detects single dye molecules; particles ≤ 35 nm 2.6 × 10 7 /mL CONVENTIONAL CYTOMETER A median diameter 83 nm higher detection limit 7.5 × 10 6 /mL CONVENTIONAL CYTOMETER B median diameter 105 nm higher detection limit ≈70-fold spread the count is set by the detection limit Source: Kim et al., J Extracell Vesicles 2024 — a split EV sample. When the counts are restricted to a common size and fluorescence range, most of the disagreement disappears.
ReadoutFocal detail
Plate 26 — The same vesicle sample, counted on three cytometers, returned concentrations spanning about seventy-fold. Each instrument counted correctly what it could see; what differed was where each one stopped seeing. A concentration without its detection limit is not a result.

A published split-sample comparison makes the point sharply. The same EV preparation, counted on a single-molecule flow cytometer and on two conventional flow cytometers, gave 5.5 × 10⁸, 2.6 × 10⁷ and 7.5 × 10⁶ particles per mL, a spread of about seventy-fold, with median sizes of 57, 83 and 105 nanometers. Each instrument counted correctly what it could see. When the authors restricted all three data sets to a common size and fluorescence range, most of the disagreement disappeared.

A concentration without a detection limit is not a result

"We measured 10⁹ EVs per mL" means little unless it states what the instrument could detect, in calibrated units of size and fluorescence, and which gates were applied. The community's reporting framework for EV flow cytometry, MIFlowCyt-EV, asks for exactly this: calibration, a stated detection limit, buffer-only and detergent-lysis controls, and serial dilution to show that single particles, not clumps or swarms, were counted.

Where EV diagnostics stand

EVs carry protected RNA and tissue-specific proteins, and their promise as biomarkers is real. Converting that promise into validated tests has been slow, largely because of the measurement problems above. A urine-based EV RNA test for prostate cancer received FDA breakthrough-device designation in 2019 and is cited in clinical guidelines. As of September 2026 we have not found an EV-based diagnostic with FDA clearance or approval. Most EV measurement remains a research activity, and the work that will move it forward is largely metrological: calibrated instruments, stated detection limits, reference materials and agreed controls.

In practice

For EV work, state with every concentration the instrument, its detection limit in calibrated units, the size and fluorescence gates, and the isolation method. Run buffer-only, detergent-lysis and dilution-series controls. Screen antibodies and fluorophores against the particle population on the instrument that will be used, because dye brightness and antibody labeling efficiency decide what a fluorescence-based counter can detect. When comparing with published values, compare within the same size and fluorescence window or not at all.

+ What this chapter established
  • Cell-free DNA and EVs carry information from tissues that cannot easily be biopsied.
  • cfDNA fragments reveal mutations, tissue of origin through methylation, and origin through fragment size.
  • In early cancer, the number of tumor fragments in the tube, not sequencer accuracy, limits detection.
  • EV counts depend on the detection limit; without calibration and stated limits they cannot be compared.

+ Part V · The case

Reading the twins.

With every layer now in hand, the case from Chapter 1 can be read the way a practitioner would read it: layer by layer, then across layers, then against its own limits. This is where the central question is answered in full.

15 — The layers together

What the layers showed together.

+ The questionWith the genome held constant, what did combining the layers reveal that no single layer could?

The design, and why it suits the question

The NASA Twins Study followed identical twins for 25 months: one spent 340 days on the International Space Station, the other stayed on Earth. Ten teams sampled blood, urine and stool before, during and after the flight, and measured the layers of this guide alongside physiology and cognition. Identical twins start from one fertilized egg, so the ground twin is as close to a genetic control as biology allows. With the inherited genome held constant, every difference the teams found had to live in another layer.

The design also has obvious limits, and reading it well means keeping them in view. It is one pair of people. Hundreds of thousands of measurements were compared, so some differences will be chance. Spaceflight bundles many exposures together, including microgravity, radiation, confinement, diet and disrupted sleep, and a single case cannot say which one caused a given change. The study is best read as a richly measured case that generates hypotheses, which is how its authors presented it.

The teams also did something that the chapters on cells anticipated: they separated blood into several immune cell fractions before measuring RNA and methylation. Without that step a change in the mixture of cells would have looked like a change inside cells, the averaging problem of Chapter 11.

Layer by layer

27 — The same genome, measured for 25 months
NASA Twins Study (Science, April 2019) · identical twins one spent 340 days on the ISS; the other stayed on Earth as the control LAYER IN FLIGHT 6 MONTHS AFTER LANDING Inherited genome sequence unchanged unchanged Epigenome (DNA methylation) shifted (immune, oxidative-stress genes) largely back Transcriptome (gene expression) many genes changed most returned; a minority stayed changed Telomeres lengthened mostly reverted; more short telomeres Proteome changed substantial recovery Metabolome changed substantial recovery Microbiome (gut) composition shifted near preflight Cognition some decline persisted the inherited sequence did not change
DNARNAProteinMetaboliteModificationFocal detail
Plate 27 — Ten measurement teams followed one pair of identical twins through a year in orbit. The inherited genome sequence was the constant; nearly every other layer moved, and most moved back. What persisted — in gene expression, telomere lengths, chromosomal rearrangements in blood cells and cognition — lies outside the inherited sequence.

Genome. The inherited sequence did not change, as expected. Two refinements keep this honest. Identical twins are not perfectly identical: each carries a small number of mutations acquired after the embryo split. And the blood cells of the flight twin showed more chromosomal rearrangements, such as inversions, during flight, some still present after return. Those are changes in some cells, acquired during life, not changes to the inherited genome. They are a reminder that "the genome is constant" is true of the germline and only approximately true of every cell.

Epigenome. DNA methylation shifted in flight in the immune cell types measured, including at genes involved in immune function and oxidative stress, and largely returned after landing.

Transcriptome. Expression changed for many genes during flight. Most returned to preflight levels within six months, and a minority remained changed. This is the finding that was misreported in 2018 as a change to the astronaut's DNA.

Telomeres. The protective ends of chromosomes lengthened in blood cells during flight, which surprised the investigators, then shortened quickly after return, leaving more short telomeres than before the mission.

Proteome and metabolome. Protein and metabolite patterns in blood and urine changed in flight, in directions consistent with fluid shifts and physiological stress, and recovered substantially after landing.

Microbiome. The composition of the gut community shifted in flight and returned close to its preflight state within six months.

Cognition. Measured performance declined after return and had not fully recovered at six months. It is not a molecular layer, but it is the phenotype that the molecular layers are ultimately trying to explain.

What the combination revealed

Three things were visible only because many layers were measured together.

First, agreement across layers is evidence. A shift in immune and stress-response genes seen in methylation, in expression and in proteins is far harder to dismiss as chance than the same shift seen in one data set of thousands of comparisons. A single-layer finding is a hypothesis; the same finding in independent layers is an argument.

Second, the layers moved on different clocks, and not simply the clocks Chapter 1 would suggest. Metabolites, gut microbes and most transcripts responded and recovered quickly, but so did most of the methylation changes. What persisted was a subset of expression programs, the distribution of telomere lengths, rearrangements carried by long-lived blood cells, and cognition. Timing across layers is information that no single layer holds.

Third, the layers disagreed, and the disagreement was informative. RNA and protein did not move in lockstep. That is not a defect of the study; it is how cells work.

Why RNA does not predict protein

Across genes, the amount of messenger RNA predicts the amount of protein only partly. Early genome-wide studies found that RNA levels explained around 40% of the variation in protein levels between genes; later analyses that corrected for measurement noise put it considerably higher, above 80% in some. Changes in RNA over time, within one gene, predict changes in its protein much less well, and those changes are what a study like the twins measures. The reasons are the control points of Chapter 1. Genes differ in how efficiently their RNA is translated. Proteins differ in how long they last, from minutes to months, so a long-lived protein can stay abundant long after its RNA has fallen. Secreted proteins leave the cell that made them. And a protein's activity can be switched by modification with no change in amount at all.

28 — RNA is not protein
protein level (log) mRNA level (log) fast protein turnover secreted — leaves the cell long-lived protein little mRNA, much protein schematic · across genes, mRNA explains roughly 40% of protein variation in early studies; later estimates correcting for measurement noise are higher
ReadoutFocal detail
Plate 28 — Knowing how much RNA a gene makes predicts its protein level only partly. Translation rates, protein lifetimes and secretion all intervene, which is why proteomics cannot be inferred from transcriptomics and must be measured.

The practical rule is short. RNA tells you what a cell is preparing to do; protein tells you what it has; activity and metabolites tell you what it is doing. None substitutes for the others.

How layers are combined

Integration methods fall into three families. Early integration joins all measurements into one table and analyzes it together, which is simple but lets the largest data set dominate. Late integration analyzes each layer separately and combines the conclusions, which respects each layer's statistics but misses relationships between them. Intermediate integration looks for a small number of hidden factors that explain variation across all layers at once, and has become the common choice for studies with many samples.

Genetics adds a powerful anchor. When a genetic variant is associated with the level of an RNA, a protein and a metabolite in the same direction, and with a disease, the chain from code to phenotype can be traced step by step across a population. This is why the large population studies of Chapters 9 and 18 measure several layers in the same people.

The central question, answered

The guide asked what each omics layer measures that the genome alone cannot. The twins supply the answer in miniature, and the chapters behind it supply the mechanism.

LayerWhat it records that the genome cannotHow it is readClock
Epigenomewhich parts of the genome a cell keeps available, and remembers through divisionsequencing after conversion, enzymes, antibodies, native long readsminutes to years
Transcriptomewhich genes are being read now, and how intenselycounting cDNA fragments by sequencinghours
Proteomewhich molecular machines exist, in which modified formsmass spectrometry, affinity bindinghours to days
Metabolomewhich chemistry is actually running, including diet, drugs and microbesmass spectrometry, NMRseconds to minutes
Cytome and single-cellwhich cells are present and in which states, rather than an averagecytometry, barcoded single-cell sequencingper cell
Spatialwhere each cell and molecule sits in the tissuecapture arrays, in situ imagingfixed at sampling
Microbiomethe genomes and activity of the organisms living alongside usmarker-gene and shotgun sequencingdays to weeks
Cell-free DNA and EVswhich tissues are releasing material now, and what it carriesdeep sequencing, methylation, single-particle countinghours

Each is a different record of state laid over the same code. They differ in cost and difficulty for a physical reason, whether the molecule can be copied and read by pairing. They differ in meaning because each sits at a different control point between instruction and function.

More layers do not mean more truth

Every added layer adds thousands of measurements and therefore thousands of chances for coincidence. Integration methods are powerful enough to find structure in pure noise. Multi-omics strengthens a conclusion when the layers were planned together, measured on the same samples, and when agreement across them was predicted before the data arrived, not discovered afterwards.

In practice

Plan a multi-omics study as one study, not several. Take all layers from the same collection at the same time, split into aliquots before any layer-specific processing, and record the metadata every layer will need: time of collection, time to processing, storage and batch. Decide in advance which cross-layer agreements would count as confirmation. And budget for the integration analysis; it is often the most expensive part.

+ What this chapter established
  • With the genome held constant, the twins' layers shifted in flight and mostly recovered; a subset persisted.
  • Agreement across independent layers turns a single-layer hypothesis into evidence.
  • RNA predicts protein only partly, because translation, lifetime, secretion and modification intervene.
  • Each layer records a different part of the cell's state over the same code: the answer to the central question.

+ Part VI · From signal to answer

How a measurement becomes a result.

The layers have been explained one principle at a time. These two chapters look at the whole instrument as a system and at the report it produces, because every number in an omics result has passed through a chain of conversions, and every conversion can fail.

16 — The instrument as a system

One architecture, many instruments.

+ The questionWhat actually happens between loading a sample and receiving a data file?

No single part is clever

Seen from the outside, a sequencer, a mass spectrometer and a flow cytometer look like three unrelated machines. Seen as systems they share one architecture. Each has a way of handling the sample and delivering it to the measurement, a physical conversion that turns molecules into something detectable, a sensor, detection electronics, a controller that keeps every step in time, and software that turns detector counts into an answer. Instrument engineers will recognize the signal chain behind every measuring instrument: sense → condition → digitize → compute → connect → decide. An omics instrument is that chain applied to a biological signal.

29 — Two instruments, one architecture
library + reagents fluidics (pumps, valves) flow cell (chemistry, temperature) optics (lasers, camera) image processing base calling reads (FASTQ) CONTROLLER sequences chemistry, temperature and imaging, cycle by cycle SEQUENCER sample liquid chromatography ion source (electrospray) mass analyzer(s) detector acquisition + spectra identification + quantity CONTROLLER schedules scans: which ions to fragment, when MASS SPECTROMETER both measure photons or ion counts — never the molecule
ReadoutFocal detail
Plate 29 — Seen as systems, a sequencer and a mass spectrometer are the same machine: sample handling, a physical conversion, a detector that counts photons or ions, a controller that keeps every step in time, and software that turns counts into an answer.

The subsystems of a short-read sequencer illustrate it. Fluidics pump reagents across the flow cell in a fixed sequence. The flow cell holds the chemistry at a controlled temperature. Lasers excite the dyes incorporated in each cycle and a camera images the surface. Image processing locates every cluster and measures its intensity in each color. Base calling converts those intensities into letters with a quality score. A mass spectrometer has the same shape: pumps and a chromatography column deliver peptides over time, electrospray converts them into ions, ion optics and mass analyzers sort them, a detector counts them, and acquisition software writes spectra that search software turns into identities.

In both, no individual component explains the result. The coordination does.

Timing is the system-engineering heart

In a sequencer, every cycle must add one base to every strand, image the whole surface, and remove the blocking group before the next cycle. If a small fraction of strands in a cluster miss an addition, or add two, they fall out of step with the rest, a problem called phasing. The cluster's signal then becomes a blend of the current position and its neighbors. Phasing accumulates cycle by cycle, which is why base quality declines along a read and why read length is limited. The controller's job is to hold flow, temperature and timing tightly enough that each cycle is as close to identical as the chemistry allows.

In a mass spectrometer, timing decides whether a peptide can be quantified at all. Each peptide emerges from the column over a few seconds. To measure its amount, the instrument must sample it several times across that peak, and every other ion it is also trying to measure competes for the same seconds.

Worked example: points across a peak

A peptide elutes over 6 seconds. The instrument's acquisition cycle, one MS1 scan plus its set of fragment scans, takes 0.6 seconds.

Points across the peak = 6 ÷ 0.6 = 10, enough to trace its shape and measure its area.

Add more fragment windows to reach more peptides and the cycle stretches to 1.5 seconds: 6 ÷ 1.5 = 4 points, too few for reliable quantification. Depth and precision trade against each other through time.

The data-conversion chain

The most useful way to look at any omics instrument is to follow one fact as it is re-represented from module to module. In sequencing: a DNA fragment becomes a cluster of copies, the cluster becomes a flash of light in each cycle, the flash becomes pixel intensities, the intensities become a base call with a quality score, the calls become a read, the read becomes an alignment, and alignments become a count or a variant call. In proteomics: a protein becomes peptides, peptides become ions, ions become a spectrum, the spectrum becomes a peptide match with a score, matches become an inferred protein, and signal becomes a quantity.

30 — The data-conversion chain
SEQUENCING DNA fragment amplification cluster of copies sequencing by synthesis flash of dye light per cycle imaging pixel intensities base calling base call + quality Q30 = 1 error in 1,000 join calls read alignment aligned read count / call count or variant call PROTEOMICS protein digestion (trypsin) peptides LC + electrospray ions mass analysis m/z spectrum database search peptide match (score, FDR) map peptides to proteins protein inference sum peptide signals quantity the step most often wrong: shared peptides make protein identity ambiguous each arrow can lose or distort information; the result is only as good as its weakest conversion
DNAProteinReadoutFocal detail
Plate 30 — No omics instrument measures its analyte directly. The quantity of interest is re-represented five or six times on its way to a report, and every conversion has a failure mode. Quality metrics exist to police those conversions, one by one.

At no point does the instrument measure the quantity of interest directly. It measures a proxy and converts it, five or six times over, and the result is only as trustworthy as the weakest conversion. Each conversion has a characteristic failure mode, and the quality metrics that accompany an omics result exist to police those conversions one by one.

ConversionWhat can go wrongThe metric that watches it
Library preparationuneven or biased sampling of moleculesduplicate rate, insert size, library yield
Signal to base callphasing, faint or crowded clustersfraction of bases at Q30 or above
Read to alignmentrepeats, missing reference sequencemapping rate, mapping quality
Alignment to calltoo few reads, systematic errorscoverage distribution, call quality
Ionization and detectionsuppression by co-eluting moleculessignal of spiked internal standards
Spectrum to peptidewrong match above thresholdfalse discovery rate from decoys
Peptide to proteinshared peptides, ambiguous identityprotein groups, unique peptides per protein

Controls travel the whole chain

A metric describes one conversion; a control tests all of them at once. A control is a known quantity passed through every step alongside the samples, so that anything lost or distorted along the way shows up as a difference between what went in and what came out. Sequencing runs carry a control genome whose sequence is known. RNA-seq experiments can include synthetic RNAs at known concentrations spanning the expected range. Targeted mass spectrometry adds isotope-labeled copies of each measured peptide. Cytometers run beads of certified brightness. In every case the logic is the same: the only way to know how well a chain of conversions worked is to push something known through it.

"Raw data" is already processed

Instrument output is often called raw data, but even the earliest file most users see is the product of models and settings. Base calls come from a model applied to images or currents, and spectra are centroided and filtered before they are written. Updating basecalling or acquisition software can change results from the same physical run. The software version is part of the measurement and belongs in the record.

In practice

For instrument and assay developers, the conversion chain maps directly onto design controls. Each conversion is a requirement with its own verification test, each quality metric is an acceptance criterion, and each control is a system-level verification carried into routine use. Signal-processing and calling software is part of the device and falls under the software-lifecycle process. For laboratories, the same list becomes the run record: instrument, software versions, metrics and controls for every batch.

+ What this chapter established
  • Sequencers, mass spectrometers and cytometers share one architecture: handle, convert, sense, detect, control, compute.
  • Timing is the system's critical property: phasing limits reads; points per peak limit quantification.
  • Every result passes through a chain of conversions, and each has a failure mode and a metric.
  • Controls are known quantities pushed through the whole chain; software versions are part of the measurement.
17 — Reading a real result

Telling a finding from noise.

+ The questionHanded an omics report, how do you tell a real finding from noise?

What a differential result contains

Most omics studies end in the same kind of table. Two groups of samples are compared, and for every feature measured, whether a gene, a protein or a metabolite, the table reports four things. It gives how large the difference is, as a fold change, and how surprising that difference would be if nothing were really different, as a p-value. It adjusts that p-value for the number of features tested. And it gives how abundant the feature was, since low-abundance features are measured with more noise. Reading the report means reading all four together and knowing what each can and cannot say.

Fold change, on the right scale

A fold change is a ratio. If a gene averages 200 normalized counts in treated samples and 50 in controls, the fold change is 200 ÷ 50 = 4. The reverse, 50 against 200, is 0.25. On a ratio scale those look unequal, although they describe the same size of change in opposite directions. Taking the logarithm to base 2 makes them symmetric: log₂(4) = 2 and log₂(0.25) = −2.

The rule of thumb practitioners carry is simple. A log₂ fold change of 1 is a doubling, −1 is a halving, and every further unit is another factor of two. A value of 3 is an eightfold increase. A value of 0.2 is a 15% increase, the kind of change that can be statistically certain and biologically unimportant.

Significance, and the price of testing everything

A p-value answers a narrow question: if there were truly no difference, how often would chance alone produce a difference at least this large? At the conventional threshold of 0.05, about one test in twenty will come out significant when nothing is going on. For one planned test that is an acceptable risk. For 20,000 genes it is a disaster: 20,000 × 0.05 = 1,000 genes expected to pass by chance alone, even if the treatment did nothing.

Omics analysis therefore controls the false discovery rate, the expected share of false findings among everything declared significant. The standard procedure, from Benjamini and Hochberg, ranks the p-values and compares each with a threshold that rises with its rank.

Worked example: controlling the false discovery rate

Ten features, target false discovery rate 5%. Sorted p-values and their thresholds, rank ÷ 10 × 0.05:

Rankp-valueThresholdBelow?
10.0010.005yes
20.0040.010yes
30.0120.015yes
40.0300.020no
50.0410.025no
6–100.20 to 0.900.030 to 0.050no

Find the largest rank still below its threshold: rank 3. Features 1 to 3 are declared discoveries. Features 4 and 5 have p-values under 0.05 and are not, because across ten tests they are too likely to be chance.

What a p-value is not

A p-value of 0.01 does not mean a 99% chance that the finding is real, and a non-significant result does not show that there is no difference. A p-value says how surprising the data would be if there were no effect. How likely an effect is to be real also depends on how many hypotheses were tested and how plausible they were beforehand, which is exactly what false-discovery control and replication address.

The volcano plot

The standard picture of a differential result plots every feature by its log₂ fold change across and its significance up, as the negative logarithm of the p-value so that stronger evidence sits higher. The shape gives the plot its name. Features in the upper corners changed a lot and reliably; the crowd at the bottom did not change detectably.

31 — Anatomy of a volcano plot
−4 −3 −2 −1 0 +1 +2 +3 +4 0 3 6 9 12 log2 fold change −log10 (p-value) schematic · 300 simulated genes 2-fold 2-fold FDR 5% threshold (after multiple- testing correction) HIGHER IN TREATED LOWER IN TREATED NOT SIGNIFICANT big change, few reads: not significant highly significant, biologically small
RNAFocal detail
Plate 31 — A volcano plot sets the size of a change against the confidence in it. The two axes answer different questions: a tiny, highly significant change and a large, unreliable one both deserve skepticism. Findings worth following sit in the upper corners.

Two regions deserve suspicion rather than excitement. High up near the middle are features with tiny changes and very small p-values; with enough samples, even trivial differences become statistically certain. Far out at the sides but low down are features with large changes and weak evidence, usually because they had few counts and therefore noisy estimates. The findings worth pursuing combine a change large enough to matter with evidence strong enough to trust.

Quality metrics decide whether any of it holds

Before interpreting the biology, check that the measurement worked. Each layer has its own gatekeeping metrics, the ones Chapter 16 tied to specific conversions.

LayerMetrics to checkWhat a problem looks like
DNA sequencingQ30 fraction, coverage distribution, duplicate and mapping ratesthin or uneven coverage over the regions that matter
RNA-seqRNA integrity, reads per sample, genes detected, ribosomal fractionone sample with far fewer genes detected than its group
Proteomicsidentifications at 1% FDR, missing values, CV of pooled QC injectionsdrifting QC signal across the run order
MetabolomicsCV of pooled QC samples, identification levelsmany features with QC CV above about 20–30%
Cytometryevents in each gate, viability, calibration beadsa rare population defined by a few dozen events
EVsdetection limit in calibrated units, buffer and lysis controlscounts that do not fall on dilution

When the batch is the biggest signal

Samples are rarely processed in one go. They are extracted on different days, run on different plates, sequenced on different flow cells, and each batch leaves its own small fingerprint on every measurement. With thousands of features, those small fingerprints add up to a signal that can be larger than the biology.

Principal component analysis shows it. It finds the directions along which samples differ most and plots them, and in a well-behaved experiment the leading components do not line up with processing batches; they may reflect the groups being compared or other biology such as sex or donor. When they separate the processing days instead, the batch dominates. If every treated sample was processed on one day and every control on another, the two are perfectly confounded and no analysis can separate them afterwards.

32 — When the batch is the biggest signal
colored by condition: no separation control (8) treated (8) PC2 (11%) PC1 (48% of variance) schematic · 16 simulated samples grouped by processing day: clean separation along PC1 PC2 (11%) PC1 (48% of variance) processed on day 1 processed on day 2 the largest source of variation is technical
ReadoutFocal detail
Plate 32 — Principal component analysis shows what varies most between samples. Here it is the day the samples were processed, not the treatment. Unless the design balances conditions across batches, a batch effect can be mistaken for biology or can hide it.

The remedy is design, not correction. Distribute each group evenly across batches, randomize the run order, and include pooled quality-control samples at intervals through every batch. Statistical batch correction works only when the design gives it something to work with.

Reading a real report

Here is an extract of a typical differential expression table, with a plain-language reading of each row.

GeneMean countlog₂ FCp-valueAdjusted pReading
Gene 15,4002.11 × 10⁻¹²4 × 10⁻⁹about fourfold higher, abundant, strong evidence: a finding
Gene 212,8000.183 × 10⁻⁸2 × 10⁻⁵reliably higher, but only by 13%: real and probably unimportant
Gene 393.40.0040.21tenfold change estimated from about nine reads: not significant after correction
Gene 4870−1.32 × 10⁻⁵0.003about 2.5-fold lower, good evidence: a finding
Gene 51,1500.050.710.93no detectable change
Spike-in control3,0000.020.880.97a known constant reads as constant: the pipeline is not inventing differences

The control row is as important as any gene. A spike-in added at the same amount to every sample should show no change; if it did, every other row would be in doubt.

In practice

When an omics report arrives, read it in this order: the quality metrics and controls; the design, meaning replicates, batches and whether they were balanced; the principal-component plot; the adjusted p-values and fold changes together, never either alone; and only then the biology. Ask for the full table, not just the significant list, and for the software and settings used. A finding that will drive a decision needs confirmation in independent samples.

+ What this chapter established
  • Read fold change on a log₂ scale: 1 is a doubling; small fold changes can be significant and unimportant.
  • Testing thousands of features requires false-discovery control; p < 0.05 alone gives about a thousand false hits per 20,000 tests.
  • Volcano plots show size and evidence together; trustworthy findings need both.
  • Quality metrics, controls and a balanced design come before any biological interpretation.

+ Part VII · Applications and the frontier

From signature to decision.

The last three chapters ask what omics already decides in practice, what it takes for a molecular signature to become a test someone can act on, and how to separate what has been shown from what has been announced.

18 — Omics in use today

Where omics already decides.

+ The questionWhere is omics already changing decisions, not just papers?

A working definition of "in use"

A technology is in use when a validated test built on it informs a decision routinely, is paid for, and appears in clinical guidelines or manufacturing specifications. By that test, omics is further along than its reputation suggests in a few areas and further behind in most. The routine uses are also narrower than the word implies. None of them measures a whole layer; each measures a validated subset of one, chosen because it answers a specific question.

33 — Where omics already decides
GENOME EPIGENOME (methylation) TRANSCRIPTOME PROTEOME METABOLOME METAGENOME Prenatal aneuploidy screening Cancer therapy selection Blood-based cancer screening Rare-disease diagnosis Newborn screening Infection diagnosis Drug–gene prescribing (pharmacogenomics) Biopharma process control Population cohorts routine since the late 1990s — omics before the word in routine clinical or industrial use in validation, pilot or research blank: not a main route
DNARNAProteinMetaboliteModificationFocal detail
Plate 33 — Omics is already inside routine decisions — prenatal screening, therapy selection, newborn screening — and the genome dominates. The hollow markers are where the next decade's evidence will be decided.

Before birth

The most widespread clinical use of genome sequencing is also one of the oldest: counting cell-free DNA fragments from maternal blood to screen for fetal chromosome abnormalities, launched in 2011. It rests on the counting property of short-read sequencing described in Chapter 5 and the fragment biology of Chapter 14. Professional guidance in the United States has recommended offering it to all pregnant patients since 2020.

Choosing cancer treatment

Cancer treatment is where omics has become most embedded, in three forms. Tumor genome profiling reads tens to hundreds of genes in a biopsy to find mutations that match a targeted drug. A broad profiling test covering 324 genes was approved by the FDA as a companion diagnostic in 2017. Two blood-based profiling tests followed in 2020, and in 2024 the FDA approved a distributable kit covering 517 genes with claims across tumor types. Expression signatures measure a panel of genes in a tumor to estimate the risk of recurrence. A 70-gene signature was cleared by the FDA in 2007 as the first multigene test of its kind. Two large trials changed practice: TAILORx, in 6,711 women with intermediate scores on a 21-gene assay, found that endocrine therapy alone was not inferior to adding chemotherapy, and MINDACT found that many women at high clinical risk but low genomic risk could forgo chemotherapy. Methylation of a single gene, MGMT, is routinely tested in brain tumors because it predicts response to a standard chemotherapy.

Two later steps point the same way. In November 2025 the FDA proposed moving nucleic-acid-based oncology companion diagnostics from its highest risk class to a lower one with special controls, a sign that the category has matured, and in May 2026 it approved a liquid-biopsy panel with a much larger footprint that adds methylation.

Diagnosing rare disease

For children with suspected genetic disease, sequencing has become the first test rather than the last. Meta-analyses put the diagnostic yield of genome sequencing at roughly 30% to 40%, compared with about 10% for the chromosomal microarrays it replaces. England's 100,000 Genomes Project pilot reported a 25% diagnostic yield, with about a quarter of diagnoses having immediate consequences for care. In intensive care, rapid genome sequencing typically returns results within a few days, and in record cases within hours, with a pooled diagnostic yield of about 37% and a change in management for about a quarter of the children tested.

Screening newborns

Newborn screening by tandem mass spectrometry, a targeted metabolomics test on a dried blood spot, has been routine since the late 1990s and screens every baby in many countries for dozens of inherited metabolic conditions. Genome sequencing at birth is now being tested at scale. England's Generation Study aims to sequence 100,000 newborns to screen for more than 200 treatable conditions, and had enrolled 25,000 babies by October 2025; a New York study reported results from its first 4,000 newborns in 2024.

Infection, drugs and manufacturing

Unbiased metagenomic sequencing is used for hard-to-diagnose infections, as Chapter 13 described, and pathogen genome sequencing has become a routine tool for tracking outbreaks and antimicrobial resistance. Pharmacogenomics uses a person's genotype to choose a drug or dose. The Clinical Pharmacogenetics Implementation Consortium maintains about thirty gene–drug guidelines, and the FDA's table of pharmacogenomic biomarkers in drug labeling runs to several hundred entries.

In biopharmaceutical manufacturing, omics methods are quality-control instruments. Mass spectrometry characterizes therapeutic proteins, monitors glycosylation as a critical quality attribute, and identifies residual host-cell proteins individually. Sequencing confirms the identity and stability of production cell lines.

Population cohorts: the research base for the next decade

Large cohorts that pair omics with health records are where most future clinical uses are being discovered. UK Biobank has whole-genome sequences for 490,640 participants. The US All of Us program's August 2026 release includes more than 535,000 short-read genomes, over 14,000 long-read genomes, and proteomics and RNA-seq on thousands of participants. GenomeIndia's 10,074 genomes give Indian research its own reference. The protein-level expansion of UK Biobank, 600,000 samples due in 2027, will link genetic risk to circulating proteins at a scale not seen before.

Routine omics is narrow omics

It is easy to read the list above as evidence that whole-layer measurement has arrived in the clinic. It has not. Each routine test measures a validated subset of one layer, chosen to answer one question, with its analysis locked and its performance established in the population it serves. Broad discovery measurements remain research tools; Chapter 19 describes what it takes to cross from one to the other.

In practice

Place any product or project on the map before planning it. Which layer does it measure, for which decision, in which population, and is that combination already routine, in validation, or still research? The answer sets the evidence needed, the competitors and precedents to study, and the likely regulatory route. A new test in a filled-in cell of the map has a precedent to follow; one in an empty cell will have to build its own case.

+ What this chapter established
  • Routine omics uses are narrow, validated subsets of one layer, each answering one question.
  • The genome dominates: prenatal screening, tumor profiling, rare-disease diagnosis, pharmacogenomics.
  • Targeted metabolomics has screened newborns since the late 1990s; newborn genome screening is being piloted at scale.
  • Population cohorts pairing omics with health records are the discovery base for the next decade.
19 — From signature to diagnostic

From signature to test.

+ The questionWhy do most omics discoveries never become tests, and what does it take for one that does?

Discovery and validation are different experiments

An omics discovery study measures thousands of features in tens or hundreds of samples and looks for a pattern that separates two groups. That situation, many more measurements than samples, is written p ≫ n by statisticians. It has a well-known property: with enough features, some combination will separate any two groups perfectly, even when every measurement is random. A model fitted to such data can learn the noise of that particular set of samples as easily as the biology.

34 — How a signature fits noise
20,000 measured features · 40 patients · pure noise case (20) control (20) TRAINING SET (the data used to choose the signature) AUC 1.00 shaded: called a case by the signature · schematic INDEPENDENT VALIDATION AUC 0.5 (coin toss) fails on new patients
ReadoutFocal detail
Plate 34 — With twenty thousand measurements and forty patients, some combination of features will separate any two groups perfectly — even when every number is random. Only data the signature has never seen can show whether it learned biology or noise.

The only reliable defense is data the signature has never seen. The signature, including every step of its processing and its decision threshold, is fixed, or locked, and then tested on independent samples, ideally collected elsewhere, processed at a different time, and drawn from the population in which the test would be used. Most omics signatures fail at this step. Their performance in discovery was real only for the samples used to find them.

History has made the point expensively. A serum proteomic pattern reported in 2002 to detect ovarian cancer with near-perfect accuracy was prepared for commercial launch within two years; the FDA said it required premarket approval, and independent reanalyses raised concerns that the separating patterns reflected differences in how and when samples were processed. It never reached the market as planned. In the late 2000s, gene-expression signatures from a US academic medical center were used to assign patients to chemotherapy in clinical trials. Independent statisticians could not reproduce the results, the trials were halted in 2010, and findings of research misconduct followed. The case led directly to a 2012 Institute of Medicine report on the evidence standards translational omics should meet.

The evidence ladder

A widely used framework for genetic and genomic tests asks three questions in order. Analytical validity: does the test measure the analyte accurately and reproducibly, in the specimen type it will receive, across operators, instruments, lots and days? Clinical validity: does the result correspond to the condition in the intended population, with known sensitivity and specificity? Clinical utility: does acting on the result improve outcomes for patients, compared with not testing? Ethical, legal and social considerations run alongside all three.

35 — From signature to test
CLINICAL UTILITY does using the result improve outcomes? CLINICAL VALIDITY does the result track the condition in the intended population? ANALYTICAL VALIDITY does the test measure the analyte accurately and reproducibly? DISCOVERY thousands of features, tens of samples INDEPENDENT VALIDATION COHORT where most omics signatures fail US LDT under CLIA (lab analytical validity) FDA 510(k) / De Novo / PMA companion diagnostic claims EU IVDR classes A–D · notified body companion diagnostics and genetic tests: class C INDIA CDSCO · MDR 2017 IVD classes A–D licensing: state A/B, central C/D + imports regulators review up to clinical validity; a CLIA-only LDT: analytical performance only utility: mostly for reimbursement and clinical guidelines
ReadoutFocal detail
Plate 35 — A signature becomes a test by climbing a ladder of evidence: it must be measured reliably, must track the condition in the people it is meant for, and must change what happens to them. Routes differ in how much of the ladder a regulator reviews.

Each rung needs a different kind of study, and they get harder going up. Analytical validation is laboratory work. Clinical validation needs well-characterized samples from the right people. Clinical utility often needs a prospective trial in which testing changes what is done and outcomes are followed. The expression signatures of Chapter 18 entered guidelines on analyses of archived trial samples; prospective trials in thousands of women later showed that most patients at low genomic risk could forgo chemotherapy with little or no loss of benefit.

The arithmetic of screening

Screening healthy people raises the bar further, because the disease is rare in the population being tested. Rarity changes what a positive result means.

Worked example: what a positive screening result is worth

Screen 100,000 people in whom 1% have undiagnosed cancer: 1,000 people with cancer, 99,000 without.

Sensitivity 50%: the test is positive in 500 of the 1,000.

Specificity 99.5%: it is falsely positive in 0.5% of 99,000 = 495.

Positive predictive value = 500 ÷ (500 + 495) ≈ 50%. Half of all positive results are false, even with a false-positive rate of only one in two hundred. And half the cancers are missed.

These numbers are close to those reported for real tests. A blood test designed to detect many cancers from one draw reported a positive predictive value of 52% and a specificity of 99.55% in a 142,000-person randomized trial in England. Its primary endpoint, a reduction in cancers diagnosed at stages III and IV, was not met, although stage IV diagnoses fell by 14%. The FDA's advisory committee voted on 23 September 2026 that its benefits outweigh its risks; as of this writing the agency's decision is pending. By contrast, the blood test for colorectal cancer approved in 2024 targets one cancer, with 83% sensitivity for colorectal cancer and about 10% false positives, and is approved for average-risk adults aged 45 and older; professional guidance positions it for people who decline stool tests and colonoscopy.

A high accuracy figure in a paper is not a useful test

Area under the curve, sensitivity and specificity measured in a discovery cohort are optimistic by construction. Even validated figures do not say whether a test is useful. That depends on how common the condition is in the tested population, what is done after a positive result, what that costs and risks, and whether it changes outcomes. A test is judged on its consequences, not its curve.

Regulatory routes: different roads, different reviews

United States. Tests can reach patients by two routes. A laboratory-developed test is designed and run within one certified laboratory under CLIA, which requires the laboratory to establish analytical performance; CLIA does not itself evaluate clinical validity, so no regulator reviews that rung for such a test. The FDA's 2024 rule asserting authority over laboratory-developed tests was vacated by a federal court on 31 March 2025, the agency did not appeal, and in September 2025 it restored its regulations to their earlier wording. A test sold as a kit goes through the FDA as a device: 510(k) clearance by comparison with an existing device, De Novo classification for a novel moderate-risk device, or premarket approval for the highest-risk class. A companion diagnostic, one essential for the safe and effective use of a particular drug, is developed and reviewed together with that drug. For sequencing-based tests for germline disease, FDA guidance since 2018 allows analytical validation against consensus standards and clinical validity to be supported by recognized public variant databases.

European Union. The In Vitro Diagnostic Regulation classifies tests from class A to D by risk. Companion diagnostics, human genetic tests and tests for cancer screening, diagnosis or staging are class C, which requires assessment by a notified body. Health institutions may make and use tests in-house under defined conditions, including a quality system and accreditation. Transition periods for devices already on the market run to the end of 2027 for class D, 2028 for class C and 2029 for class B, subject to conditions and deadlines. A Commission proposal of December 2025 to simplify parts of the regulation, including the in-house rules, is under negotiation, with adoption expected in 2027.

India. Under the Medical Devices Rules 2017, CDSCO classifies in vitro diagnostics from class A to D. Licensing became mandatory for classes A and B in October 2022 and for classes C and D in October 2023. Manufacturing licenses for A and B are issued by state authorities and for C and D centrally, and imports of every class are licensed centrally. Products labeled for research use only fall outside the rules, provided they are not intended, promoted or used for diagnosis. Prenatal testing, including cell-free DNA screening, is also governed by the PC&PNDT Act of 1994, which requires registration of facilities and prohibits disclosure of fetal sex. In 2026 CDSCO finalized guidance on medical device software, including software used with diagnostics, with provisions for AI-based changes. National ethical guidelines from ICMR require informed consent and genetic counseling around genetic testing.

The pipeline is part of the test

In an omics test, the analysis software does much of the measuring: it calls the variants, normalizes the counts, applies the model and produces the result. Regulators treat it accordingly. The FDA's sequencing guidance treats the bioinformatics pipeline as part of the test, to be documented, versioned and validated end to end and revalidated after changes. A signature defined by a model must be locked before validation. If the model is meant to change after release, US law since 2022 allows a predetermined change control plan, finalized in guidance in December 2024, that sets out in advance what may change and how each change will be verified. The international software-as-a-medical-device framework and India's 2026 software guidance take compatible positions.

In practice

Write the intended use before the first validation sample: the population, the specimen, the analyte, the result and the decision it informs. Everything else follows from it: the rungs of evidence needed, the size and source of the validation cohort, the risk class in each market, and whether a laboratory-developed test, a kit or a companion diagnostic is the right route. Lock the pipeline and the threshold before independent validation, version every component, and plan the regulatory file for each target market at the start. Regulation belongs among the design inputs, not at the final hurdle.

+ What this chapter established
  • Discovery with many more features than samples overfits easily; only locked, independent validation shows a real signature.
  • Tests climb a ladder of evidence: analytical validity, clinical validity, then clinical utility.
  • In screening, rarity makes even a very specific test produce many false positives.
  • US, EU and Indian routes differ in path and in how much of the ladder a regulator reviews; the pipeline is part of the test.
20 — The frontier

What changes next, and what does not.

+ The questionWhat will change in omics over the next decade, and what is still unsolved?

A rule for reading the future

Claims about the future of omics arrive faster than evidence, and they arrive dressed alike: a press release announcing a shipment reads much like one announcing a plan. The most useful habit this chapter can offer is a sorting rule. For any claim, ask which of four columns it belongs in: in routine use; available and in early use; demonstrated in research; or announced, projected or under review. Then read it with the evidence that column carries. The plate below applies the rule to the developments discussed in this guide, as of September 2026.

36 — Demonstrated, available, announced
ROUTINE USE short-read genome and exome sequencing cfDNA prenatal screening newborn screening by tandem MS clinical MS assays AVAILABLE, EARLY USE long-read clinical sequencing — first regulatory clearance of a clinical long-read system (China, 2025) spatial whole-transcriptome imaging blood-based colorectal cancer screening (FDA-approved 2024) 5-base / 6-base sequencing: DNA + methylation in one read single-cell at millions of cells DEMONSTRATED IN RESEARCH nanopore reading of protein strands single-EV multiplexed profiling multi-omics integration at cohort scale ANNOUNCED OR IN REVIEW multi-cancer blood test (PMA under FDA review, Sep 2026) single-molecule protein sequencing with all 20 amino acids population proteomics of 600,000 samples (full data expected 2027) AI ‘virtual cell’ models — do not yet beat simple baselines the honesty line: evidence to the left, promise to the right
DNARNAProteinMetaboliteModificationFocal detail
Plate 36 — Much of what is called the future of omics is already routine; much of what is announced is not yet shown. Reading a claim well starts with placing it in the right column.

Sequencing: cheaper, longer, and read with its marks

The cost of reading DNA will keep falling, but the easy gains have been taken, and the remaining costs sit increasingly in sample handling, analysis, storage and interpretation rather than reagents. The larger changes are in what a read contains. Long reads are entering the clinic; in November 2025 a long-read sequencing system received a Class III registration in China for thalassemia testing, described by its developers as the first regulatory clearance of a clinical long-read sequencer. Methods that report methylation and sequence from the same library, launched in 2024 and 2025, make the epigenome a routine by-product of the genome. Pangenome references are beginning to supplement single references, and population programs now include more than ten thousand long-read genomes alongside hundreds of thousands of short-read ones.

Proteins: population scale, and single molecules

Protein measurement is where the next decade will change most, because it starts furthest behind. Population proteomics will move from tens of thousands to hundreds of thousands of people, with the UK Biobank expansion's full data expected in 2027. Mass spectrometers keep gaining speed; manufacturers report hundreds of samples a day on current instruments. Single-molecule protein sequencing is the development to watch most carefully, because it attacks the limit set out in Chapter 2. It has moved from academic demonstration to shipping instruments with partial amino-acid coverage, and one developer reported detecting 18 of the 20 amino acids with a developmental kit in September 2026. Reading a plasma proteome molecule by molecule, at depth and at cost, remains a goal rather than a result.

Cells and space at atlas scale

Single-cell experiments now profile millions of cells per run and datasets of a hundred million cells exist. Much of that scale is being spent on perturbation atlases: measuring how cells respond when each gene is switched off or each drug is applied, systematically. Spatial methods are converging on whole-transcriptome coverage at cellular resolution, and on measuring RNA, protein and chromatin in the same section. The limiting factors are shifting from chemistry to segmentation, data volume and cost per tissue area.

Models trained on omics

Machine learning has already changed one part of molecular biology decisively: protein structure prediction, extended in 2024 to complexes of proteins with nucleic acids and small molecules. Models that read long stretches of DNA sequence and predict regulatory activity, or score the likely effect of a variant, were published in 2025 and 2026 and are in wide research use.

The more ambitious goal, a "virtual cell" that predicts how any cell will respond to any perturbation, is announced rather than demonstrated. Independent benchmarks published in 2025 found that several large single-cell foundation models did not outperform simple linear baselines at predicting the effects of genetic perturbations. In the first open challenge on the problem, completed in December 2025, the winning models did not beat naive baselines on every metric either. The binding constraint appears to be data designed to answer the question, not model size, which is why the perturbation atlases above matter.

Liquid biopsy and vesicles

Blood tests for many cancers at once are under regulatory review, and their evidence, including a large randomized trial that missed its primary endpoint while reducing stage IV diagnoses, will set the terms for the field. EV-based diagnostics are advancing more slowly, and the obstacle is largely metrological, as Chapter 14 argued: calibrated detection limits, reference materials and reporting standards. Progress there will look unglamorous and will matter more than any single biomarker.

What still needs to be solved

Most of the limits described in this guide are physical or statistical, not engineering problems waiting for the next platform. Naming them is the most durable part of any forecast.

  • The protein dynamic range. Ten orders of magnitude in plasma, no copying, no pairing. Enrichment, affinity and single-molecule methods each move the limit; none has removed it.
  • Metabolite identification. More than half of the real compounds in a typical untargeted dataset still have no confident identity.
  • Reproducibility across laboratories. Batch effects, platform differences and unreported processing choices still stop many results from transferring between sites.
  • Detection limits for particles. Vesicle and nanoparticle measurements need calibrated, stated detection limits before counts can be compared.
  • Diversity of reference data. Variant interpretation remains weakest for populations underrepresented in reference cohorts, including many in South Asia.
  • From validity to utility. Many signatures have analytical and clinical validity; far fewer have evidence that acting on them improves outcomes.
  • Interpretation and consent. The cost of analyzing, interpreting and counseling now exceeds the cost of generating the data, and genomic data raise privacy questions that outlast any single study.
The next platform will not solve the physics

It is tempting to treat every limit in omics as temporary, waiting for a better instrument. Some are. But copying and pairing, dynamic range, detection limits, the multiple-testing burden and the gap between association and utility are properties of molecules, samples and statistics. New platforms move these limits, sometimes a long way; they rarely abolish them. A claim that one has should be read with that in mind.

In practice

For any claim you meet, including your own, write down its column and its evidence: routine, with guidelines and payment; available, with early users; demonstrated, with a publication; or announced, with a date. Keep the four apart in proposals, marketing and technical files. Stating demonstrated and projected separately costs nothing and is what makes the rest of a document credible.

+ What this chapter established
  • Sort every claim into routine, available, demonstrated or announced, and weigh it by that column.
  • Sequencing gains are shifting from cost to content: long reads, native methylation, pangenomes.
  • Protein measurement will change most, with population scale now and single-molecule methods still emerging.
  • The hardest limits are physical and statistical, and platforms move them more often than they remove them.
§ Lessons

Lessons.

The chapters reduce to a short set of working rules for anyone who designs, runs, buys or reviews omics measurements. Each one traces back to a mechanism explained earlier.

  1. Name the layer before the technology. Decide whether the question concerns the code, its accessibility, its use, its products or its chemistry, and let that choose the method (Chapters 1 and 15).
  2. Ask where the signal per molecule comes from and how identity is assigned. Amplification, a bright label or single-molecule detection; sequence, mass or a tested binder. The answers predict depth and failure modes (Chapters 2 and 3).
  3. Specify minimums where they matter, not averages. Coverage over the target regions, reads per CpG, events in the gate (Chapters 4, 6 and 11).
  4. State the detection limit with every count or concentration, in calibrated units, with the gates and controls that produced it (Chapters 11 and 14).
  5. Treat every name in a results table as a claim with a confidence level. Protein groups are inferred; metabolite annotations have levels; clusters are interpretations (Chapters 8, 10 and 12).
  6. Balance groups across batches and include pooled quality controls. Confounding cannot be removed after the fact (Chapter 17).
  7. Read fold change and adjusted significance together, never either alone (Chapter 17).
  8. Push something known through the whole chain. Spike-ins, standards and calibration beads test every conversion at once (Chapter 16).
  9. Version the software as part of the instrument, and record it with every result (Chapters 16 and 19).
  10. Keep aliquots for orthogonal confirmation by a targeted method before any finding drives a decision (Chapters 7, 9 and 10).
  11. Validate a locked signature on samples it has never seen before anyone acts on it (Chapter 19).
  12. Write the intended use first: population, specimen, analyte, result and decision. The evidence plan and the regulatory route follow from it (Chapter 19).
  13. Keep demonstrated and announced apart in every claim, including your own (Chapter 20).
§ Glossary

Glossary.

Terms are defined as they are used in this guide.

Adapter
A short synthetic DNA sequence attached to the ends of each fragment so that the sequencer can grip and read it.
Affinity reagent
A molecule, usually an antibody or aptamer, that binds a target by shape. Its specificity must be shown for each target.
Alignment
Placing each sequencing read at the position in a reference genome that it best matches.
Analytical validity
Whether a test measures its analyte accurately and reproducibly in the specimens it will receive.
Aptamer
A short single strand of DNA or RNA, often chemically modified, selected to fold into a shape that binds a particular target.
ATAC-seq
A method that maps open chromatin by letting an adapter-inserting enzyme act only where DNA is accessible.
Base calling
The conversion of a sequencer's raw signal, light or current, into letters with quality scores.
Batch effect
Systematic differences between groups of samples processed at different times or places, unrelated to biology.
Bisulfite conversion
A chemical treatment that turns unmethylated cytosine into uracil, read as T, while methylated cytosine stays C.
cDNA
Complementary DNA: a DNA copy of RNA made by reverse transcriptase, so that RNA can be amplified and sequenced.
Cell barcode
A DNA sequence shared by every molecule captured from one cell, used to sort pooled reads back into cells.
Cell-free DNA (cfDNA)
Short DNA fragments released into blood by dying cells, mostly about 167 base pairs long.
Circulating tumor DNA (ctDNA)
The fraction of cell-free DNA released by a tumor.
Clinical utility
Whether acting on a test result improves outcomes compared with not testing.
Clinical validity
Whether a test result corresponds to the condition of interest in the intended population.
Companion diagnostic
A test essential for the safe and effective use of a specific drug, developed and reviewed together with it.
Compensation
The mathematical correction of fluorescence spillover between detectors in flow cytometry.
Coverage
The number of sequencing reads overlapping a position, or its average across a region or genome, written as 30× and so on.
CpG
A cytosine followed by a guanine in the DNA sequence: the usual site of DNA methylation in human cells.
Data-independent acquisition (DIA)
A mass spectrometry mode that fragments all ions within successive mass windows rather than selecting the most intense.
Detection limit
The smallest amount or signal an instrument or assay can distinguish reliably from background.
Dynamic range
The ratio between the largest and smallest quantities present in a sample, or measurable by an instrument in one run.
Electrospray ionization
Spraying a liquid from a charged needle so that dissolved molecules become gas-phase ions for mass spectrometry.
Epigenome
The chemical marks and packaging on DNA that control which genes a cell can use, copied through cell division.
Extracellular vesicle (EV)
A particle enclosed by a lipid bilayer, released by cells, carrying protein, RNA and lipid cargo.
False discovery rate (FDR)
The expected share of false findings among all results declared significant.
Feature
In metabolomics, a signal defined by a mass-to-charge ratio and a retention time, which may or may not be a distinct compound.
Fold change
The ratio of a quantity between two conditions, usually reported on a log₂ scale.
Gating
Selecting a population of cells or particles by drawing boundaries on plots of their measured parameters.
Hydrodynamic focusing
Narrowing a sample stream inside a faster sheath flow so that cells pass a detector in single file.
Isomers
Molecules with the same atoms, and therefore the same mass, arranged differently.
Laboratory-developed test (LDT)
A test designed, validated and run within a single laboratory rather than sold as a kit.
Library
The collection of adapter-tagged fragments prepared from a sample for sequencing.
Long read
A sequencing read from a single native molecule, typically thousands to hundreds of thousands of bases long.
Mass cytometry
Cytometry that labels antibodies with metal isotopes and reads them by mass instead of fluorescence.
Mass-to-charge ratio (m/z)
The quantity a mass spectrometer actually measures: an ion's mass divided by its charge.
Metagenomics
Sequencing the mixed DNA of a microbial community without culturing its members.
Methylation level
The fraction of reads at a site that carry a methyl mark: a population average, not a property of one cell.
MISEV2023
The International Society for Extracellular Vesicles' guidelines on nomenclature, isolation, characterization and reporting.
Multiple testing
The inflation of chance findings when many hypotheses are tested at once.
Nanoparticle tracking analysis (NTA)
A method that sizes and counts particles by following the scattered or fluorescent light of each as it drifts.
Nucleosome
A spool of histone proteins around which about 147 base pairs of DNA are wrapped.
Omics
The measurement of a complete set of molecules of one kind, a whole layer, in a single experiment.
Phasing (sequencing)
Loss of synchrony among strands in a cluster, which degrades base quality along a read.
Positive predictive value
The share of positive test results that are true positives. It depends strongly on how common the condition is.
Principal component analysis (PCA)
A method that finds the directions of greatest variation among samples, used to spot batches and outliers.
Protein inference
Deducing which proteins were present from the peptides identified; ambiguous when peptides are shared.
Proteoform
A specific molecular form of a protein, including its sequence variant, splicing and modifications.
Proximity extension assay
An affinity method in which two DNA-tagged antibodies must bind the same protein for their tags to form a countable barcode.
Quality score (Q)
A logarithmic estimate of a base call's error probability: Q30 is one error in 1,000.
Reference genome
An agreed genome sequence used as the coordinate system for aligning reads.
Retention time
The time a molecule takes to emerge from a chromatography column, used as an identity coordinate.
Reverse transcriptase
An enzyme that copies RNA into DNA.
Spectral cytometry
Cytometry that records each cell's full emission spectrum and unmixes the dyes computationally.
Spike-in
A known quantity of a reference molecule added to every sample to test and calibrate the whole measurement chain.
Structural variant
A genomic change of more than about 50 base pairs: a deletion, duplication, inversion or insertion.
Transcriptome
The complete set of RNA molecules in a cell or sample at a given time.
Unique molecular identifier (UMI)
A random tag attached to each captured molecule so that amplification copies are counted once.
Variant calling
Identifying positions where the reads consistently differ from the reference, with a confidence for each call.
§ Sources

Sources.

Sources are listed by the chapter in which they are first used. Figures from manufacturers are identified as such in the text. Status and prices are as of September 2026.

  1. Ch. 1, 15 — Garrett-Bakelman F.E. et al. The NASA Twins Study: a multidimensional analysis of a year-long human spaceflight. Science 364, eaau8650 (2019). doi.org/10.1126/science.aau8650
  2. Ch. 1 — National Geographic. No, Scott Kelly's year in space didn't mutate his DNA (15 March 2018).
  3. Ch. 1, 15 — Schwanhäusser B. et al. Global quantification of mammalian gene expression control. Nature 473, 337–342 (2011).
  4. Ch. 2, 9 — Anderson N.L., Anderson N.G. The human plasma proteome: history, character, and diagnostic prospects. Mol Cell Proteomics 1, 845–867 (2002).
  5. Ch. 4 — Nurk S. et al. The complete sequence of a human genome. Science 376, 44–53 (2022).
  6. Ch. 4 — Liao W.-W. et al. A draft human pangenome reference. Nature 617, 312–324 (2023); Human Pangenome Reference Consortium, Data Release 2 (May 2025). humanpangenome.org
  7. Ch. 4, 18 — GenomeIndia Consortium. Nature Genetics (April 2025). doi.org/10.1038/s41588-025-02153-x
  8. Ch. 5 — National Human Genome Research Institute. DNA Sequencing Costs: Data (last updated May 2023). genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data
  9. Ch. 5 — Manufacturer specifications and announcements for short-read, long-read and sequencing-by-expansion systems, 2022–2026 (Illumina, Ultima Genomics, PacBio, Oxford Nanopore Technologies, Roche). Figures quoted are the manufacturers' own.
  10. Ch. 6 — Buenrostro J.D. et al. ATAC-seq. Nature Methods 10, 1213–1218 (2013); Kaya-Okur H.S. et al. CUT&Tag. Nature Communications 10, 1930 (2019); Lieberman-Aiden E. et al. Hi-C. Science 326, 289–293 (2009).
  11. Ch. 7 — ENCODE Consortium. RNA-seq data standards. encodeproject.org/data-standards/rna-seq/long-rnas/
  12. Ch. 8 — Guzman U.H. et al. Ultra-fast label-free quantification and comprehensive proteome coverage with narrow-window data-independent acquisition. Nature Biotechnology (2024). doi.org/10.1038/s41587-023-02099-7
  13. Ch. 8 — Deutsch E.W. et al. Human Proteome Project 2025 report. Journal of Proteome Research 25(2), 539 (2026).
  14. Ch. 8 — Al Siblani H., Armengaud J., Lozano C. Independent comparison of plasma preparation and enrichment methods for DIA proteomics. Journal of Proteome Research (June 2026).
  15. Ch. 9 — Sun B.B. et al. Plasma proteomic associations with genetics and health in the UK Biobank. Nature 622, 329–338 (2023); UK Biobank, launch of the expanded Pharma Proteomics Project (January 2025).
  16. Ch. 9 — Feng W. et al. Multiplexed immunoassay with attomolar sensitivity (NULISA). Nature Communications 14, 7238 (2023).
  17. Ch. 9 — Motone K. et al. Multi-pass, single-molecule nanopore reading of long protein strands. Nature 633, 662–669 (2024); Reed B.D. et al. Real-time dynamic single-molecule protein sequencing. Science 378, 186–192 (2022).
  18. Ch. 10 — Wishart D.S. et al. HMDB 5.0. Nucleic Acids Research 50, D622–D631 (2022); Sumner L.W. et al. Proposed minimum reporting standards for chemical analysis. Metabolomics 3, 211–221 (2007).
  19. Ch. 10 — Chi et al. (S. Li laboratory). Study separating real compounds from adducts, isotopes and in-source fragments in untargeted LC-MS data. bioRxiv doi.org/10.1101/2025.02.04.636472 (2025); Metabolomics (2026).
  20. Ch. 11 — Konecny A.J. et al. OMIP-102: 50-color phenotyping of the human immune system. Cytometry A (2024); Pillai et al., mass cytometry review, Frontiers in Immunology (2022).
  21. Ch. 11 — Cytek Biosciences. Announcement of a 60-color spectral instrument in early access (June 2026). Manufacturer claim.
  22. Ch. 12 — Human Cell Atlas, Nature collection (November 2024); CZ CELLxGENE Census, long-term-support release 2025-11-08; Stoeckius M. et al. CITE-seq. Nature Methods 14, 865–868 (2017).
  23. Ch. 12 — Arc Institute and Vevo Therapeutics. Tahoe-100M single-cell perturbation atlas (February 2025).
  24. Ch. 13 — Sender R., Fuchs S., Milo R. Revised estimates for the number of human and bacteria cells in the body. PLoS Biology 14, e1002533 (2016).
  25. Ch. 13 — Benoit P. et al. Seven-year performance of a clinical metagenomic next-generation sequencing test for diagnosis of central nervous system infections. Nature Medicine (2024). doi.org/10.1038/s41591-024-03275-1
  26. Ch. 13 — Karius. FDA Breakthrough Device designation for a plasma cell-free DNA pathogen test (May 2024).
  27. Ch. 14 — Welsh J.A. et al. Minimal information for studies of extracellular vesicles (MISEV2023). Journal of Extracellular Vesicles 13, e12404 (2024).
  28. Ch. 14 — Welsh J.A. et al. MIFlowCyt-EV: a framework for standardized reporting of extracellular vesicle flow cytometry experiments. Journal of Extracellular Vesicles 9, 1713526 (2020).
  29. Ch. 14 — Kim et al. Split-sample EV concentration measured on a custom single-molecule flow cytometer and two conventional flow cytometers. Journal of Extracellular Vesicles 13(8), e12498 (2024).
  30. Ch. 14 — Moss J. et al. Comprehensive human cell-type methylation atlas reveals origins of circulating cell-free DNA. Nature Communications 9, 5068 (2018).
  31. Ch. 14 — van der Pol E. et al. Standardization of extracellular vesicle measurements by flow cytometry through vesicle diameter approximation (ISTH multicenter study of 46 flow cytometers). Journal of Thrombosis and Haemostasis 16, 1236–1245 (2018).
  32. Ch. 14, 19 — US FDA. Shield (P230009), approved July 2024. fda.gov/medical-devices/recently-approved-devices/shield-p230009
  33. Ch. 14 — Bio-Techne. FDA Breakthrough Device Designation for an exosome-based prostate cancer test (June 2019).
  34. Ch. 15 — Vogel C., Marcotte E.M. Insights into the regulation of protein abundance from proteomic and transcriptomic analyses. Nature Reviews Genetics 13, 227–232 (2012).
  35. Ch. 18 — Sparano J.A. et al. TAILORx. New England Journal of Medicine 379, 111–121 (2018); Cardoso F. et al. MINDACT. NEJM 375, 717–729 (2016).
  36. Ch. 18 — 100,000 Genomes Project Pilot Investigators. New England Journal of Medicine 385, 1868–1880 (2021); meta-analyses of diagnostic yield in npj Genomic Medicine (2018, 2024) and Genetics in Medicine (2025).
  37. Ch. 18 — Genomics England. 25,000 babies join the Generation Study (October 2025); UK Biobank whole-genome sequencing, Nature (2025); All of Us Research Program, Curated Data Repository v9 (August 2026).
  38. Ch. 18 — CPIC guidelines (cpicpgx.org); US FDA, Table of Pharmacogenomic Biomarkers in Drug Labeling (updated August 2026).
  39. Ch. 18 — US FDA and CMS. FoundationOne CDx approval (November 2017); FDA approval of TruSight Oncology Comprehensive, P230011 (August 2024); FDA clearance of the 70-gene MammaPrint signature (February 2007).
  40. Ch. 18 — US Federal Register. Proposed reclassification of nucleic acid-based test systems for use with oncology therapeutics (25 November 2025); Guardant Health, FDA approval of an expanded liquid-biopsy companion diagnostic with methylation (May 2026).
  41. Ch. 19 — CDC. ACCE model for evaluating genetic tests. Institute of Medicine. Evolution of Translational Omics: Lessons Learned and the Path Forward (2012).
  42. Ch. 19 — US Federal Register. Medical devices; laboratory developed tests (19 September 2025); American Clinical Laboratory Association v. FDA, E.D. Tex. (31 March 2025).
  43. Ch. 19 — US FDA. Considerations for design, development, and analytical validation of NGS-based IVDs; Use of public human genetic variant databases (both April 2018). Marketing submission recommendations for a predetermined change control plan for AI-enabled device software functions (December 2024).
  44. Ch. 19 — Regulation (EU) 2017/746 (IVDR); Regulation (EU) 2024/1860; European Commission proposal COM(2025) 1023 (December 2025).
  45. Ch. 19 — CDSCO. Medical Devices Rules 2017, IVD FAQs; Guidance document on medical device software (2026). ICMR National Ethical Guidelines for Biomedical and Health Research (2017), Section 10.
  46. Ch. 19 — GRAIL. NHS-Galleri trial results (ASCO, May 2026); FDA advisory committee vote (23 September 2026).
  47. Ch. 19 — India. Pre-Conception and Pre-Natal Diagnostic Techniques (Prohibition of Sex Selection) Act, 1994.
  48. Ch. 20 — Abramson J. et al. AlphaFold 3. Nature 630, 493–500 (2024); Arc Institute, Virtual Cell Challenge 2025 wrap-up (December 2025); benchmark of single-cell foundation models for perturbation prediction, Nature Methods (August 2025).
  49. Ch. 20 — PacBio. First regulatory clearance of a clinical long-read sequencer, NMPA Class III (4 November 2025); Quantum-Si, detection of 18 amino acids on a developmental kit (15 September 2026).
§ PDF edition

Take the PDF with you.

Reading online needs no sign-up. For the PDF edition, laid out for print and offline reading, tell us where to send it and we will email you a download link.

We use your details to send the link, and to tell you about new guides only if you tick the box. See the privacy policy.

A joint guide of Unplex® Technologies LLP and bioparticle. Created with Claude (Anthropic).
Educational reference material, not engineering, regulatory or clinical advice. Product and company names referenced are the property of their respective owners.