NGS in Cancer: How Tumour Sequencing Works and What It Finds

Next generation sequencing reads the DNA of a tumour and compares it against the patient’s healthy DNA. The differences that are left are the mutations the cancer acquired. A handful of them explain why the tumour grows and which drugs might stop it.

That is the whole idea. Everything below is how it is done and where it goes wrong.

Why cancer needs its own kind of sequencing

A tumour is not a separate organism. It is the patient’s own tissue with damage in it. So reading the tumour alone tells you very little. Every person carries millions of harmless differences from the reference genome and those differences show up in the tumour sample too.

The way around this is simple. You sequence two samples from the same patient. One is the tumour. The other is normal tissue or a blood sample. Anything present in both came with the patient and is called a germline variant. Anything present only in the tumour was acquired and is called a somatic variant. Cancer sequencing is the hunt for the somatic ones.

This is the single biggest difference from rare disease sequencing and it shapes the whole pipeline.

How a tumour is sequenced, step by step

Tumour sequencing workflow: tumour and normal samples go through library prep and sequencing, are aligned, compared at somatic variant calling, then annotated and interpreted
Steps 1 is lab work. Steps 2 to 5 are bioinformatics.

The lab fragments the DNA and prepares a library. The sequencer reads millions of short fragments. From there the work is computational. Reads are aligned to the reference genome. The aligned tumour is compared against the aligned normal. Variant callers written for cancer look for positions where the two disagree. What comes out is a long list of candidate mutations that then has to be filtered and interpreted.

A typical exome run produces tens of thousands of raw candidates. A clinical report names perhaps five. The distance between those two numbers is the job.

Panel or exome or whole genome

Three choices, and they trade depth against breadth.

Targeted panels read a few hundred genes that are already known to matter in cancer. They are cheap and they read each position hundreds or thousands of times, so they can detect a mutation present in only a small fraction of the cells. Most hospital testing uses panels.

Whole exome sequencing reads every protein coding region, about 1% of the genome. It finds mutations in genes nobody thought to put on a panel. Depth is lower so very rare subclones can be missed.

Whole genome sequencing reads everything including the regions between genes. It is the only way to see large structural rearrangements and mutations in regulatory regions properly. It costs the most and produces the most data to handle. We compare the two genome-wide options in detail in whole genome versus whole exome sequencing for cancer.

Why tumour samples are harder to analyse

Three problems come up in every project.

Purity. A biopsy is a mixture. Normal cells, immune cells and connective tissue come along with the tumour. If only 30% of the sample is cancer then a mutation present in every cancer cell appears in roughly 15% of the reads. Callers tuned for germline work expect 50% and will throw it away.

Heterogeneity. A tumour is not one clone. Different regions carry different mutations. A mutation in a quarter of the cancer cells sits at a very low fraction of reads and looks exactly like noise.

Artefacts. Formalin fixed tissue, which is what most pathology archives hold, damages DNA in a way that creates false mutations. So does PCR. Real cancer pipelines spend a lot of effort removing errors that look convincing.

This is why somatic calling uses different tools from germline calling. Reading the output needs the same awareness.

Drivers and passengers

Most somatic mutations do nothing. The cell was already dividing out of control and the copying errors piled up. Those are passengers.

A few mutations are the reason the cancer grows. Those are drivers, and they are what everybody wants. Telling them apart is partly a database question and partly a biology question. Is the gene a known oncogene or tumour suppressor. Does the mutation hit a functional domain. Has it been reported in other patients with the same cancer. Does a drug exist that targets it.

Annotation tools attach all of this to the variant list. The judgement at the end is human.

Where this is used in practice

Choosing a treatment. Some drugs only work when a specific mutation is present. Sequencing finds out before the patient starts a course that would not help them.

Understanding resistance. When a treatment stops working, sequencing the tumour again often shows a new mutation that explains why.

Watching for return. Tumour DNA circulates in blood. Sequencing a blood sample, often called liquid biopsy, can pick up traces of cancer coming back before a scan shows anything.

Research. Large sequencing projects are how the field learned which genes drive which cancers in the first place.

What you need to run this analysis yourself

The pipeline is not secret. Every step uses tools you can install today. What takes time is learning the order of them, the parameters that matter and how to tell a real mutation from an artefact.

You need three things. The Linux command line, because every tool runs there and the files are too big for a laptop spreadsheet. A working knowledge of alignment and quality control, so you can tell whether your data is sound before you trust the output. And practice on real tumour data with a matched normal, because that comparison is where cancer analysis differs from everything else.

Our hands-on courses follow exactly the pipeline in the figure above. Cancer Genomics: whole exome variant calling runs an exome case end to end. Cancer Genomics: whole genome variant calling does the same for a genome. Both start from raw reads and finish with an annotated and filtered variant list you can defend.

If you want the ground underneath them first, the NGS and variant calling path puts the courses in order, and a BioCode membership opens all of them along with the rest of the library.

New to the biology behind this? Start with molecular biology basics and come back.

Common questions

Does cancer sequencing need a blood sample as well as the tumour?
For somatic analysis it is strongly preferred. Without a matched normal you cannot tell an acquired mutation from one the patient was born with, and you fall back on population databases, which is far less reliable.

How deep does cancer sequencing need to be?
Deeper than germline work. Exomes are often run at 100x or more for the tumour, panels at several hundred to several thousand. The low fraction of reads carrying a real mutation is the reason.

Can RNA sequencing be used in cancer too?
Yes, and it answers a different question. DNA sequencing finds the mutations. RNA sequencing shows which genes are being expressed and can detect fusion transcripts that DNA alone can miss.

How long does the analysis take?
On a reasonable server an exome pipeline runs in hours and a genome in a day or so. The interpretation afterwards takes longer than the computing.