Python for bioinformatics
Right now you paste one sequence at a time into a web form. Python does the same job for ten thousand sequences while you make coffee. This path takes you from your first line of code to real analysis in BioPython.
Two courses. 119 lessons. About seventeen hours in total. Written for biologists who have never programmed.
Who this is for
You work at the bench
You can design an experiment and read a paper. The analysis goes to somebody else. This path hands that half of the job back to you.
You are doing an MSc or a PhD
Your data outgrew the spreadsheet months ago. Your supervisor keeps saying you should write a script. Start at zero here and finish able to do exactly that.
You already analyse data by hand
You use BLAST and Clustal and PDB through their websites. Learn to drive all three from Python and run a hundred jobs at once.
What you will learn
Write Python that reads like plain English
You start with variables and strings and loops and functions. Every example uses biological data. Nothing here is an abstract puzzle about cars or bank accounts.
Handle sequences the proper way
Seq objects behave like strings until you need them to behave like biology. You will transcribe and translate and mutate sequences. Alphabets stop you mixing DNA with protein by accident.
Read and write the file formats you already use
FASTA and GenBank and compressed archives all open through SeqIO. You will pull out the records you want and write the rest back to disk.
Search NCBI from your own code
Entrez gives you ESearch and EFetch and ESummary and ELink. BLAST runs and parses without the website. Your search becomes a script you can repeat.
Align sequences and build trees
You will call ClustalW from Python and align pairs with pairwise2. Distance matrices turn into phylogenetic trees you can draw and convert and edit in Bio.Phylo.
Open structures and find motifs
Bio.PDB reads a structure file straight into Python. Bio.motifs finds motifs and draws a WebLogo and reads MEME output.
What you'll be able to do
You will write your own Python instead of borrowing someone else's script. You will read and write sequence files of any size. You will search NCBI and run BLAST from your own code. You will align sequences and build phylogenetic trees and open PDB structures. Work that used to take an afternoon will take seconds.
The path · 2 courses in order
-
1
$49.00 Add to cart -
2
$89.00 Add to cart
Tools you will be using
- Python 3
- BioPython
- Anaconda & Jupyter
- Bio.Seq
- SeqIO
- AlignIO
- Bio.Entrez
- NCBI BLAST
- ClustalW
- Bio.PDB
- Bio.Phylo
- Bio.motifs
What you will have built by the end
- A script that opens a multi-record FASTA file and pulls out every sequence that matches your pattern
- A script that searches NCBI for a gene and downloads the records and summarises them in a table
- An alignment run through ClustalW from Python and inspected without leaving your editor
- A phylogenetic tree built from a distance matrix and exported as a figure
Every lesson in order
1 Python for Bioinformatics: Fundamentals
Python 23 lessons
- Why Python in Bioinformatics
- Introduction to Python and it's Installation
- Comments
- Basic Input and output
- Mathematical Operations
- Strings
- Dictionaries
- Lists
- + 15 more lessons in this section
BIoPython 48 lessons
- Introduction to BioPython & Installation
- Bio.Seq Create a Seq Object
- Bio.Seq Seq Object Behaves Like a String
- Bio.Seq Central Dogma in Play Through Python
- Bio.Seq Unknown & Mutable Sequences
- Bio.Alphabet Understanding the Alphabets of Biology
- Bio.Alphabet IUPAC and Types of Sequence Representations
- Bio.Alphabet Concatenation of Multiple Seq Records Using Generic Alphabets
- + 40 more lessons in this section
2 Advanced Python for Bioinformatics
Introduction 1 lessons
- Introduction to BioPython & Installation
Sequence Analysis 12 lessons
- Bio.Seq Create a Seq Object
- Bio.Seq Seq Object Behaves Like a String
- Bio.Seq Central Dogma in Play Through Python
- Bio.Seq Unkown & Mutable Sequences
- Bio.Alphabet Understanding the Alphabets of Biology
- Bio.Alphabet IUPAC and Types of Sequence Representations
- Bio.Alphabet Concatenation of Multiple Seq Records Using Generic Alphabets
- SeqRecord Creating Seq Records
- + 4 more lessons in this section
Sequence Data Parsing 4 lessons
- SeqIO Reading a Sequence File
- SeqIO Parsing a Sequence File
- SeqIO Parsing a Compressed Sequence File & Creating a Dictionary of Sequences
- SeqIO - Write Sequences and SeqRecords Into Files
Sequence Data Extraction 1 lessons
- SeqIO Extracting Annotations and Pattern-wise Sequence Data Extraction
Alignment Parsing and Analysis 9 lessons
- AlignIO - Reading and Parsing a Multiple Sequence Alignment File
- AlignIO - Writing Alignments and Multiple Sequence Alignment Records
- AlignIO - Conversion of Alignment Formats
- AlignIO - Manipulating Alignments
- AlignIO - ClustalW Python Wrapper - Align Multiple Sequences
- AlignIO - Pairwise2 - Align Two Sequences
- AlignIO - Information Mapping of Alignments
- AlignIO - Format Alignments
- + 1 more lessons in this section
BLAST Database Searching 1 lessons
- Bio.Blast - Querying NCBI BLAST Through Python
Parsing BLAST results 1 lessons
- Bio.Blast - Parsing BLAST Results
Biological Data Retrieval 10 lessons
- Bio.Entrez - Accessing ENTREZ Using Python
- Bio Entrez Use Esummary To Get Summary Of Your Accessions
- Bio.Entrez - Use EFetch to Download Complete Records
- Bio.Entrez - Use EGQuery to Do Global Queries for Search Counts
- Bio.Entrez - Use Elink To Search For Database Links Of Records
- Bio.Entrez - Use ESearch to Search the Entrez Databases
- Bio.Entrez - Use Espell To Get Correct Spellings For Your Search Terms
- Bio.Entrez - Download GenBank and Entrez Records
- + 2 more lessons in this section
Parsing a PDB Structure file 1 lessons
- Bio.PDB - Reading a PDB (3D Structure) File
Phylogenetic Analysis 6 lessons
- Bio.Phylo - Calculating Distance Matrix Between Sequences For Phylogenetic Analysis
- Bio.Phylo - Converting Phylogenetic Tree Data Formats
- Bio.Phylo - Printing Out Phylogenetic Tree in ASCII
- Bio.Phylo - Reading Phylogenetic Trees
- Bio.Phylo - Visualization And Manipulation Of Phylogenetic Trees
- Bio.Phylo - Writing Out Phylogenetic Data
Protein Sequence Analysis 2 lessons
- Bio.motifs - Creating a WebLogo of Motifs
- Bio.motifs - MEME Analysis
Questions people ask before they start
I have never written a line of code. Is this too advanced?
No. The first course begins by installing Python and printing one line of output. Every lesson after that uses biological data so you are never learning programming in the abstract.
Do I need both courses?
Begin with Fundamentals. Advanced picks up where it stops and covers the BioPython modules you need for real work. Buy the two together and the bundle price applies by itself. A membership includes both.
Will this work on Windows?
Yes. Windows and macOS and Linux are all covered. Setup uses Anaconda so you never have to fight your system installation of Python.
How long will it take me?
There are about seventeen hours of video. Most people finish inside a month at five hours a week. Your access never expires so you can take longer.
Do I get a certificate?
Yes. Each course issues a PDF certificate with its own verification number when you finish it.
What if the course is not for me?
Write to support@biocode.org.uk inside seven days and we will refund you. A membership can be cancelled whenever you like.
Ready to start?
Buy the courses on their own or open the whole library with a membership. This path is included in both plans.