10: Introduction to Bioinformatics
Chapter 10: Introduction to Bioinformatics with Python
Bioinformatics connects computer science, statistics, and molecular biology. This chapter covers the industry-standard Biopython toolkit, programmatically querying primary biological databases (NCBI, Ensembl, UniProt), executing automated BLAST searches, and parsing 3D macromolecular structures (PDB).
1. Overview of Bioinformatics & Biological Databases
Modern biological discovery relies on interconnected public repositories: NCBI GenBank (nucleotides), UniProt (protein functions & domains), Ensembl (annotated vertebrate genomes), and the Protein Data Bank (PDB, experimental 3D structures).
2. Querying Biological Databases Programmatically
Biopython’s Bio.Entrez module automates literature and sequence retrieval from NCBI directly from Python:
from Bio import Entrez, SeqIO
Entrez.email = "researcher@university.edu" # Always provide an email per NCBI policy
# Search PubMed for CRISPR and Sickle Cell disease
handle = Entrez.esearch(db="pubmed", term="CRISPR Sickle Cell Disease", retmax=5)
record = Entrez.read(handle)
handle.close()
print("PubMed Article IDs:", record["IdList"])
# Download GenBank record for Human Beta-Globin (HBB)
fetch_handle = Entrez.efetch(db="nucleotide", id="NM_000518.5", rettype="gb", retmode="text")
seq_record = SeqIO.read(fetch_handle, "genbank")
fetch_handle.close()
print(f"Gene: {seq_record.name} | Description: {seq_record.description[:60]}... | Length: {len(seq_record)} bp")3. Biopython Library for Sequence Manipulation
Using Bio.Seq and Bio.SeqIO for rapid FASTA processing, transcription, and translation:
from Bio.Seq import Seq
coding_dna = Seq("ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG")
messenger_rna = coding_dna.transcribe()
protein = coding_dna.translate(to_stop=True)
print("DNA:", coding_dna)
print("mRNA:", messenger_rna)
print("Peptide:", protein)4. BLAST Searches & Sequence Alignment
Running automated BLAST queries against NCBI databases and parsing alignment hits:
from Bio.Blast import NCBIWWW, NCBIXML
# Query unknown sequence against NCBI non-redundant nucleotide database
query_seq = "TGGGCCTCATCTTCCTCTTCCTCTTCTT"
result_handle = NCBIWWW.qblast("blastn", "nt", query_seq)
blast_record = NCBIXML.read(result_handle)
for alignment in blast_record.alignments[:3]:
for hsp in alignment.hsps:
print(f"Hit: {alignment.title[:50]} | E-value: {hsp.expect:.2e} | Identities: {hsp.identities}/{hsp.align_length}")5. Protein Structure Analysis & Prediction (PDB)
Parsing 3D atomic coordinates from Protein Data Bank (.pdb / .cif) files and calculating spatial distances between residues:
from Bio.PDB import PDBParser
parser = PDBParser(QUIET=True)
structure = parser.get_structure("1A8O", "1A8O.pdb")
for model in structure:
for chain in model:
print(f"Chain {chain.id} contains {len(chain)} residues")
for residue in list(chain)[:3]:
ca_atom = residue["CA"] if "CA" in residue else None
if ca_atom:
print(f" Residue {residue.get_resname()}{residue.id[1]} Alpha-Carbon: {ca_atom.get_coord()}")