omics.co.in

10: Introduction to Bioinformatics

🐍 Downloadable Lesson Python Script 2.8 KB

ch10_biopython_toolkit.py

Biopython Seq and SeqRecord objects, automated NCBI Entrez fetching, BLAST XML hit parsing, and PDB coordinate exploration.

$ python ch10_biopython_toolkit.py

Chapter 10: Introduction to Bioinformatics with Python

Bioinformatics connects computer science, statistics, and molecular biology. This chapter covers the industry-standard Biopython toolkit, programmatically querying primary biological databases (NCBI, Ensembl, UniProt), executing automated BLAST searches, and parsing 3D macromolecular structures (PDB).

1. Overview of Bioinformatics & Biological Databases

Modern biological discovery relies on interconnected public repositories: NCBI GenBank (nucleotides), UniProt (protein functions & domains), Ensembl (annotated vertebrate genomes), and the Protein Data Bank (PDB, experimental 3D structures).

2. Querying Biological Databases Programmatically

Biopython’s Bio.Entrez module automates literature and sequence retrieval from NCBI directly from Python:

from Bio import Entrez, SeqIO

Entrez.email = "researcher@university.edu"  # Always provide an email per NCBI policy

# Search PubMed for CRISPR and Sickle Cell disease
handle = Entrez.esearch(db="pubmed", term="CRISPR Sickle Cell Disease", retmax=5)
record = Entrez.read(handle)
handle.close()
print("PubMed Article IDs:", record["IdList"])

# Download GenBank record for Human Beta-Globin (HBB)
fetch_handle = Entrez.efetch(db="nucleotide", id="NM_000518.5", rettype="gb", retmode="text")
seq_record = SeqIO.read(fetch_handle, "genbank")
fetch_handle.close()

print(f"Gene: {seq_record.name} | Description: {seq_record.description[:60]}... | Length: {len(seq_record)} bp")

3. Biopython Library for Sequence Manipulation

Using Bio.Seq and Bio.SeqIO for rapid FASTA processing, transcription, and translation:

from Bio.Seq import Seq

coding_dna = Seq("ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG")
messenger_rna = coding_dna.transcribe()
protein = coding_dna.translate(to_stop=True)

print("DNA:", coding_dna)
print("mRNA:", messenger_rna)
print("Peptide:", protein)

4. BLAST Searches & Sequence Alignment

Running automated BLAST queries against NCBI databases and parsing alignment hits:

from Bio.Blast import NCBIWWW, NCBIXML

# Query unknown sequence against NCBI non-redundant nucleotide database
query_seq = "TGGGCCTCATCTTCCTCTTCCTCTTCTT"
result_handle = NCBIWWW.qblast("blastn", "nt", query_seq)

blast_record = NCBIXML.read(result_handle)
for alignment in blast_record.alignments[:3]:
    for hsp in alignment.hsps:
        print(f"Hit: {alignment.title[:50]} | E-value: {hsp.expect:.2e} | Identities: {hsp.identities}/{hsp.align_length}")

5. Protein Structure Analysis & Prediction (PDB)

Parsing 3D atomic coordinates from Protein Data Bank (.pdb / .cif) files and calculating spatial distances between residues:

from Bio.PDB import PDBParser

parser = PDBParser(QUIET=True)
structure = parser.get_structure("1A8O", "1A8O.pdb")

for model in structure:
    for chain in model:
        print(f"Chain {chain.id} contains {len(chain)} residues")
        for residue in list(chain)[:3]:
            ca_atom = residue["CA"] if "CA" in residue else None
            if ca_atom:
                print(f"  Residue {residue.get_resname()}{residue.id[1]} Alpha-Carbon: {ca_atom.get_coord()}")
← Return to Home Hub Scroll to Top ↑