omics.co.in
🐍 25 Verified Python Scripts Course Code + Genomics Library Standalone CLI & Modular

Python Scripts for Bioinformatics & Genomics

Download tested, production-grade Python 3 scripts for next-generation sequencing, variant filtering, RNA-Seq transcriptomics, structural modeling, and all 13 chapters of the Python for Next Gen Biologist curriculum. Each script includes full documentation, standalone command-line arguments, and pure-Python fallbacks.

⚡
Showing 25 of 25 Python scripts

📚 Course Curriculum Scripts (Chapters 1–13)

13 Lessons
Chapter 1 ch01_dna_basics.py (3 KB)

Chapter 1: DNA Basics, GC Content & Reverse Complement

Foundational genomics script calculating GC percentage, nucleotide frequencies, transcription to mRNA (T->U), and 5' to 3' reverse complement generation.

Dependencies: Python 3 standard library
$ python ch01_dna_basics.py --seq ATGCGATCGATC
Download .py
Chapter 2 ch02_data_structures.py (3.4 KB)

Chapter 2: Genomic Data Structures (Codons, Motifs & k-mers)

Implements dictionaries for genetic codon translation, sets for unique k-mer extraction, and lists of tuples for restriction enzyme motif scanning.

Dependencies: Python 3 standard library
$ python ch02_data_structures.py --seq ATGCGATCGATCGATCGATAGCTAGCTA
Download .py
Chapter 3 ch03_functions_orfs.py (3.4 KB)

Chapter 3: Modular Functions & 6-Frame ORF Scanner

Defines reusable functions to scan all 6 reading frames for Open Reading Frames (ORFs) and recursive Hamming edit distance calculation.

Dependencies: Python 3 standard library
$ python ch03_functions_orfs.py --seq ATGAAACCCGGGTTTTAA --min-aa 10
Download .py
Chapter 4 ch04_fastq_stream_parser.py (3.4 KB)

Chapter 4: Streaming FASTQ Parser & Quality Filter

Stream-parses raw 4-line FASTQ records, converts ASCII Phred+33 scores, computes mean read quality, and filters reads exceeding Q30 thresholds.

Dependencies: Python 3 standard library (gzip)
$ python ch04_fastq_stream_parser.py -i reads.fastq.gz -q 30 -l 50
Download .py
Chapter 5 ch05_fasta_csv_io.py (3.6 KB)

Chapter 5: Multi-line FASTA & Expression Matrix File I/O

Memory-efficient buffered generator for multi-line FASTA records and CSV/TSV expression matrix parser calculating per-gene mean and variance.

Dependencies: Python 3 standard library (csv, math)
$ python ch05_fasta_csv_io.py --fasta sample.fasta
Download .py
Chapter 6 ch06_dna_oop.py (6.4 KB)

Chapter 6: Object-Oriented Biological Sequences (OOP)

DnaSequence, ProteinSequence, and GenomicVariant classes implementing object inheritance, molecular weight calculations, and variant mutagenesis.

Dependencies: Python 3 standard library
$ python ch06_dna_oop.py
Download .py
Chapter 7 ch07_modules_packages.py (2.9 KB)

Chapter 7: Modular Genomics Toolkit Architecture

Demonstrates modular library design, IUPAC degenerate nucleotide expansion, and dinucleotide frequency profiling (CpG island detection).

Dependencies: Python 3 standard library
$ python ch07_modules_packages.py --seq ATGCGATCGATC --motif RGYW
Download .py
Chapter 8 ch08_regex_data_cleaning.py (3.6 KB)

Chapter 8: Genomic Regular Expressions & Metadata Cleaning

Pattern matching for C2H2 zinc fingers and bacterial Pribnow boxes, plus regex standardization of messy clinical patient metadata.

Dependencies: Python 3 standard library (re, csv)
$ python ch08_regex_data_cleaning.py
Download .py
Chapter 9 ch09_numpy_pandas_viz.py (4.3 KB)

Chapter 9: NumPy Matrices, Pandas DataFrames & Visualization

NumPy Position Weight Matrix (PWM) log-odds scoring, Pandas differential expression filtering, and Volcano plot generation.

Dependencies: numpy, pandas, matplotlib
Install: pip install numpy pandas matplotlib
$ python ch09_numpy_pandas_viz.py -o volcano.png
Download .py
Chapter 10 ch10_biopython_toolkit.py (2.8 KB)

Chapter 10: Biopython Sequence Analysis & Entrez Retrieval

Biopython Seq and SeqRecord objects, automated NCBI Entrez fetching, BLAST XML hit parsing, and PDB coordinate exploration.

Dependencies: biopython
Install: pip install biopython
$ python ch10_biopython_toolkit.py
Download .py
Chapter 11 ch11_rnaseq_deseq_analysis.py (3.4 KB)

Chapter 11: RNA-Seq Differential Expression & Statistical Testing

Two-sample Welch's t-test, Benjamini-Hochberg False Discovery Rate (FDR) correction, and protein co-expression network graph creation.

Dependencies: numpy, pandas, scipy, networkx
Install: pip install numpy pandas scipy networkx
$ python ch11_rnaseq_deseq_analysis.py
Download .py
Chapter 12 ch12_ml_cancer_classifier.py (2.9 KB)

Chapter 12: Machine Learning for Cancer Subtype Classification

Scikit-learn pipeline for gene expression: PCA dimensionality reduction, Random Forest Classifier, 5-fold cross-validation, and predictive biomarkers.

Dependencies: scikit-learn, numpy, pandas
Install: pip install scikit-learn numpy pandas
$ python ch12_ml_cancer_classifier.py
Download .py
Chapter 13 ch13_hpc_parallel_genomics.py (2.7 KB)

Chapter 13: High-Performance Computing & Genetic Drift Simulation

Multi-core parallel computing using multiprocessing.Pool, and Wright-Fisher population genetics simulation of neutral genetic drift.

Dependencies: Python 3 standard library
$ python ch13_hpc_parallel_genomics.py
Download .py

🧬 General Bioinformatics & Genomics Tools

12 Production Utilities
Sequencing & Assembly fastq_qc_trimmer.py (4.1 KB)

Streaming FASTQ Quality Trimmer & Adapter Clipper

High-throughput FASTQ sliding-window quality trimmer with 3' adapter removal and length thresholding. Supports gzipped streams.

Dependencies: Python 3 standard library (gzip)
$ python fastq_qc_trimmer.py -i reads.fastq.gz -o trimmed.fastq.gz --min-q 25 --window 4
Download .py
Sequencing & Assembly fasta_assembly_metrics.py (3.6 KB)

Genome Assembly Metrics Calculator (N50, L50, GC Skew)

Computes essential de novo genome assembly quality metrics: N50, L50, N90, L90, min/max/mean contig lengths, and overall GC content.

Dependencies: Python 3 standard library
$ python fasta_assembly_metrics.py -i contigs.fasta --min-len 500
Download .py
Genomics & Variants vcf_variant_filter.py (4.6 KB)

VCF Variant Filter & Ti/Tv Ratio Calculator

Stream-parses VCF files, filters variants by depth (DP), quality (QUAL), and allele frequency (AF), computes Ti/Tv ratio, and outputs a clean TSV.

Dependencies: Python 3 standard library (gzip)
$ python vcf_variant_filter.py -i variants.vcf.gz -o filtered.tsv --min-dp 20 --min-qual 30
Download .py
Sequencing & Assembly sam_bam_depth_coverage.py (2.7 KB)

SAM/BAM Alignment Depth & Coverage Analyzer

Analyzes sorted BAM alignments to compute genome-wide mapping rate, mean coverage depth, and MAPQ score distributions.

Dependencies: pysam (optional, pure-Python fallback included)
Install: pip install pysam
$ python sam_bam_depth_coverage.py -i alignment.bam
Download .py
Transcriptomics rnaseq_matrix_normalizer.py (2.8 KB)

RNA-Seq Count Matrix Normalizer (TPM, RPKM, CPM)

Normalizes raw RNA-Seq read counts matrix to Counts Per Million (CPM), RPKM, and Transcripts Per Million (TPM) with optional log2 scaling.

Dependencies: numpy, pandas
Install: pip install numpy pandas
$ python rnaseq_matrix_normalizer.py -c raw_counts.tsv -l lengths.tsv --method tpm --log2
Download .py
Transcriptomics volcano_ma_plot.py (3.5 KB)

Publication-Grade Volcano & MA Plot Generator

Reads differential expression results (DESeq2/edgeR/limma) and renders high-resolution Volcano and MA plots with automatic gene labeling.

Dependencies: matplotlib, pandas, numpy
Install: pip install matplotlib pandas numpy
$ python volcano_ma_plot.py -i deseq_results.tsv -o volcano.png --fc 1.5 --fdr 0.05
Download .py
Genomics & Variants gtf_gff_transcript_parser.py (3.3 KB)

GTF/GFF3 Feature Extractor & Longest Isoform Filter

Parses Ensembl/GENCODE annotations, computes exon and CDS spans, selects the principal longest transcript isoform, and exports BED tracks.

Dependencies: Python 3 standard library
$ python gtf_gff_transcript_parser.py -i annotation.gtf -o transcripts.bed
Download .py
Phylogenetics & Structure phylo_upgma_distance_matrix.py (3.4 KB)

Phylogenetic Jukes-Cantor Distance & UPGMA Tree Builder

Calculates pairwise Jukes-Cantor evolutionary distances from multiple sequence alignments and builds hierarchical UPGMA trees in Newick format.

Dependencies: Python 3 standard library
$ python phylo_upgma_distance_matrix.py -i alignment.fasta
Download .py
Genomics & Variants pwm_motif_promoter_scanner.py (2.6 KB)

Position Weight Matrix (PWM) Regulatory Motif Scanner

Constructs log2-odds Position Weight Matrices from transcription factor binding sites and scans promoter regions for regulatory motifs.

Dependencies: Python 3 standard library
$ python pwm_motif_promoter_scanner.py -m motifs.txt -s promoter.fasta
Download .py
Phylogenetics & Structure pdb_rmsd_contact_map.py (3.1 KB)

Protein PDB Structural RMSD & Contact Map Analyzer

Extracts 3D C-alpha coordinates from PDB structure files, computes conformational RMSD, and generates residue-residue spatial contact maps.

Dependencies: Python 3 standard library (math)
$ python pdb_rmsd_contact_map.py -p structure.pdb
Download .py
Public Databases & APIs ncbi_entrez_batch_fetcher.py (3.7 KB)

NCBI Entrez Automated Batch Downloader

Batch retrieves nucleotide FASTA, GenBank records, and PubMed abstracts directly from NCBI E-utilities API with rate-limiting.

Dependencies: Python 3 standard library (urllib, xml)
$ python ncbi_entrez_batch_fetcher.py --db nuccore --ids NM_000518.5 -o out.fasta
Download .py
Genomics & Variants kmer_entropy_profiler.py (2.7 KB)

Genomic k-mer Frequency & Shannon Entropy Profiler

Counts canonical k-mer frequency spectra (k=2-6), identifies overrepresented sequences, and computes sliding-window Shannon entropy.

Dependencies: Python 3 standard library
$ python kmer_entropy_profiler.py -i genome.fasta -k 4
Download .py