omics.co.in

File Input and Output

🐍 Downloadable Lesson Python Script 3.6 KB

ch05_fasta_csv_io.py

Memory-efficient buffered generator for multi-line FASTA records and CSV/TSV expression matrix parser calculating per-gene mean and variance.

$ python ch05_fasta_csv_io.py --fasta sample.fasta

File Input and Output in Bioinformatics

Biological data formats such as FASTA, FASTQ, SAM/BAM, BED, GFF3, and VCF are primarily text or compressed text files. Effective reading, streaming, and writing of files prevents memory overflows during large-scale analysis.

# Reading large TSV gene expression files line by line
with open("expression_matrix.tsv", "r") as infile, open("significant_genes.tsv", "w") as outfile:
    header = infile.readline()
    outfile.write(header)
    
    for line in infile:
        parts = line.strip().split("\t")
        gene_id = parts[0]
        log2_fc = float(parts[1])
        padj = float(parts[2])
        
        # Filter for significantly upregulated genes
        if log2_fc >= 2.0 and padj <= 0.01:
            outfile.write(line)
← Return to Home Hub Scroll to Top ↑