omics.co.in

12: Machine Learning for Biology

🐍 Downloadable Lesson Python Script 2.9 KB

ch12_ml_cancer_classifier.py

Scikit-learn pipeline for gene expression: PCA dimensionality reduction, Random Forest Classifier, 5-fold cross-validation, and predictive biomarkers.

$ python ch12_ml_cancer_classifier.py

Chapter 12: Machine Learning for Biology

Machine Learning (ML) has become central to modern genomics—powering neural-network variant callers (DeepVariant), AlphaFold structural predictions, and clinical tumor classification. This chapter covers the machine learning workflow in biological research using Scikit-Learn.

1. Machine Learning Algorithms for Genomics

Biological data presents unique ML challenges: high dimensionality (thousands of genes but tens of patients), class imbalance, and biological noise. Standard models include Random Forests, Support Vector Machines (SVM), and Lasso/Ridge Regularized Linear Models.

2. The Scikit-Learn Pipeline

Constructing reproducible ML pipelines with preprocessing, feature scaling, and stratified cross-validation:

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import StratifiedKFold, cross_val_score

# Construct clinical classification pipeline
pipeline = Pipeline([
    ("scaler", StandardScaler()),
    ("classifier", RandomForestClassifier(n_estimators=100, random_state=42))
])

3. Applications of Machine Learning to Biology

Real-world applications in biology include: predicting promoter and splice sites from sequence k-mers, classifying antibiotic resistance from bacterial genomes, and scoring variant pathogenicity.

4. Case Study: Classification of Cancer Subtypes

Complete script predicting breast cancer tumor subtypes from RNA-seq expression profiles:

import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report, confusion_matrix

# Synthetic genomic expression data: 100 samples, 50 biomarker genes
np.random.seed(42)
X = np.random.randn(100, 50)
# Subtypes: 0 = Luminal A, 1 = Basal-like
y = np.random.choice([0, 1], size=100)

# Stratified split into training (80%) and test (20%) cohorts
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)

# Train Random Forest Classifier
clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_train, y_train)

# Evaluate predictions
y_pred = clf.predict(X_test)
print("Classification Report:\n", classification_report(y_test, y_pred, target_names=["Luminal A", "Basal-like"]))

# Extract top biomarker genes by Gini importance
importances = clf.feature_importances_
top_gene_indices = np.argsort(importances)[::-1][:5]
print("Top 5 Predictive Biomarker Feature Indices:", top_gene_indices)
← Return to Home Hub Scroll to Top ↑