12: Machine Learning for Biology
Chapter 12: Machine Learning for Biology
Machine Learning (ML) has become central to modern genomics—powering neural-network variant callers (DeepVariant), AlphaFold structural predictions, and clinical tumor classification. This chapter covers the machine learning workflow in biological research using Scikit-Learn.
1. Machine Learning Algorithms for Genomics
Biological data presents unique ML challenges: high dimensionality (thousands of genes but tens of patients), class imbalance, and biological noise. Standard models include Random Forests, Support Vector Machines (SVM), and Lasso/Ridge Regularized Linear Models.
2. The Scikit-Learn Pipeline
Constructing reproducible ML pipelines with preprocessing, feature scaling, and stratified cross-validation:
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import StratifiedKFold, cross_val_score
# Construct clinical classification pipeline
pipeline = Pipeline([
("scaler", StandardScaler()),
("classifier", RandomForestClassifier(n_estimators=100, random_state=42))
])3. Applications of Machine Learning to Biology
Real-world applications in biology include: predicting promoter and splice sites from sequence k-mers, classifying antibiotic resistance from bacterial genomes, and scoring variant pathogenicity.
4. Case Study: Classification of Cancer Subtypes
Complete script predicting breast cancer tumor subtypes from RNA-seq expression profiles:
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report, confusion_matrix
# Synthetic genomic expression data: 100 samples, 50 biomarker genes
np.random.seed(42)
X = np.random.randn(100, 50)
# Subtypes: 0 = Luminal A, 1 = Basal-like
y = np.random.choice([0, 1], size=100)
# Stratified split into training (80%) and test (20%) cohorts
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
# Train Random Forest Classifier
clf = RandomForestClassifier(n_estimators=100, random_state=42)
clf.fit(X_train, y_train)
# Evaluate predictions
y_pred = clf.predict(X_test)
print("Classification Report:\n", classification_report(y_test, y_pred, target_names=["Luminal A", "Basal-like"]))
# Extract top biomarker genes by Gini importance
importances = clf.feature_importances_
top_gene_indices = np.argsort(importances)[::-1][:5]
print("Top 5 Predictive Biomarker Feature Indices:", top_gene_indices)