The Platform
How Genomarker works: upload and ingest, exploration, differential expression, enrichment, and machine learning — with confounder checks enforced at submit.
Genomarker organizes a biomarker study as one pipeline — upload, explore, analyze, enrich, predict — with rigor enforced between the stages instead of left to the user’s discipline.
Upload & ingest
Datasets arrive as expression matrices plus sample metadata: up to 1 GB via chunked, resumable upload; GEO series matrices directly from NCBI; microarray data with GPL probeset annotation; Excel files with automatic repair; and single-cell h5ad. Ingest screens for PHI, harmonizes gene symbols to current HGNC nomenclature, and routes count data (RNA-seq) and intensity data (microarray) down their correct statistical paths — rejecting inappropriate inputs, such as TPM values for count-based differential expression, instead of silently accepting them.
Explore
Before any model is fit, the data has to be understood. Interactive PCA and k-means and hierarchical clustering with heatmaps and dendrograms surface outliers, batch structure, and promising signals early, when they are cheap to act on. Normalization produces derived datasets with full lineage, so every downstream analysis traces back to exactly the data it ran on.
Analyze
Differential expression runs on two estimators matched to the assay: DESeq2 for RNA-seq counts and limma-trend for microarray intensities — the latter a pure-Python implementation value-verified against committed R 3.68.4 goldens. Covariate designs, repeated-measures blocking, and LFC shrinkage are first-class. Single-cell studies ingest h5ad and analyze through pseudobulk differential expression or MAST.
Confounder Pre-Flight
Every analysis passes through a submit-time gate. Confounder Pre-Flight screens the design for confounding — a surrogate-variable analysis, batch/outcome association tests, remeasurement detection — and hard-refuses a confounded design, returning the evidence and the fix instead of a quietly wrong result. The wow is the stop, not the speed.
Enrich
Over-representation analysis places significant features in biological context against Reactome, GO (BP/MF/CC), and MSigDB Hallmark gene sets, shipped as versioned reference bundles — an enrichment result always names the exact reference version it used.
Predict
Classifier development runs XGBoost and random forest with cross-validation, feature importance and stability, and model comparison. Every classifier carries rigor primitives by default: isotonic and Platt calibration, Brier score, decision-curve analysis, subgroup AUROC, holdout and external-validation detection, and cohort minimums.
Try Genomarker
Genomarker is live and free during early access. Upload a dataset and run your first analysis in minutes — no bioinformatics team required. Your data stays yours.
The source repository is currently private.