Sign inSign up

etycksen/expresso-explorer

By etycksen

•Updated 7 days ago

Shiny app for visualizing expresso results from v3, v4, and nf version of coffee+eXpresso pipelines

Image
Data science
0

165

etycksen/expresso-explorer repository overview

⁠eXpresso eXplorer

A lightweight R Shiny viewer for the results of the full GTAC@MGI eXpresso differential expression analysis pipeline. It reads the merged result workbooks the pipeline writes and lets you browse differential expression and GAGE gene set results for every contrast, with volcano plots and clustered expression heatmaps.

Nothing is re-analysed: every statistic comes from the pipeline, and the sliders only filter what is displayed. The app is built to run in 1 GB of memory or less.

⁠Running eXpresso eXplorer

Online (no installation)

A shared test deployment runs at https://etycksen-gtac-mgi.shinyapps.io/eXpresso_eXplorer/⁠. It cannot see files on your computer, so choose Upload files on the import page and select Merged_differential_expression_results.xlsx together with any Merged_GAGE_*.xlsx workbooks.

In Docker docker run -d -p 3838:3838 -m 1g -v /path/to/results:/data:ro etycksen/expresso-explorer:1.0.0

Then open http://localhost:3838⁠. The container starts in /data, opens on Folder on the server, and cannot browse outside /data. On Apple Silicon Macs add --platform linux/amd64 (it runs under emulation, so imports are about twice as slow).

⁠Environment variables
VariableEffect
EXPRESSO_RESULTS_DIRFolder the import page starts in (default: the working directory).
EXPRESSO_ROOTWhen set, scanning and browsing are confined to this folder.

⁠How to use eXpresso eXplorer

  1. Import Results: Upload files (the default) and select the result workbooks:

    • Two or more contrasts: Merged_differential_expression_results.xlsx, any Merged_GAGE_*.xlsx, and Sample_Weights.xlsx.
    • A single contrast: its …Differential_Expression.xlsx, any …GAGE_*_Results.xlsx, and Sample_Weights.xlsx. These GAGE workbooks each sit in their own subfolder, so in the file dialog open the results folder and search for .xlsx to list them all at once. Files the app does not use are ignored.

    Where the app runs alongside the results (locally, or in Docker with the results mounted), choose Folder on the server instead, then enter or Browse… to the results folder and click Scan folder. With Search subfolders ticked, folders up to four levels down are searched; Nextflow work folders are skipped.

  2. Pick the files: the differential expression results, and whichever gene set collections you want. Sample_Weights.xlsx is used automatically when present. Click ☕ Load eXpresso!

  3. One tab per contrast, each with two sub-tabs:

    • Differential Expression: summary tiles, volcano plot, heatmap of the top significant genes, and the full results table with downloads.
    • Gene Set Enrichment: choose a collection, then select a gene set to see a heatmap of its member genes and their statistics.
  4. Contrast Summary (two or more contrasts): which genes are shared or unique across contrasts, with direction tracked separately.

With more than five contrasts, the contrast tabs are grouped under a Contrasts menu.

⁠Adjusted or unadjusted p-values

Each Differential Expression tab, Gene Set Enrichment tab and the Contrast Summary has a Significance based on choice:

  • Adjusted p-value (FDR), the default: the pipeline's Benjamini–Hochberg adjusted p-value (adj.P.Val, or GAGE's FDR).
  • Unadjusted p-value: the raw P.Value (GAGE P.value). This is often more useful for small pilot studies, where FDR correction can leave nothing significant.

Genes or gene sets at or below the p-value threshold are significant, and the volcano plot, heatmap, tables, tiles and downloads all follow the choice. With unadjusted p-values the app shows how many genes would pass by chance alone (tested genes × threshold), as a reminder to treat those hits as leads to confirm.

⁠Input files

FileWhat the app uses
Merged_differential_expression_results.xlsxFeature_ID and gene annotation; <contrast>_{logFC, linearFC, CI.L, CI.R, P.Value, adj.P.Val} for each contrast; group_* group means; sample.* per-sample log2 expression (the heatmap matrix).
Merged_GAGE_*.xlsx (any number)Term_ID; <contrast>_{mean_logFC_v.background, P.value, FDR} for each contrast; Genes (comma-separated member symbols).
Sample_Weights.xlsx (optional)Sample and Group: the pipeline's own sample → group table.

Contrasts are discovered from the column names; nothing is hard-coded.

Single-contrast runs. The pipeline only writes the merged workbooks when a run has two or more contrasts. When a folder has no merged DE workbook, the app uses the per-contrast workbooks every run writes instead:

FileWhat the app uses
…Differential_Expression.xlsx (one per contrast)The same annotation and sample.* columns, with unprefixed logFC, linearFC, CI.L, CI.R, P.Value, adj.P.Val.
…GAGE_<collection>_Results.xlsx (one per contrast and collection)Term_ID, set.size, mean_logFC_v.background, P.value, FDR, Genes.

The contrast name comes from the file name (group_KO_Differential_Expression.xlsx → KO). A single contrast whose files carry no name (just Differential_Expression.xlsx) is named after its folder. These files are combined in memory into the merged layout, so everything else in the app works the same. A merged workbook always takes precedence when one exists.

Contrast name matching. Pipeline versions before the fix in coffee_eXpresso-nf 743642c / brbseq-nf 344ba6c stripped group_ from every DE contrast name but only from the start of GAGE contrast names, so the same contrast appears as KO-control in the DE file and KO-group_control in the GAGE files. The app matches them with group_ removed, so results from either version load, and reports any GAGE contrast it cannot match on the import page.

Sample groups. Heatmaps default to the samples in the two groups being compared. Groups come from the pipeline's Sample_Weights.xlsx when the results folder has one. Otherwise samples are assigned by name where the name matches a group_* column, and then to the group whose mean expression profile they are closest to (for example Ctrl_1 becomes control). A contrast that names only one group, such as a reference-design coefficient (KO, meaning KO vs the reference level), shows all samples. The assignment is shown on the import page, and you can double-click a group to correct it.

Gene set members are matched to the DE table by gene symbol, falling back to a case-insensitive match when an exact match finds fewer than half of the genes.

Tag summary

Content type

Image

Digest

sha256:c7f34e6fe…

Size

620.9 MB

Last updated

7 days ago

docker pull etycksen/expresso-explorer