Pipelines & Software
CCBR develops and maintains a suite of standardized, reproducible bioinformatics pipelines for NGS data. All pipelines are version-controlled, containerized, and designed to run on NIH Biowulf using Snakemake or Nextflow workflow managers.
More information about Pipeliner is available at ccbr.github.io/Pipeliner. For the broader scientific software catalog maintained by BTEP, visit the BTEP Software List.
Available Pipelines on Biowulf
| # | Data Type | Pipeline GitHub link | Full Name | Notes |
|---|---|---|---|---|
| 1 | RNASeq | RENEE | Rna sEquencing aNalysis pipElinE | Comprehensive RNA-seq workflow that combines contamination screening, adapter trimming, two-pass STAR alignment, and RSEM quantification. It generates gene and isoform count matrices, fusion calls, and MultiQC-style quality summaries for cohort-level review. |
| 2 | WESeq | XAVIER | eXome Analysis and Variant explorER | Whole-exome sequencing pipeline aligned to Broad-style best practices using a reproducible Snakemake execution model. It accepts FASTQ or BAM inputs and supports germline and somatic variant calling, CNV inference, and downstream annotation. |
| 3 | ATACSeq | ASPEN | Atac Seq PipEliNe | ATAC-seq analysis workflow for preliminary QC, peak calling, and differential chromatin accessibility testing. It is configured for Biowulf usage and provides standardized outputs for regulatory genomics interpretation. |
| 4 | ChIPSeq | CHAMPAGNE | CHromAtin iMmuno PrecipitAtion sequencinG aNalysis pipEline | ChIP-seq analysis pipeline with support for configurable reference genomes and optional spike-in genome handling. It produces reproducible peak-centric outputs from aligned sequencing data and can be run through CCBR module-based workflows. |
| 5 | CRISPRSeq | CRISPIN | CRISPr screen sequencing analysis pipelINe | CRISPR screen sequencing workflow designed for reproducible run setup and execution on Biowulf. It processes screening libraries into analysis-ready outputs suitable for guide-level and gene-level interpretation. |
| 6 | CUT&RunSeq | CARLISLE | Cut And Run anaLysIS pipeLinE | CUT&RUN pipeline that performs peak calling with MACS2, SEACR, and GoPeaks using spike-in-normalized FASTQ inputs. It is built in Snakemake and tuned for reproducible chromatin profiling on Biowulf. |
| 7 | circRNASeq | CHARLIE | Circrnas in Host And viRuses anaLysis pIpEline | circRNA pipeline for detection, annotation, and quantification of host and viral circular RNAs. It orchestrates multiple circRNA callers in parallel, including CIRCExplorer2 and CIRI2-based paths, to support robust cross-method discovery. |
| 8 | scRNASeq | SINCLAIR | SINgle CelL AnalysIs Resource | Single-cell analysis resource supporting multiple next-generation modalities in a reproducible Nextflow framework. It starts from FASTQ or h5-aligned inputs, performs per-sample QC, and generates per-contrast integration reports. |
| 9 | WGSeq | LOGAN | whoLe genOme-sequencinG Analysis pipeliNe | Whole-genome sequencing pipeline based on Broad-aligned processing patterns and Nextflow orchestration. It calls and annotates germline and somatic variants, CNVs, and SVs with reproducible containerized execution. |
| 10 | EV-Seq | ESCAPE | Extracellular veSiCles rnAseq PipelinE | Extracellular vesicle RNA-seq workflow for EV-derived libraries from raw reads through analysis-ready quantification outputs. It emphasizes consistent QC and standardized processing for comparative EV transcriptomics. |
| 11 | spatialSeq | SPENCER | SPactial sequENCing Resources | Spatial sequencing workflow for processing location-aware transcriptomic measurements across tissue regions. It is designed to produce structured outputs for downstream spatial feature analysis, integration, and visualization. |
For any other data type or pipeline not listed here, email us directly to start the conversation.
Experimental Design Best Practices
Talk to us early. We often cannot save a poorly designed experiment, but we can help you design a successful one. We are committed to helping you determine the optimal experimental design given your budget, including sample size, number of replicates, sources of batch effects, sequencing depth, and overall bioinformatics considerations for the experiment.
General Principles
- Minimum 3 biological replicates per condition
- Plan for confounders: batch effects, sex, passage number, time of collection
- Include appropriate controls: input, IgG, vehicle-treated, matched normal
- Discuss sample size and statistical power with CCBR before proceeding
- Plan metadata collection early so study design, sample annotations, processed outputs, and raw files are ready for downstream analysis and GEO submission
- ≥3 biological replicates per condition (≥5–6 for primary patient tumors)
- ≥30–50M reads/sample; ≥100M for fusions or rare transcripts
- Poly-A for RIN ≥7; ribo-depletion for degraded RNA
- Never pool biological replicates
- Discuss batch layout if samples span multiple sequencing runs
- Use biological replicates (≥2 minimum, ≥3 preferred) and process all samples under identical conditions.
- Use fresh or properly preserved samples with high viability and minimize mitochondrial contamination.
- Avoid PCR over-amplification and monitor library complexity early.
- Use paired-end sequencing and aim for approximately 50M paired-end reads per sample.
- Distribute samples across batches and lanes to reduce confounding batch effects.
- Use biological replicates (minimum 2, preferably 3 or more) under the same workflow conditions.
- Use ChIP-grade, preferably ENCODE-validated antibodies and track lot numbers carefully.
- Always include input DNA and matched controls; use spike-ins when comparing global changes.
- Choose depth and read strategy based on target type, expected signal, and genome complexity.
- Consider CUT&RUN or CUT&Tag when working with low-input or weak-signal samples.
- Use biological replicates (minimum 2, preferably 3 or more) and process samples under identical conditions.
- Use validated antibodies and consistent input material with healthy, gently handled cells.
- Include IgG or no-antibody controls and maintain consistent spike-in ratios across samples.
- Prefer paired-end sequencing and design library prep to capture the short fragments produced by CUT&RUN.
- Balance samples across batches and avoid PCR over-amplification.
- Matched tumor and germline pairs are ideal for somatic variant calling.
- For mouse studies, matched tumor and germline samples from the same individual remain the preferred design; cohort-matched controls can be used with care.
- Aim for ≥100× mean target depth for tumor samples and ≥50× for germline samples.
- Always document the exome capture kit and retain the corresponding target BED file.
- For familial studies, consult before sequencing because sample selection strongly affects downstream analysis.
- Target ≥30× mean depth for germline analysis and ≥60–100× for somatic variant detection.
- Always include a matched normal for somatic analyses whenever possible.
- WGS is preferred when structural variation or copy-number variation is a major objective.
- PCR-free library preparation is preferred for reducing GC bias and improving coverage uniformity.
- For mouse studies, matched tumor and germline from the same individual remains the strongest design.
- Use at least 2 biological replicates per condition, with 3 or more preferred when feasible.
- Start with the highest-quality sample possible and aim for viability >80%.
- A practical target is approximately 5,000–10,000 recovered cells per sample.
- Start with a minimum of 20,000 read pairs per cell for gene-expression libraries.
- Use multiplexing strategies when appropriate to reduce batch effects and improve throughput.
- Select the appropriate workflow for the specimen type and perform a pilot study for new tissue types.
- Optimize section thickness, tissue placement, and avoid damaged or low-quality regions.
- Use independent biological replicates and distribute conditions across slides and batches.
- Plan sequencing depth based on assay type, tissue complexity, and tissue-covered spots.
- Capture metadata consistently and consider matched scRNA-seq or snRNA-seq reference data for downstream interpretation.
For further assistance in planning your NGS experiment, please contact nciccbr@mail.nih.gov or visit our Bethesda office hours on Thursdays from 3:00 to 5:00 PM. For cost information and assistance with experiment setup, please visit the Frederick Sequencing and Genomics Core (FSGC) website or contact Bao Tran at bao.tran@nih.gov.
NIDAP: NIH Integrated Data Analysis Platform
NIDAP is a cloud-based collaborative analysis platform that hosts user-friendly bioinformatics workflows, analysis tools, and visualization tools developed by the NCI community for researchers across NIH.
NIDAP 1.0, now retired, was a cloud-based analysis platform hosted by Palantir at NCI, providing a secure collaborative environment for running CCBR workflows, sharing data, and exploring results interactively. The Palantir-hosted NIDAP 1.0 platform retired on September 28, 2025.
NIDAP 2.0 currently under evaluation and is being developed to restore and improve standardized bioinformatics workflows for CCR researchers, with an emphasis on reproducibility, modular execution, and user-friendly scientific analysis. This next-generation approach is intended to preserve what researchers valued most in NIDAP: graphical access to sophisticated analyses, shareable workflows, and results that can be reproduced with confidence over time.
What You Can Do on NIDAP 2.0
- Run pre-built CCBR workflows such as bulk RNA-seq and scRNA-seq through a point-and-click interface
- Securely share results and processed data with collaborators inside and outside NCI
- Explore interactive visualizations such as heatmaps, volcano plots, and UMAP embeddings
- Access supported CCBR workflows in a collaborative cloud environment
What to Expect with NIDAP 2.0
- Standardized workflows for bulk RNA-seq, single-cell RNA-seq, spatial transcriptomics, and related analyses
- Self-contained execution environments designed to improve reproducibility over time
- User-focused workflow design for review, sharing, and reuse of results
- Lower technical barriers for researchers while improving the stability, transparency, and reusability of computational analyses
Why it matters: NIDAP 2.0 aims to lower technical barriers for researchers while improving the stability, transparency, and reusability of computational analyses.
Workflow Directions Under Evaluation for NIDAP 2.0
Galaxy
Galaxy is being evaluated as a framework for standardized bioinformatics pipelines in NIDAP 2.0. The current prototype emphasizes structured, multi-step workflows for count processing, normalization, differential analysis, and visualization of bulk RNA-seq data.
Code Ocean
Code Ocean is being explored for modular, self-contained workflows that support reproducible scientific analysis and more flexible user interaction. This direction supports current results review and future interactive workflow development.
To learn more about NIDAP genomic workflows and how to cite NIDAP for analyses conducted using bulk RNA-seq, single-cell RNA-seq, or Digital Spatial Profiling workflows, visit our NCI NIDAP Community SharePoint Site.