Skip to content

Command reference

Choose a command below to see when to use it, how to prepare its inputs, an example run, and the files it writes. The complete options follow each guide and are also available with <command> --help.

Command Purpose
prepare-atac Linux CLI or Linux container only. Download public ATAC-seq reads or use local FASTQ files, then trim, align, filter, call peaks, calculate alignment coverage, and write QC files. The GUI and native macOS/Windows installations start from coordinate-sorted BAM/BAI and matching peak BED files.
bulk-footprinting Run a complete bulk ATAC-seq analysis from aligned reads and peak regions to footprint scores, motif comparisons, and interactive reports.
atac-correct Correct ATAC-seq cut-site signal for Tn5 sequence bias. Start with coordinate-sorted BAM files, adjacent BAI indexes, and matching peak BED files. Use the corrected signal as the input to call-footprints.
call-footprints Calculate a footprint score at each position in accessible regions using the corrected bigWigs from atac-correct. The score measures local cut-site depletion relative to the surrounding signal; higher scores indicate stronger footprint evidence.
match-motifs Find motif matches in accessible regions and measure the footprint score at each match. This produces a motif summary for each sample and classifies sites as predicted bound or unbound.
diff-footprints Find motifs whose footprint scores differ between conditions. Start with the per-sample results from match-motifs. You can also compare two sets of genomic regions measured in the same samples.
normalize-bigwig Put corrected cut-site tracks on a comparable signal scale using the same background regions for every sample. This is an optional step after atac-correct; use the resulting tracks for downstream scoring or plotting when you need this normalization.
plot-aggregate Plot the average cut-site signal around motif sites to inspect the shape of a footprint. Start with corrected bigWigs from atac-correct and motif results from match-motifs. Outputs can be static figures or an interactive HTML report.
review-multi-comparisons Review several diff-footprints comparisons in one place. Combine the existing reports into a browser bundle or a single HTML file, then switch between comparisons to explore motif statistics and aggregate profiles.
run-yaml-workflow Run one or more fp-tools commands from a saved YAML configuration. Use this to repeat a GUI run or apply the same settings to several samples or comparisons.
fp-tools-gui Launch the browser interface for configuring and running fp-tools commands. The Windows and Apple Silicon desktop downloads present the same interface in a native fp-tools application window.
fp-tools-runtime Check or install the external programs used by prepare-atac and discover-motifs. fp-tools manages these programs separately from your other software. Linux supports read preparation and motif discovery; macOS and Windows support motif discovery.
discover-motifs Find recurring DNA sequence patterns in candidate footprint regions, without starting from a known motif list. Provide candidate intervals and a reference genome, or use sequences you have already extracted into a FASTA file.
summarize-motifs Turn completed motif-discovery results into a table of motif sequences, significance values, and optional matches to known motifs. Add an HTML report to review the results in a browser.
pseudobulk-fragments Combine single-cell ATAC-seq fragments into groups such as cell types or donor–cell-type pairs. Each group becomes a pseudobulk sample for downstream analysis.
find-signature-fp Score selected motif sites in individual cells and plot the results on a UMAP and in heatmaps. The command pools signal from nearby cells to reduce sparsity, so the scores describe relative footprint signatures rather than independent binding calls for every cell.
sc-footprinting Analyze single-cell ATAC-seq by combining cells into pseudobulk groups, calling footprints and motif differences between groups, then mapping selected footprint signatures back to individual cells.

prepare-atac

Linux CLI or Linux container only. Download public ATAC-seq reads or use local FASTQ files, then trim, align, filter, call peaks, calculate alignment coverage, and write QC files. The GUI and native macOS/Windows installations start from coordinate-sorted BAM/BAI and matching peak BED files.

Example command

prepare-atac --samples metadata.tsv --genome hg38 --outdir project

Primary inputs

The workflow uses all available cores by default, sharing its core budget across concurrent samples. Tool diagnostics appear live in the terminal and remain in the sample logs; alignment data continue to go to their output files.

  • --samples — TSV or CSV sample sheet. Provide sample, condition, and fastq_1 paths or URLs; add fastq_2 for paired-end reads. A public sequencing run can instead be supplied in run_accession.
  • --genome — managed hg38 or mm10 reference label, or a custom label used with explicit reference options.
  • --outdir — project directory represented by {project} below.

For paired-end local files, metadata.tsv can contain:

sample  condition   fastq_1 fastq_2
control_1   control reads/control_1_R1.fastq.gz reads/control_1_R2.fastq.gz
treated_1   treated reads/treated_1_R1.fastq.gz reads/treated_1_R2.fastq.gz

Repeated rows with the same sample, condition, and replicate combine technical sequencing runs. Different sample values sharing a condition are biological replicates.

Main outputs

For each {sample}, the default modern profile writes:

Path Meaning
{project}/samples/{sample}/alignment/{sample}.filtered.bam Coordinate-sorted, filtered ATAC-seq alignment used downstream.
{project}/samples/{sample}/alignment/{sample}.filtered.bam.bai Samtools index for the filtered BAM.
{project}/samples/{sample}/peaks/{sample}.narrowPeak MACS3 narrow-peak calls before project-level merging.
{project}/samples/{sample}/tracks/{sample}.rp10m.bw Sequencing-depth-normalized alignment coverage bigWig for viewing read coverage.
{project}/samples/{sample}/qc/{sample}.fastp.html Interactive Fastp read-trimming QC report.
{project}/samples/{sample}/qc/{sample}.fastp.json Machine-readable Fastp metrics.
{project}/samples/{sample}/qc/flagstat.tsv Samtools alignment and filtering counts.
{project}/samples/{sample}/qc/fragment_lengths.tsv Fragment-length distribution used to inspect ATAC-seq periodicity.
{project}/samples/{sample}/qc/metrics.json Consolidated per-sample QC metrics.
{project}/samples/{sample}/qc/commands.log External commands used for that sample.

Project-level files include:

Path Meaning
{project}/peaks/merged_peaks.bed Union of sample peak intervals.
{project}/peaks/merged_peaks_filtered.bed Analysis peak set after excluded chromosomes are removed.
{project}/metadata/resolved_runs.tsv Resolved local/downloaded FASTQ files and run grouping.
{project}/metadata/samples.tsv sample, condition, bam, and peaks table ready for bulk-footprinting.
{project}/reports/qc_summary.tsv Cross-sample QC summary.

Reference options

With --genome hg38 or --genome mm10, the command downloads and verifies the matching reference and blacklist as needed. Use --reference-dir to choose where references are cached.

For a custom genome label, provide --fasta and --macs-genome-size. Supply --bowtie2-index if you already have an index; otherwise the command builds one from the FASTA. --blacklist replaces the managed hg38/mm10 blacklist, while --no-blacklist disables it. Custom FASTA inputs use no blacklist unless you provide one.

Complete options

usage: prepare-atac [-h] [--samples SAMPLES] [--genome GENOME]
                    [--outdir OUTDIR] [--config CONFIG]
                    [--profile {modern,homer-atac}] [--id-column ID_COLUMN]
                    [--sample-column SAMPLE_COLUMN]
                    [--condition-column CONDITION_COLUMN]
                    [--include INCLUDE [INCLUDE ...]]
                    [--reference-dir REFERENCE_DIR] [--fasta FASTA]
                    [--bowtie2-index BOWTIE2_INDEX]
                    [--blacklist BLACKLIST | --no-blacklist] [--tss TSS]
                    [--macs-genome-size MACS_GENOME_SIZE] [--cores CORES]
                    [--max-parallel-samples MAX_PARALLEL_SAMPLES]
                    [--memory-gb MEMORY_GB] [--keep-intermediates]
                    [--no-resume] [--fail-fast] [--dry-run] [--doctor]
                    [--write-default-config PATH]
                    [--runtime {auto,managed,system,container}]

Download single- or paired-end ATAC-seq reads and prepare filtered BAM, peak
BED, sequencing-depth-normalized alignment coverage bigWig, and QC files.

options:
  -h, --help            show this help message and exit
  --samples SAMPLES     TSV/CSV sample metadata with an ID/run_accession or
                        fastq_1 column.
  --genome GENOME       Named genome (hg38/mm10) or custom genome label.
  --outdir OUTDIR       Output project directory.
  --config CONFIG       Optional preprocessing YAML; CLI values override
                        packaged defaults.
  --profile {modern,homer-atac}
                        Processing method: modern uses fastp, samtools, and
                        MACS3; homer-atac uses Trim Galore, Picard, and HOMER
                        (default: modern).
  --id-column ID_COLUMN
                        Explicit accession column name.
  --sample-column SAMPLE_COLUMN
                        Explicit sample-name column.
  --condition-column CONDITION_COLUMN
                        Explicit condition column.
  --include INCLUDE [INCLUDE ...]
                        Only process these sample names or accessions.
  --reference-dir REFERENCE_DIR
                        Reference cache root (default: ~/.cache/fp-
                        tools/references).
  --fasta FASTA         Custom reference FASTA.
  --bowtie2-index BOWTIE2_INDEX
                        Existing Bowtie2 index prefix.
  --blacklist BLACKLIST
                        Custom blacklist BED.
  --no-blacklist        Disable the managed blacklist for hg38 or mm10.
  --tss TSS             Optional TSS BED for enrichment QC.
  --macs-genome-size MACS_GENOME_SIZE
                        MACS3 genome size or hs/mm shorthand for custom
                        genomes.
  --cores CORES         Optional total core limit (default: all available
                        cores).
  --max-parallel-samples MAX_PARALLEL_SAMPLES
                        Maximum samples processed concurrently (default:
                        config value).
  --memory-gb MEMORY_GB
                        Enforced total memory budget in GiB; reserves 8 GiB
                        for the host.
  --keep-intermediates  Keep trimmed FASTQs and intermediate alignment files
                        under each sample .work directory.
  --no-resume           Recompute completed samples even when fingerprints
                        match.
  --fail-fast           Stop after the first failed sample.
  --dry-run             Validate and list the planned samples without
                        downloading or processing.
  --doctor              Report external preprocessing dependencies and exit.
  --write-default-config PATH
                        Write the fully documented default YAML and exit.
  --runtime {auto,managed,system,container}
                        External-tool runtime: auto/managed provisions the
                        pinned fp-tools runtime, system uses PATH, and
                        container uses the complete image (default: auto).

bulk-footprinting

Run a complete bulk ATAC-seq analysis from aligned reads and peak regions to footprint scores, motif comparisons, and interactive reports.

The bulk workflow guide provides minimal sample and comparison tables for a two-condition analysis.

Example command

bulk-footprinting --sample-table samples.tsv --comparison-table comparisons.tsv --genome hg38 \
  --outdir project

Primary inputs

The workflow uses all available cores by default. It shows each stage and its command output live in the terminal, while keeping the log files listed below. The complete options include an optional core limit.

In the graphical user interface (GUI), leave the Cores field blank to select all available cores when the workflow runs. Enter a number only to limit it. Saved configurations retain automatic selection across machines.

  • --sample-table — TSV with sample, condition, bam, and peaks columns. Each BAM must be coordinate-sorted and have a matching BAI index.
  • --comparison-table — TSV with comparison, cond1, and cond2 columns. Use condition names from the sample table.
  • --genome — managed hg38 or mm10 assembly, or a reference FASTA matching the BAM and peak coordinates.
  • --outdir — project output directory.

Use the same genome assembly and chromosome names for every BAM and BED file.

Main outputs

{project} is the --outdir, {sample} comes from the sample table, and {comparison} comes from the comparison table:

Path Meaning
{project}/samples/{sample}/atac_correct/{sample}_corrected.bw Bias-corrected cut-site signal.
{project}/samples/{sample}/footprints/{sample}_footprints.bw Footprint score signal.
{project}/samples/{sample}/match_motifs/motif_matches_results.txt Per-sample motif summary and binding calls.
{project}/comparisons/{comparison}/diff_footprints_results.txt Motif-level differential statistics.
{project}/comparisons/{comparison}/diff_footprints_{cond1}_{cond2}.html Portable interactive comparison report.
{project}/reports/review_multi_comparisons/index.html Static browser combining every requested comparison.
{project}/reports/review_multi_comparisons.html Aggregate-free portable review written when standalone HTML review mode is selected.
{project}/logs/bulk_footprinting/bulk_footprinting_commands.sh Exact commands generated for the workflow stages.
{project}/logs/bulk_footprinting/{stage}.stdout.log and {stage}.stderr.log Stage-specific logs for troubleshooting.

Start by opening the comparison HTML report. Use the combined review to compare motif results across all requested comparisons, and inspect aggregate profiles alongside the statistics.

Add --dry-run to check the inputs and inspect the commands before starting.

For a combined review, include each condition pair only once. Reversing the conditions still counts as the same pair. Repeated pairs are rejected before reference downloads or analysis, with the comparison IDs and table line numbers shown in the error. To run separate comparisons with the same condition labels but different sample subsets, use --review-format none.

Reference and motif options

Choosing hg38 or mm10 downloads and verifies the matching reference and blacklist as needed. Use --reference-dir to choose where they are cached. --blacklist replaces a managed assembly's blacklist, while --no-blacklist disables it. Custom FASTA inputs never infer a blacklist.

Choose a packaged motif database with --motif-db, provide custom files with --motifs, or combine both options. With neither option, the workflow uses jaspar2026_vertebrates. Run bulk-footprinting --list-motif-dbs to list the packaged databases.

If you have FASTQ files on Linux, run prepare-atac first, then provide its generated metadata/samples.tsv to this command.

Complete options

usage: bulk-footprinting [-h] --sample-table SAMPLE_TABLE --comparison-table
                         COMPARISON_TABLE --genome GENOME
                         [--reference-dir REFERENCE_DIR]
                         [--blacklist BLACKLIST | --no-blacklist]
                         [--motifs [MOTIFS ...]] [--motif-db MOTIF_DB]
                         [--list-motif-dbs]
                         [--normalization {none,condition-quantile,sample-quantile}]
                         [--plot-aggregate {sig,all,top,off}]
                         [--review-format {auto,bundle,standalone,none}]
                         --outdir OUTDIR [--cores CORES] [--resume] [--force]
                         [--dry-run] [--fail-fast]
                         [--runtime {auto,managed,system,container}]

Run the complete bulk ATAC-seq footprinting workflow for explicit comparisons.

options:
  -h, --help            show this help message and exit
  --sample-table SAMPLE_TABLE
                        TSV with sample, condition, BAM, and peak BED columns.
  --comparison-table COMPARISON_TABLE
                        TSV with comparison, cond1, and cond2 columns.
  --genome GENOME       Managed assembly (hg38 or mm10) or a reference FASTA
                        matching the BAM and peak coordinates.
  --reference-dir REFERENCE_DIR
                        Managed reference cache root (default: ~/.cache/fp-
                        tools/references).
  --blacklist BLACKLIST
                        Blacklist BED overriding the managed assembly
                        blacklist.
  --no-blacklist        Disable the automatic hg38/mm10 blacklist.
  --motifs [MOTIFS ...]
                        Optional motif files.
  --motif-db MOTIF_DB   Built-in motif database to use alone or combine with
                        --motifs (default when neither is provided:
                        jaspar2026_vertebrates).
  --list-motif-dbs      List available built-in motif databases and exit.
  --normalization {none,condition-quantile,sample-quantile}
                        Differential-stage normalization (default: none).
  --plot-aggregate {sig,all,top,off}
                        Aggregate profiles generated by diff-footprints
                        (default: all).
  --review-format {auto,bundle,standalone,none}
                        Combined-review output; auto uses standalone when
                        aggregation is off and bundle otherwise (default:
                        auto).
  --outdir OUTDIR       Project output directory.
  --cores CORES         Optional total worker core limit (default: all
                        available cores).
  --resume              Skip stages whose expected outputs are complete.
  --force               Rerun stages even when outputs already exist.
  --dry-run             Validate inputs and print the commands without running
                        them.
  --fail-fast           Stop at the first failed stage.
  --runtime {auto,managed,system,container}
                        External-tool runtime: auto/managed provisions the
                        pinned fp-tools runtime, system uses PATH, and
                        container uses the complete image (default: auto).

atac-correct

Correct ATAC-seq cut-site signal for Tn5 sequence bias. Start with coordinate-sorted BAM files, adjacent BAI indexes, and matching peak BED files. Use the corrected signal as the input to call-footprints.

Example command

atac-correct \
  --sample-table project/metadata/samples.tsv \
  --genome hg38.fa.gz \
  --blacklist hg38.blacklist.bed \
  --outdir project

Primary inputs

  • --sample-table — tab-separated file with one row per sample and the columns sample, condition, bam, and peaks. Give each sample a unique name and use the same condition label for biological replicates.
  • --genome — reference FASTA whose chromosome names and assembly match every BAM and peak BED.
  • --blacklist — optional assembly-matched BED of regions to exclude from bias estimation and corrected output.
  • --outdir — project directory represented by {project} below.

Use the sample-table example to prepare your inputs. Replace the FASTA and blacklist paths with your own files. A custom FASTA does not automatically select a blacklist.

Main outputs

For each {sample} in the table, this command writes:

Path Meaning
{project}/samples/{sample}/atac_correct/{sample}_corrected.bw Bias-corrected cut-site signal. Positive positions have more observed cuts than expected; negative positions have fewer.
{project}/samples/{sample}/atac_correct/{sample}_atacorrect.pdf Diagnostic plots comparing learned Tn5 sequence bias before and after correction. Omitted with --skip-qc.
{project}/samples/{sample}/atac_correct/{sample}_AtacBias.pickle Saved bias model for reuse. Downstream commands use the corrected bigWig and do not need this file.

With --write-tracks all, the same directory also contains:

Path Meaning
{sample}_uncorrected.bw Observed base-resolution cut-site signal after the configured forward/reverse read shifts and sequencing-depth normalization.
{sample}_bias.bw Tn5 sequence-bias score predicted from the reference sequence.
{sample}_expected.bw Expected cut-site signal after the sequence-bias score is scaled to local observed cuts.

Project-level peak outputs are {project}/peaks/merged_peaks.bed and {project}/peaks/merged_peaks_filtered.bed. A direct single-BAM run writes the same {prefix}_*.bw, {prefix}_atacorrect.pdf, and {prefix}_AtacBias.pickle patterns directly under {outdir}.

Open the QC PDF to inspect the sequence-bias diagnostics, then use the corrected bigWig for footprint scoring. A negative corrected value means fewer cuts than the bias model expected; it does not by itself identify a bound TF.

Complete options

usage: atac-correct [-h] [--bams [<bam> ...]] [--fragments [<fragments.tsv.gz> ...]]
                    [-g <fasta>] [-p [<bed> ...]] [--regions-in <bed>]
                    [--regions-out <bed>] [--blacklist <bed>] [--extend <int>]
                    [--split-strands] [--norm-off] [--write-tracks [<track> ...]]
                    [--track-off [<track> ...]] [--skip-qc]
                    [--scale-corrected {auto,none,q95}] [--scale-background <bed>]
                    [--scale-corrected-bigwigs [<bigwig> ...]]
                    [--scale-target {median,mean}] [--scale-chrom-sizes <chrom.sizes>]
                    [--merged-peaks-out <bed>] [--drop-chroms [<chrom> ...]]
                    [--k_flank <int>] [--read_shift <int> <int>] [--bg_shift <int>]
                    [--window <int>] [--score_mat <mat>] [--bias-pkl <obj>]
                    [--prefix <prefix>] [--sample-names [<name> ...]]
                    [--sample-table <tsv>] [--layout {custom,project}]
                    [--sample-output-root <directory>] [--outdir <directory>]
                    [--cores <int>] [--sample-workers <int>] [--split <int>]
                    [--verbosity <int>]

__________________________________________________________________________________________

                                  fp-tools atac-correct
__________________________________________________________________________________________

atac-correct corrects ATAC-seq cutsite signal for Tn5 sequence bias.

Usage:
atac-correct --bams <reads.bam> [<more_reads.bam> ...] --genome <genome.fa> --peaks
<merged_peaks.bed> [<sample_peaks.bed> ...]

Output files:
- <outdir>/<sample>/<sample>_corrected.bw for multi-BAM runs
- optional auxiliary tracks with --write-tracks
- <outdir>/<prefix>_atacorrect.pdf

------------------------------------------------------------------------------------------

Required arguments:
  --bams [<bam> ...]               One or more .bam files containing reads to be corrected
  --fragments [<fragments.tsv.gz> ...]
                                   One or more 10x-style fragment files; uses cut sites
                                   directly without creating pseudo-BAMs
  -g <fasta>, --genome <fasta>     A .fasta-file containing whole genomic sequence
  -p [<bed> ...], --peaks [<bed> ...]
                                   One shared merged peak BED, or multiple per-sample peak
                                   BEDs to merge internally

Optional arguments:
  --regions-in <bed>               Input regions for estimating bias (default: regions not
                                   in peaks.bed)
  --regions-out <bed>              Output regions (default: peaks.bed)
  --blacklist <bed>                Blacklisted regions in .bed-format (default: None)
  --extend <int>                   Extend output regions with basepairs
                                   upstream/downstream (default: 100)
  --split-strands                  Write out tracks per strand
  --norm-off                       Switches off normalization based on number of reads
  --write-tracks [<track> ...]     Cut-site signal bigWigs to write (default: corrected;
                                   use all for corrected, observed/uncorrected, sequence-
                                   bias, and expected signals)
  --track-off [<track> ...]        Compatibility option to switch off individual bigWig
                                   tracks after --write-tracks is resolved
  --skip-qc                        Skip atac-correct diagnostic PDF and pre/post bias
                                   verification counts. Corrected bigWig output is
                                   unchanged.
  --scale-corrected {auto,none,q95}
                                   Optionally q95-scale corrected bigWigs after
                                   correction. In auto mode this runs only when --scale-
                                   corrected-bigwigs has more than one track (default:
                                   auto)
  --scale-background <bed>         Shared BED regions used to estimate q95 scaling for
                                   --scale-corrected
  --scale-corrected-bigwigs [<bigwig> ...]
                                   Corrected bigWigs to q95-scale together. Include the
                                   current sample's corrected bigWig or omit to scale only
                                   the current output
  --scale-target {median,mean}     Across-sample q95 target for --scale-corrected
                                   (default: median)
  --scale-chrom-sizes <chrom.sizes>
                                   Optional chromosome sizes file for scaled bigWig output
                                   validation
  --merged-peaks-out <bed>         Path for internally merged peak BED when multiple
                                   --peaks files are supplied (default:
                                   <outdir>/merged_peaks.bed)
  --drop-chroms [<chrom> ...]      Drop any chromosomes in the list from the correction.
                                   The default is to drop the mitochondrial chromosome.
                                   Default: ['chrM', 'chrMT', 'M', 'MT', 'Mito']

Advanced options (defaults are suitable for most analyses):
  --k_flank <int>                  Flank +/- of cutsite to estimate bias from (default:
                                   12)
  --read_shift <int> <int>         Read shift for forward and reverse reads (default: 4
                                   -5)
  --bg_shift <int>                 Read shift for estimation of background frequencies
                                   (default: 100)
  --window <int>                   Window size for calculating expected signal (default:
                                   100)
  --score_mat <mat>                Type of matrix to use for bias estimation (PWM/DWM)
                                   (default: DWM)
  --bias-pkl <obj>                 Path to a pre-calculated AtacBias.pkl-object, as output
                                   from a previous atac-correct run (default: None). Can
                                   be used to bypass the internal bias estimation.

Run arguments:
  --prefix <prefix>                Prefix for output files in single-BAM runs (default:
                                   BAM filename stem)
  --sample-names [<name> ...]      Sample labels for --bams (default: BAM filename stems)
  --sample-table <tsv>             Project sample table with sample, condition, bam, and
                                   peaks columns
  --layout {custom,project}        Use fp-tools standard project output layout under
                                   --outdir (default: project when --sample-table is
                                   provided)
  --sample-output-root <directory>
                                   Sample output root; writes each sample under
                                   <root>/<sample>/atac_correct, typically
                                   <project>/samples
  --outdir <directory>             Output directory for files (default: current working
                                   directory)
  --cores <int>                    Number of cores to use for computation (default: all
                                   available cores)
  --sample-workers <int>           Number of samples to process concurrently for multi-BAM
                                   runs (default: auto when --cores is set)
  --split <int>                    Split of multiprocessing jobs (default: 100)
  --verbosity <int>                Level of output logging (0: silent, 1: errors/warnings,
                                   2: info, 3: stats, 4: debug, 5: spam) (default: 3)

call-footprints

Calculate a footprint score at each position in accessible regions using the corrected bigWigs from atac-correct. The score measures local cut-site depletion relative to the surrounding signal; higher scores indicate stronger footprint evidence.

Example command

call-footprints \
  --signals A_corrected.bw B_corrected.bw \
  --sample-names A B \
  --regions merged_peaks.bed \
  --sample-output-root project/samples

Primary inputs

  • --signals — one bias-corrected cut-site signal bigWig per sample.
  • --sample-names — labels in the same order as --signals.
  • --regions — BED intervals in which scores are calculated; normally the project merged, filtered peaks.
  • --sample-output-root — root represented by {sample_root} below.

Replace A_corrected.bw and B_corrected.bw with your corrected tracks. In a project, these are under project/samples/{sample}/atac_correct/; use project/peaks/merged_peaks_filtered.bed for the regions.

Main outputs

Path Meaning
{sample_root}/{sample}/footprints/{sample}_footprints.bw Base-resolution footprint score bigWig used by match-motifs and diff-footprints.
{sample_root}/{sample}/footprints/{sample}_candidate_footprints.bed Optional local score maxima for de novo motif discovery; written with --call-candidates.
user-selected *.npz Optional compressed scale-by-position score arrays when --score multiscale is used with an NPZ output option.

In direct mode, --output result.bw writes exactly result.bw; multiple signals written through --outdir {outdir} use {outdir}/{signal_stem}_footprints.bw.

Pass the footprint score tracks to match-motifs to summarize evidence at known motif sites. Keep the corrected cut-site tracks for aggregate plots, which show the signal shape around those sites.

Complete options

usage: call-footprints [-h] [-s <bigwig>] [--signals [<bigwig> ...]] [-o <bigwig>]
                       [--outputs [<bigwig> ...]] [-r <bed>] [--score <score>]
                       [--absolute] [--extend <int>] [--smooth <int>]
                       [--min-limit <float>] [--max-limit <float>] [--scales [<int> ...]]
                       [--multiscale-summary <method>] [--output-multiscale-npz <npz>]
                       [--output-multiscale-npzs [<npz> ...]] [--output-bed <bed>]
                       [--output-beds [<bed> ...]] [--output-bed-dir <directory>]
                       [--call-candidates] [--top-n <int>] [--min-score <float>]
                       [--call-width <bp>] [--min-distance <bp>] [--fp-min <int>]
                       [--fp-max <int>] [--flank-min <int>] [--flank-max <int>]
                       [--footprint-kernel {fast,reference}] [--window <int>]
                       [--sample-names [<name> ...]] [--sample-table <tsv>]
                       [--layout {custom,project}] [--sample-output-root <directory>]
                       [--outdir <directory>] [--cores <int>] [--sample-workers <int>]
                       [--split <int>] [--verbosity <int>]

__________________________________________________________________________________________

                                 fp-tools call-footprints
__________________________________________________________________________________________

call-footprints calculates footprint, sum, mean, or pass-through scores from one or more
bigWig signals and can optionally call ranked footprint candidate intervals.

Usage: call-footprints --signals <cutsites.bw> [<more_cutsites.bw> ...] --regions
<regions.bed> --outdir <output_dir>
   or: call-footprints --signal <cutsites.bw> --regions <regions.bed> --output <output.bw>

Output:
- footprint score bigWig(s)
- optional candidate BED from --output-bed/--output-beds for de novo motif discovery

------------------------------------------------------------------------------------------

Required arguments:
  -s <bigwig>, --signal <bigwig>        A .bw file of ATAC-seq cutsite signal
  --signals [<bigwig> ...]              One or more .bw files of ATAC-seq cutsite signal
  -o <bigwig>, --output <bigwig>        Full path to output bigwig
  --outputs [<bigwig> ...]              Output bigWig path per --signals input
  -r <bed>, --regions <bed>             Genomic regions to run footprinting within

Optional arguments:
  --score <score>                       Type of scoring to perform on cutsites
                                        (footprint/sum/mean/none/multiscale) (default:
                                        footprint)
  --absolute                            Convert bigwig signal to absolute values before
                                        calculating score
  --extend <int>                        Extend input regions with bp (default: 100)
  --smooth <int>                        Smooth output signal by mean in <bp> windows
                                        (default: no smoothing)
  --min-limit <float>                   Limit input bigwig value range (default: no lower
                                        limit)
  --max-limit <float>                   Limit input bigwig value range (default: no upper
                                        limit)

Parameters for score == multiscale:
  --scales [<int> ...]                  Window sizes for multiscale depletion scoring
                                        (default: 8 16 24 32 64 100 147)
  --multiscale-summary <method>         How to collapse scale-specific scores into the
                                        output bigWig (default: max)
  --output-multiscale-npz <npz>         Optional compressed NumPy sidecar with per-region
                                        scale-by-position multiscale scores (only for
                                        --score multiscale)
  --output-multiscale-npzs [<npz> ...]  Output multiscale NPZ sidecar per --signals input

Optional footprint candidate BED calling:
  --output-bed <bed>                    Optional BED-like file of genomic coordinates for
                                        footprint peaks used by de novo motif discovery
  --output-beds [<bed> ...]             Candidate BED path per --signals input
  --output-bed-dir <directory>          Directory for candidate BED files derived from
                                        --signals names
  --call-candidates                     In project/sample-output-root mode, also write
                                        candidate footprint BEDs for de novo motif
                                        discovery
  --top-n <int>                         Keep only the top N footprint calls by score
                                        (default: keep all)
  --min-score <float>                   Minimum footprint score for candidate BED calls
                                        (default: no threshold)
  --call-width <bp>                     Width of candidate BED intervals centered on local
                                        maxima (default: 50)
  --min-distance <bp>                   Minimum distance between retained local footprint
                                        centers within a region (default: 20)

Parameters for score == footprint:
  --fp-min <int>                        Minimum footprint width (default: 20)
  --fp-max <int>                        Maximum footprint width (default: 50)
  --flank-min <int>                     Minimum range of flanking regions (default: 10)
  --flank-max <int>                     Maximum range of flanking regions (default: 30)
  --footprint-kernel {fast,reference}   Footprint scoring kernel (default: fast; use
                                        reference for the original scalar implementation)

Parameters for score == sum:
  --window <int>                        The window for calculation of sum (default: 100)

Run arguments:
  --sample-names [<name> ...]           Sample labels for --signals when using project
                                        layout
  --sample-table <tsv>                  Project sample table with sample, condition, bam,
                                        and peaks columns
  --layout {custom,project}             Use fp-tools standard project output layout under
                                        --outdir (default: project when --sample-table is
                                        provided)
  --sample-output-root <directory>      Sample output root; writes each sample under
                                        <root>/<sample>/footprints, typically
                                        <project>/samples
  --outdir <directory>                  Output directory used with --signals when
                                        --outputs is not supplied
  --cores <int>                         Number of cores to use for computation (default:
                                        all available cores)
  --sample-workers <int>                Number of input signals to process concurrently
                                        for multi-signal runs (default: auto when --cores
                                        is set)
  --split <int>                         Split of multiprocessing jobs (default: 100)
  --verbosity <int>                     Level of output logging (0: silent, 1:
                                        errors/warnings, 2: info, 3: stats, 4: debug, 5:
                                        spam) (default: 3)

match-motifs

Find motif matches in accessible regions and measure the footprint score at each match. This produces a motif summary for each sample and classifies sites as predicted bound or unbound.

Example command

match-motifs \
  --signals A_footprints.bw B_footprints.bw \
  --sample-names A B \
  --genome hg38.fa.gz \
  --peaks merged_peaks.bed \
  --motif-db jaspar2026_vertebrates \
  --sample-output-root project/samples

Primary inputs

  • --signals — one footprint score bigWig per sample.
  • --sample-names — sample labels in the same order as --signals.
  • --genome — assembly-matched reference FASTA used to scan motif sequences.
  • --peaks — accessible-region BED searched for motif instances.
  • --motif-db — packaged motif collection; the example uses JASPAR 2026 vertebrates.
  • --sample-output-root — root represented by {sample_root} below.

Use the score tracks from call-footprints. In a project, they are under project/samples/{sample}/footprints/; use the same reference FASTA and filtered peaks as the earlier steps.

Main outputs

For each {sample}, the example writes to {sample_root}/{sample}/match_motifs/:

Path Meaning
motif_matches_results.txt Tab-separated motif summary with site counts and per-sample mean scores.
motif_matches_distances.txt Motif-similarity distances used for motif clustering.
cache/motif_sites.tsv.gz Compact scanned motif-site cache reusable by differential analysis.
cache/background_scores.tsv.gz Compact background-score cache.
{motif}/beds/{motif}_{sample}_all.bed All scanned instances for one motif.
{motif}/beds/{motif}_{sample}_bound.bed Instances classified as bound in the sample.
{motif}/beds/{motif}_{sample}_unbound.bed Instances classified as unbound in the sample.

{motif} follows the selected --naming convention, such as CTCF_MA0139.2. --motif-outputs summary omits the per-motif BED files but keeps the summary and reusable caches.

The default shared scan in the example writes tab-separated summaries. An independent sample scan can also write motif_matches_results.xlsx unless --skip-excel is used. A joint analysis with repeated condition labels can write motif_matches_replicate_motif_score_matrix.tsv; separate per-sample folders do not contain that joint matrix.

Start with motif_matches_results.txt to review site counts and mean scores. Bound and unbound are model-based classifications, not direct measurements of TF binding. Related TFs can recognize similar motifs, so a motif name alone does not distinguish every TF in a family.

To use your own motifs, replace --motif-db jaspar2026_vertebrates with --motifs your_motifs.meme. Custom motifs alone do not add the default database; JASPAR 2026 vertebrates is used when neither option is supplied.

Complete options

usage: match-motifs [-h] [--signals [<bigwig> ...]] [--peaks <bed>] [--genome <fasta>]
                    [--motifs [<motifs> ...]] [--motif-db <name>] [--list-motif-dbs]
                    [--sample-names [<name> ...]] [--cond-names [<name> ...]]
                    [--sample-table <tsv>] [--layout {custom,project}]
                    [--match-scan-mode {auto,shared,per-sample}]
                    [--sample-output-root <directory>] [--peak-header <file>]
                    [--naming <string>] [--motif-pvalue <float>] [--bound-pvalue <float>]
                    [--cluster-threshold <float>] [--pseudo <float>] [--skip-excel]
                    [--output-peaks <bed>] [--norm-off]
                    [--normalization {condition-quantile,sample-quantile,none}]
                    [--aggregate-signals [<bigwig> ...]]
                    [--plot-aggregate {sig,all,top,off}] [--plot-aggregate-top-n <int>]
                    [--plot-aggregate-motifs <motif> [<motif> ...]]
                    [--default-aggregate-plots <int>]
                    [--aggregate-pvalue-threshold <float>] [--aggregate-flank <bp>]
                    [--aggregate-normalization {match,none,sample-quantile,size-factor}]
                    [--aggregate-site-set {all,bound}]
                    [--motif-outputs {auto,summary,full}] [--report-label <text>]
                    [--outdir <directory>] [--prefix <prefix>] [--cores <int>]
                    [--sample-workers <int>] [--split <int>] [--debug] [--verbosity <int>]

__________________________________________________________________________________________

                                  fp-tools match-motifs
__________________________________________________________________________________________

match-motifs scans motifs in open chromatin regions for one or more footprint score tracks
and infers sample-specific bound and unbound motif sites.

Usage:
match-motifs --signals <footprints.bw> [<more_footprints.bw> ...] --genome <genome.fasta>
--peaks <peaks.bed> [--motif-db jaspar2026_vertebrates | --motifs <motifs.txt>]

Output files:
- <outdir>/<prefix>_results.{txt,xlsx}
- <outdir>/<prefix>_distances.txt
- <outdir>/cache/* compact reuse caches
- <outdir>/<TF>/beds/*_all.bed, *_bound.bed, and *_unbound.bed by default
- optional <outdir>/<TF>/<TF>_overview.{txt,xlsx} with --motif-outputs full

------------------------------------------------------------------------------------------

Required arguments:
  --signals [<bigwig> ...]         One or more footprint score bigWigs (.bigwig format)
  --peaks <bed>                    Peaks.bed containing open chromatin regions
  --genome <fasta>                 Genome .fasta file

Optional arguments:
  --motifs [<motifs> ...]          Motif file(s) in pfm/jaspar/meme/transfac format; if
                                   omitted, the built-in JASPAR 2026 vertebrates set is
                                   used
  --motif-db <name>                Built-in motif database to use or add to --motifs
                                   (default when --motifs is omitted:
                                   jaspar2026_vertebrates)
  --list-motif-dbs                 List available built-in motif databases and exit
  --sample-names [<name> ...]      Sample labels for --signals (default: prefix of each
                                   --signals file)
  --cond-names [<name> ...]        Optional condition labels for --signals (default:
                                   prefix of each --signals file)
  --sample-table <tsv>             Project sample table with sample and condition columns
                                   plus bam/peaks for upstream steps
  --layout {custom,project}        Use fp-tools standard project output layout under
                                   --outdir (default: project when --sample-table is
                                   provided)
  --match-scan-mode {auto,shared,per-sample}
                                   Project match-motifs scan mode. auto uses one shared
                                   motif scan for multi-sample project runs; per-sample
                                   preserves independent sample scans.
  --sample-output-root <directory>
                                   Sample output root; writes one match_motifs folder
                                   under <root>/<sample> for each input signal, typically
                                   <project>/samples
  --peak-header <file>             File containing the header of --peaks separated by
                                   whitespace or newlines (default: peak columns are named
                                   "_additional_<count>")
  --naming <string>                Naming convention for TF output files ('id', 'name',
                                   'name_id', 'id_name') (default: 'name_id')
  --motif-pvalue <float>           Set p-value threshold for motif scanning (default:
                                   1e-4)
  --bound-pvalue <float>           Set p-value threshold for bound/unbound split (default:
                                   0.001)
  --cluster-threshold <float>      Set the clustering threshold. Motifs below this
                                   threshold will be assigned to one cluster (default:
                                   0.5)
  --pseudo <float>                 Pseudocount for calculating log2fcs (default: estimated
                                   from data)
  --skip-excel                     Skip creation of Excel files to speed up large motif
                                   analyses
  --output-peaks <bed>             Gives the possibility to set the output peak set
                                   differently than the input --peaks. This will limit all
                                   analysis to the regions in --output-peaks. NOTE:
                                   --peaks must still be set to the full peak set!
  --norm-off                       Turn off normalization of footprint scores
  --normalization {condition-quantile,sample-quantile,none}
                                   Signal normalization mode (default: none; --norm-off
                                   maps to none)
  --aggregate-signals [<bigwig> ...]
                                   Corrected cut-site bigWigs used for embedded aggregate
                                   profiles
  --plot-aggregate {sig,all,top,off}
                                   Embed aggregate profiles in HTML reports for
                                   significant, all, top-N, or no motifs (default: sig)
  --plot-aggregate-top-n <int>     Maximum number of motifs to aggregate when --plot-
                                   aggregate sig/top or fallback selection is used
                                   (default: 20)
  --plot-aggregate-motifs <motif> [<motif> ...]
                                   Ordered motif IDs, names, or output prefixes to embed
                                   as aggregate profiles; overrides automatic aggregate-
                                   motif selection
  --default-aggregate-plots <int>  Number of aggregate profiles initially displayed in
                                   interactive reports (default: 4; maximum: 12)
  --aggregate-pvalue-threshold <float>
                                   P-value threshold for --plot-aggregate sig (default:
                                   0.05)
  --aggregate-flank <bp>           Flank around motif centers for embedded aggregate
                                   profiles (default: 100)
  --aggregate-normalization {match,none,sample-quantile,size-factor}
                                   Normalization for embedded aggregate profiles (default:
                                   match --normalization)
  --aggregate-site-set {all,bound}
                                   Motif-site BEDs used for embedded aggregate profiles:
                                   all motif hits or sample-specific bound sites (default:
                                   all)
  --motif-outputs {auto,summary,full}
                                   Per-motif output mode. For match-motifs, auto writes
                                   compact caches plus per-motif BED files; summary writes
                                   only main result/report tables and caches; full writes
                                   per-motif BED and overview files synchronously. For
                                   diff-footprints, auto writes full motif outputs only
                                   when aggregate reports need them.
  --report-label <text>            Optional method label shown under the report subtitle
                                   in interactive HTML reports
  --prefix <prefix>                Prefix for overview files in --outdir folder (default:
                                   motif_matches)

Run arguments:
  --outdir <directory>             Output directory to place motif tables, BED files, and
                                   plots in (default: motif_matches_output)
  --cores <int>                    Number of cores to use for computation (default: all
                                   available cores)
  --sample-workers <int>           Number of input samples to process concurrently in
                                   project/sample-output-root mode (default: auto when
                                   --cores is set)
  --split <int>                    Split of multiprocessing jobs (default: 100)
  --debug                          Creates an additional '_debug.pdf'-file with debug
                                   plots
  --verbosity <int>                Level of output logging (0: silent, 1: errors/warnings,
                                   2: info, 3: stats, 4: debug, 5: spam) (default: 3)

diff-footprints

Find motifs whose footprint scores differ between conditions. Start with the per-sample results from match-motifs. You can also compare two sets of genomic regions measured in the same samples.

Example command

diff-footprints \
  --sample-table project/metadata/samples.tsv \
  --comparison-table project/metadata/comparisons.tsv \
  --genome hg38.fa.gz \
  --peaks project/peaks/merged_peaks_filtered.bed \
  --motif-db jaspar2026_vertebrates \
  --outdir project

Primary inputs

  • --sample-table — TSV with sample and condition columns. Reuse the table from the earlier steps; the command finds results under {project}/samples/{sample}/.
  • --comparison-table — TSV with comparison, cond1, and cond2 columns. Each row names one comparison and its two condition labels.
  • --genome — the reference FASTA used for motif matching.
  • --peaks — accessible-region BED file.
  • --motif-db — built-in motif database name.
  • --outdir — project directory for statistics, figures, and HTML reports.

Condition labels must match the sample table. Include each biological replicate as its own sample row. See the sample and comparison tables for a minimal example.

Main outputs

Each comparison is written below {project}/comparisons/{comparison}/, where {comparison} is taken from the comparison table and {prefix} defaults to diff_footprints:

Path Meaning
{prefix}_results.txt Tab-separated motif-level differential footprint statistics; change direction is cond1 - cond2.
{prefix}_results.xlsx Excel copy in direct runs unless --skip-excel is used. Project comparison-table runs, including the example above, omit it.
{prefix}_distances.txt Motif distances used for clustering related motifs.
{prefix}_{cond1}_{cond2}.html Portable interactive report with volcano, motif, and embedded aggregate-profile views.
{prefix}_replicate_report.tsv Long-form per-replicate diagnostic data when replicate reporting is active.
{prefix}_replicate_summary.tsv Motif-level replicate agreement summary.
{prefix}_replicate_report.png Replicate diagnostic figure.
{prefix}_figures.pdf and {prefix}_clusters.pdf Optional static summaries written with --static-plots.
{motif}/beds/{motif}_{condition}_bound.bed Motif instances classified as bound for a condition when full motif outputs are required.

Open the HTML report to explore motifs, then use the text result table for further analysis. Positive changes favor cond1; negative changes favor cond2. Check replicate agreement and the aggregate cut-site profiles along with the statistics when interpreting a difference.

Compare region sets

Region-set analyses use the same result/report patterns and add confidence intervals, motif prevalence, region counts, per-replicate effects, and matching balance to the result tables.

For a region-set comparison, use --comparison-axis regions, provide two or more BED files with --regions, and name them with --region-labels. An optional --region-strata-column preserves accessibility or other matching strata during resampling. One sample uses a stratified label-permutation test; two or more biological replicates use a paired empirical-Bayes model. Use --plot-aggregate-motifs to choose an ordered aggregate panel without limiting the motifs tested, and --default-aggregate-plots to set its initial size.

Complete options

usage: diff-footprints [-h] [--signals [<bigwig> ...]] [--peaks <bed>] [--genome <fasta>]
                       [--motifs [<motifs> ...]] [--motif-db <name>] [--list-motif-dbs]
                       [--sample-names [<name> ...]] [--cond-names [<name> ...]]
                       [--comparison-axis {conditions,regions}] [--regions [<bed> ...]]
                       [--region-labels [<name> ...]] [--region-strata-column <int>]
                       [--region-permutations <int>] [--region-bootstrap <int>]
                       [--min-regions-per-set <int>] [--random-seed <int>]
                       [--sample-table <tsv>] [--comparison-table <tsv>]
                       [--layout {custom,project}] [--sample-dirs [<directory> ...]]
                       [--project-dir <directory>] [--peak-header <file>]
                       [--naming <string>] [--motif-pvalue <float>]
                       [--bound-pvalue <float>] [--cluster-threshold <float>]
                       [--pseudo <float>] [--time-series] [--time-course] [--skip-excel]
                       [--output-peaks <bed>] [--norm-off]
                       [--normalization {condition-quantile,sample-quantile,none}]
                       [--replicate-report {auto,on,off}] [--replicate-map <tsv>]
                       [--replicate-report-out <tsv>] [--replicate-summary-out <tsv>]
                       [--replicate-figure-out <figure>]
                       [--aggregate-signals [<bigwig> ...]]
                       [--plot-aggregate {sig,all,top,off}] [--plot-aggregate-top-n <int>]
                       [--plot-aggregate-motifs <motif> [<motif> ...]]
                       [--default-aggregate-plots <int>]
                       [--aggregate-pvalue-threshold <float>] [--aggregate-flank <bp>]
                       [--aggregate-normalization {match,none,sample-quantile,size-factor}]
                       [--aggregate-site-set {all,bound}] [--reuse-existing-results]
                       [--motif-outputs {auto,summary,full}] [--static-plots]
                       [--per-motif-plots] [--skew-report] [--report-label <text>]
                       [--outdir <directory>] [--prefix <prefix>] [--cores <int>]
                       [--split <int>] [--debug] [--verbosity <int>]

__________________________________________________________________________________________

                                 fp-tools diff-footprints
__________________________________________________________________________________________

diff-footprints takes motifs, footprint signals, and genome sequence as input to compare
motif-associated footprint evidence across biological conditions or user-defined region
sets. Region-set mode gives every region equal weight and supports matching strata and
paired biological replicates.

Usage:
diff-footprints --signals <bigwig1> (<bigwig2> (...)) --genome <genome.fasta> --peaks
<peaks.bed> [--motif-db jaspar2026_vertebrates | --motifs <motifs.txt>]
diff-footprints --comparison-axis regions --signals <bigwig1> [<replicate2> ...] --regions
<set1.bed> <set2.bed> --genome <genome.fasta> [--region-labels <set1> <set2>]

Output files:
- <outdir>/<prefix>_results.{txt,xlsx}
- <outdir>/<prefix>_distances.txt
- <outdir>/<prefix>_<condition1>_<condition2>.html
- optional <outdir>/<prefix>_figures.pdf with --static-plots
- optional <outdir>/<prefix>_clusters.pdf with --static-plots
- optional <outdir>/<TF>/plots/<TF>_log2fcs.pdf with --per-motif-plots
- optional <outdir>/<TF>/<TF>_overview.{txt,xlsx} with --motif-outputs full
- <outdir>/<TF>/beds/<TF>_<condition>_bound.bed (per motif-condition pair)
- <outdir>/<TF>/beds/<TF>_<condition>_unbound.bed (per motif-condition pair)

------------------------------------------------------------------------------------------

Required arguments:
  --signals [<bigwig> ...]         Signal per condition (.bigwig format)
  --peaks <bed>                    Peaks.bed containing open chromatin regions across all
                                   conditions (not used with --comparison-axis regions)
  --genome <fasta>                 Genome .fasta file

Optional arguments:
  --motifs [<motifs> ...]          Motif file(s) in pfm/jaspar/meme/transfac format; if
                                   omitted, the built-in JASPAR 2026 vertebrates set is
                                   used
  --motif-db <name>                Built-in motif database to use or add to --motifs
                                   (default when --motifs is omitted:
                                   jaspar2026_vertebrates)
  --list-motif-dbs                 List available built-in motif databases and exit
  --sample-names [<name> ...]      Sample labels for --signals; distinct from --cond-names
                                   and used for per-sample score columns, replicate
                                   reports, and aggregate profiles (default: prefix of
                                   each --signals file)
  --cond-names [<name> ...]        Condition labels for --signals; repeat names to define
                                   biological replicates (default: prefix of each
                                   --signals file)
  --comparison-axis {conditions,regions}
                                   Compare biological conditions or region sets measured
                                   in the same sample(s) (default: conditions)
  --regions [<bed> ...]            Two or more non-overlapping BED files for --comparison-
                                   axis regions
  --region-labels [<name> ...]     Labels for --regions (default: BED filename stems)
  --region-strata-column <int>     Optional 1-based BED column containing matching strata
  --region-permutations <int>      Within-stratum label permutations for a one-sample
                                   region comparison (default: 100000)
  --region-bootstrap <int>         Within-set, within-stratum bootstrap samples for
                                   confidence intervals (default: 1000)
  --min-regions-per-set <int>      Minimum motif-containing regions required in each set
                                   (default: 10)
  --random-seed <int>              Random seed for region-set resampling (default: 1)
  --sample-table <tsv>             Project sample table with sample and condition columns
                                   plus bam/peaks for upstream steps
  --comparison-table <tsv>         Project comparison table with comparison, cond1, and
                                   cond2 columns
  --layout {custom,project}        Use fp-tools standard project output layout under
                                   --outdir (default: project when --sample-table is
                                   provided)
  --sample-dirs [<directory> ...]  Sample output folders containing match_motifs/ outputs
                                   to reuse for differential analysis
  --project-dir <directory>        Parent folder containing sample output folders for
                                   folder-based differential analysis, typically
                                   <project>/samples
  --peak-header <file>             File containing the header of --peaks separated by
                                   whitespace or newlines (default: peak columns are named
                                   "_additional_<count>")
  --naming <string>                Naming convention for TF output files ('id', 'name',
                                   'name_id', 'id_name') (default: 'name_id')
  --motif-pvalue <float>           Set p-value threshold for motif scanning (default:
                                   1e-4)
  --bound-pvalue <float>           Set p-value threshold for bound/unbound split (default:
                                   0.001)
  --cluster-threshold <float>      Set the clustering threshold. Motifs below this
                                   threshold will be assigned to one cluster (default:
                                   0.5)
  --pseudo <float>                 Pseudocount for calculating log2fcs (default: estimated
                                   from data)
  --time-series                    Will only compare signals1<->signals2<->signals3 (...)
                                   in order of input, and skip all-against-all comparison.
  --time-course                    Alias for --time-series; compare adjacent ordered
                                   conditions only.
  --skip-excel                     Skip creation of Excel files to speed up large motif
                                   analyses
  --output-peaks <bed>             Gives the possibility to set the output peak set
                                   differently than the input --peaks. This will limit all
                                   analysis to the regions in --output-peaks. NOTE:
                                   --peaks must still be set to the full peak set!
  --norm-off                       Turn off normalization of footprint scores across
                                   conditions
  --normalization {condition-quantile,sample-quantile,none}
                                   Cross-sample normalization mode (default: none; --norm-
                                   off maps to none)
  --replicate-report {auto,on,off}
                                   Write replicate-aware differential-footprint
                                   diagnostics (default: auto for repeated condition names
                                   or --replicate-map)
  --replicate-map <tsv>            Optional TSV with condition/replicate or
                                   condition/n_replicates columns
  --replicate-report-out <tsv>     Output long-form replicate diagnostic TSV (default:
                                   <outdir>/<prefix>_replicate_report.tsv)
  --replicate-summary-out <tsv>    Output replicate diagnostic summary TSV (default:
                                   <outdir>/<prefix>_replicate_summary.tsv)
  --replicate-figure-out <figure>  Output replicate diagnostic figure (default:
                                   <outdir>/<prefix>_replicate_report.png)
  --aggregate-signals [<bigwig> ...]
                                   Corrected cut-site bigWigs used for embedded aggregate
                                   profiles
  --plot-aggregate {sig,all,top,off}
                                   Embed aggregate profiles in HTML reports for
                                   significant, all, top-N, or no motifs (default: sig)
  --plot-aggregate-top-n <int>     Maximum number of motifs to aggregate when --plot-
                                   aggregate sig/top or fallback selection is used
                                   (default: 20)
  --plot-aggregate-motifs <motif> [<motif> ...]
                                   Ordered motif IDs, names, or output prefixes to embed
                                   as aggregate profiles; overrides automatic aggregate-
                                   motif selection
  --default-aggregate-plots <int>  Number of aggregate profiles initially displayed in
                                   interactive reports (default: 4; maximum: 12)
  --aggregate-pvalue-threshold <float>
                                   P-value threshold for --plot-aggregate sig (default:
                                   0.05)
  --aggregate-flank <bp>           Flank around motif centers for embedded aggregate
                                   profiles (default: 100)
  --aggregate-normalization {match,none,sample-quantile,size-factor}
                                   Normalization for embedded aggregate profiles (default:
                                   match --normalization)
  --aggregate-site-set {all,bound}
                                   Motif-site BEDs used for embedded aggregate profiles:
                                   all motif hits or condition-specific bound sites
                                   (default: all)
  --reuse-existing-results         Regenerate final diff-footprints reports from existing
                                   <prefix>_results.txt and per-motif BEDs without
                                   rescanning motifs
  --motif-outputs {auto,summary,full}
                                   Per-motif output mode. For match-motifs, auto writes
                                   compact caches plus per-motif BED files; summary writes
                                   only main result/report tables and caches; full writes
                                   per-motif BED and overview files synchronously. For
                                   diff-footprints, auto writes full motif outputs only
                                   when aggregate reports need them.
  --static-plots                   Also write static volcano and cluster PDF summaries. By
                                   default diff-footprints writes the interactive HTML
                                   report without these PDFs.
  --per-motif-plots                Also write one diagnostic log2 fold-change PDF per
                                   motif. Disabled by default to keep diff-footprints
                                   fast.
  --skew-report                    Also write the optional skew/shift PDF report. Disabled
                                   by default.
  --report-label <text>            Optional method label shown under the report subtitle
                                   in interactive HTML reports
  --prefix <prefix>                Prefix for overview files in --outdir folder (default:
                                   diff_footprints)

Run arguments:
  --outdir <directory>             Output directory to place motif tables, BED files, and
                                   plots in (default: diff_footprints_output)
  --cores <int>                    Number of cores to use for computation (default: all
                                   available cores)
  --split <int>                    Split of multiprocessing jobs (default: 100)
  --debug                          Creates an additional '_debug.pdf'-file with debug
                                   plots
  --verbosity <int>                Level of output logging (0: silent, 1: errors/warnings,
                                   2: info, 3: stats, 4: debug, 5: spam) (default: 3)

normalize-bigwig

Put corrected cut-site tracks on a comparable signal scale using the same background regions for every sample. This is an optional step after atac-correct; use the resulting tracks for downstream scoring or plotting when you need this normalization.

Example command

normalize-bigwig \
  --sample-table project/metadata/samples.tsv \
  --background project/peaks/merged_peaks_filtered.bed \
  --outdir project \
  --method background-scale \
  --stat q95 \
  --target median

Primary inputs

  • --sample-table — TSV with sample and condition columns; reuse the table from atac-correct. Tracks are read from {project}/samples/{sample}/atac_correct/.
  • --background — shared BED intervals used to calculate comparable background statistics.
  • --outdir — project directory represented by {project} below.
  • --method — transformation; background-scale multiplies each signal by a shared-target scale factor.
  • --stat — within-sample background statistic; the example uses the 95th percentile.
  • --target — across-sample target for the selected statistic; the example uses the median.

Main outputs

Path Meaning
{project}/samples/{sample}/normalize/{sample}_corrected_q95_scaled.bw Q95-scaled bias-corrected cut-site signal bigWig for one sample.
{project}/normalize_bigwig_qc.tsv Background statistics, selected statistic, target, and scale factor for every sample.
{project}/normalize_bigwig_manifest.tsv Sample-to-input/output signal mapping for downstream use.

Check the QC table to see how much each track was scaled. The manifest lists the output paths to pass to the next command. Scaling changes signal magnitude; it does not by itself provide evidence of TF binding.

In custom layout, default outputs use {outdir}/{input_stem}.background_scale_{stat}.bw, plus the two QC tables in {outdir}. background-zscore instead writes standardized signal and uses a method-specific filename suffix.

Complete options

usage: normalize-bigwig [-h] [--bigwigs BIGWIGS [BIGWIGS ...]]
                        [--background BACKGROUND] [--outdir OUTDIR]
                        [--sample-names [SAMPLE_NAMES ...]]
                        [--sample-table SAMPLE_TABLE]
                        [--layout {custom,project}]
                        [--sample-output-root SAMPLE_OUTPUT_ROOT]
                        [--method {background-scale,background-zscore,none}]
                        [--stat STAT] [--target {median,mean}]
                        [--chrom-sizes CHROM_SIZES] [--workers WORKERS]

Normalize input signal bigWigs using robust statistics from shared background
BED regions. For corrected cut-site bigWigs, the recommended method is
background-scale.

options:
  -h, --help            show this help message and exit
  --bigwigs BIGWIGS [BIGWIGS ...]
                        Input bigWig files to normalize together.
  --background BACKGROUND
                        Shared background BED used to estimate sample
                        statistics.
  --outdir OUTDIR       Output directory for normalized bigWig QC tables and
                        default outputs.
  --sample-names [SAMPLE_NAMES ...]
                        Sample labels for --bigwigs when using project layout.
  --sample-table SAMPLE_TABLE
                        Project sample table with sample, condition, bam, and
                        peaks columns.
  --layout {custom,project}
                        Use fp-tools standard project output layout under
                        --outdir (default: project when --sample-table is
                        provided).
  --sample-output-root SAMPLE_OUTPUT_ROOT
                        Sample output root; writes each sample under
                        <root>/<sample>/normalize, typically
                        <project>/samples.
  --method {background-scale,background-zscore,none}
                        Normalization method (default: background-scale).
  --stat STAT           Background statistic used by background-scale
                        (default: q90). Use median, iqr, or quantiles such as
                        q90, q95, q97.5, or q99.
  --target {median,mean}
                        Across-sample target statistic for background-scale
                        (default: median).
  --chrom-sizes CHROM_SIZES
                        Optional chromosome sizes file for output
                        validation/header.
  --workers WORKERS     Number of input signal bigWigs to normalize
                        concurrently (default: all available cores, capped by
                        input count).

plot-aggregate

Plot the average cut-site signal around motif sites to inspect the shape of a footprint. Start with corrected bigWigs from atac-correct and motif results from match-motifs. Outputs can be static figures or an interactive HTML report.

Example command

plot-aggregate \
  --sample-table project/metadata/samples.tsv \
  --motifs SPIB CEBPB \
  --site-set bound \
  --outdir project

Primary inputs

  • --sample-table — TSV with sample and condition columns; the command reads corrected tracks and motif results from {project}/samples/{sample}/.
  • --motifs — motif names or identifiers to plot.
  • --site-set — sites to average; the example uses sites classified as bound in each sample.
  • --outdir — project directory containing motif results and receiving plots.

Replace SPIB CEBPB with motifs present in your motif results.

Main outputs

  • {project}/reports/plot_aggregate.html — default project-layout interactive aggregate report with motif-centered signal profiles.
  • the exact --output path — static PDF/PNG/SVG or interactive HTML in custom layout.
  • the exact --output-txt path — optional per-position aggregate values.
  • the exact --output_aggregated_signals, --output_aggregated_scores, and --output_aggregated_stats paths — optional source tables when requested.
  • the exact --output path in --motif-grid mode — multipage motif-by-comparison PDF built from a review bundle.

When both signal types are available, use footprint score bigWigs for motif statistics and bias-corrected cut-site signal bigWigs for observed aggregate profiles; label the chosen signal explicitly in figure captions.

Look for central cut-site depletion relative to the flanking signal. Also compare the number of sites and the sample-to-sample consistency. Bound-site plots can use different sites in each sample; choose --site-set all to inspect all matched motif sites instead.

For your own site sets, supply BED files with --TFBS. Use --regions to restrict or compare regions of interest.

Export a grid from a review report

plot-aggregate \
  --input-html project/reports/review_multi_comparisons/index.html \
  --motif-grid \
  --output project/reports/motif_aggregate_grid.pdf

Complete options

usage: plot-aggregate [-h] [--TFBS [<bed> ...]] [--signals [<bigwig> ...]]
                      [--match-dir [<directory> ...]] [--sample-dirs [<directory> ...]]
                      [--sample-table <tsv>] [--layout {custom,project}]
                      [--manifest <tsv>] [--input-html [<html> ...]]
                      [--regions [<bed> ...]] [--whitelist [<bed> ...]]
                      [--blacklist [<bed> ...]] [--output] [--outdir <directory>]
                      [--output-txt] [--output-csv] [--output_aggregated_signals]
                      [--output_aggregated_scores] [--multiscale-npz <npz>]
                      [--output-multiscale-aggregate] [--title] [--format {auto,pdf,html}]
                      [--flank] [--motifs [<motif> ...]] [--site-set {bound,all,unbound}]
                      [--top-n <int>] [--default-layout {1x1,1x2,2x2,2x3}]
                      [--hide-summary] [--TFBS-labels [...]] [--signal-labels [...]]
                      [--cond-names [<name> ...]] [--region-labels [...]]
                      [--control-label <label>] [--grid <rows>x<cols>] [--share-y]
                      [--normalize]
                      [--normalization {none,condition-quantile,sample-quantile}]
                      [--normalization-comparison-output] [--output_aggregated_stats]
                      [--show-replicate-sd] [--negate] [--smooth <int>] [--log-transform]
                      [--plot-boundaries] [--signal-on-x] [--remove-outliers <float>]
                      [--motif-grid] [--rows-per-page ROWS_PER_PAGE]
                      [--order-htmls [ORDER_HTMLS ...]] [--fill-missing-profiles]
                      [--recompute-missing-profiles] [--repeat-column-labels {none,row}]
                      [--cores <int>] [--verbosity <int>]

__________________________________________________________________________________________

                                 fp-tools plot-aggregate
__________________________________________________________________________________________

Input / output arguments:
  --TFBS [<bed> ...]                    TFBS sites (*required)
  --signals [<bigwig> ...]              Signals in bigwig format (*required)
  --match-dir [<directory> ...]         match-motifs output directory or directories to
                                        use as the motif-site source
  --sample-dirs [<directory> ...]       Alias for --match-dir in HTML mode; sample or
                                        differential output directories containing motif
                                        BEDs
  --sample-table <tsv>                  Project sample table with sample and condition
                                        columns
  --layout {custom,project}             Use fp-tools standard project output layout under
                                        --outdir (default: project when --sample-table is
                                        provided)
  --manifest <tsv>                      TSV with sample, signal, and match_dir/sample_dir
                                        columns for HTML mode
  --input-html [<html> ...]             Existing aggregate or diff-footprints HTML
                                        payload(s) to merge in HTML mode
  --regions [<bed> ...]                 Regions to overlap with TFBS (optional)
  --whitelist [<bed> ...]               Only plot sites overlapping whitelist (optional)
  --blacklist [<bed> ...]               Exclude sites overlapping blacklist (optional)
  --output                              Path to output plot (default: fp-
                                        tools_aggregate.pdf)
  --outdir <directory>                  Project directory used with --layout project
  --output-txt                          Path to output file for aggregates in .txt-format
                                        (default: None)
  --output-csv                          Path to aggregated signal CSV output (default:
                                        None)
  --output_aggregated_signals           Path to CSV file for per-base aggregated signals
                                        (default: None)
  --output_aggregated_scores            Path to CSV file for aggregated footprint-score
                                        table (default: None)
  --multiscale-npz <npz>                Optional call-footprints --output-multiscale-npz
                                        sidecar to render as a scale-by-position aggregate
                                        figure
  --output-multiscale-aggregate         Path for the optional multiscale aggregate figure
                                        (default: <output stem>_multiscale.<output ext>)

Plot arguments:
  --title                               Title of plot (default: "Aggregated signals")
  --format {auto,pdf,html}              Output format for --output. auto uses the output
                                        file extension (default: auto)
  --flank                               Flanking base pairs (+/-) to show in plot (counted
                                        from middle of the TFBS) (default: 60)
  --motifs [<motif> ...]                Motif prefixes, names, or IDs to plot from
                                        --match-dir
  --site-set {bound,all,unbound}        Motif-site BED set to use from --match-dir
                                        (default: bound)
  --top-n <int>                         Number of motifs to plot from --match-dir when
                                        --motifs is omitted (default: 12)
  --default-layout {1x1,1x2,2x2,2x3}    Initial HTML subplot layout (default: 2x2)
  --hide-summary                        Hide the TF site-count summary sidebar in HTML
                                        mode
  --TFBS-labels [ ...]                  Labels used for each TFBS file (default: prefix of
                                        each --TFBS)
  --signal-labels [ ...]                Labels used for each signal file (default: prefix
                                        of each --signals)
  --cond-names [<name> ...]             Condition names for --signals; repeated names are
                                        averaged as replicates
  --region-labels [ ...]                Labels used for each regions file (default: prefix
                                        of each --regions)
  --control-label <label>               Overlay each non-control signal against this
                                        control signal label (must match one of --signal-
                                        labels)
  --grid <rows>x<cols>                  Explicit grid layout for subplots, e.g. 2x5 or
                                        3x4. Panels fill in order of the input signal
                                        files.
  --share-y                             Share y-axis range across plots
                                        (none/signals/sites/both). Use "--share-y signals"
                                        if bigwig signals have similar ranges. Use "--
                                        share-y sites" if sites per bigwig are comparable,
                                        but bigwigs themselves aren't comparable (default:
                                        none)
  --normalize                           Normalize the aggregate signal(s) to be between
                                        0-1 (default: the true range of values is shown)
  --normalization {none,condition-quantile,sample-quantile}
                                        diff-footprints-compatible quantile normalization
                                        before aggregate plotting (default: none)
  --normalization-comparison-output     Optional paired raw-vs-normalized aggregate figure
  --output_aggregated_stats             Path to CSV file for aggregate mean/SD/stat
                                        summaries (default: None)
  --show-replicate-sd                   Draw replicate SD ribbons when --cond-names
                                        contains repeated condition names
  --negate                              Negate overlap with regions
  --smooth <int>                        Smooth output signal by taking the mean of
                                        <smooth> bp windows (default: 1; no smoothing)
  --log-transform                       Log transform the signals before aggregation
  --plot-boundaries                     Plot TFBS boundaries (Note: estimated from first
                                        region in each --TFBS)
  --signal-on-x                         Show signals on x-axis and TFBSs on y-axis
                                        (default: signal is on y-axis)
  --remove-outliers <float>             Value between 0-1 indicating the percentile of
                                        regions to include, e.g. 0.99 to remove the sites
                                        with 1% highest values (default: 1)
  --motif-grid                          Create a multi-page motif-by-comparison PDF from
                                        one review-multi-comparisons report
  --rows-per-page ROWS_PER_PAGE         Motif rows per page in --motif-grid mode (default:
                                        16)
  --order-htmls [ORDER_HTMLS ...]       Optional review reports used to define one shared
                                        motif order in --motif-grid mode
  --fill-missing-profiles               Fill missing motif profiles from profiles embedded
                                        elsewhere in the review report
  --recompute-missing-profiles          Recompute missing motif profiles from project
                                        bigWigs and motif BEDs
  --repeat-column-labels {none,row}     Repeat comparison labels in every motif row
                                        (default: none)

Run arguments:
  --cores <int>                         Worker processes for recomputing missing motif
                                        profiles
  --verbosity <int>                     Level of output logging (0: silent, 1:
                                        errors/warnings, 2: info, 3: stats, 4: debug, 5:
                                        spam) (default: 3)

review-multi-comparisons

Review several diff-footprints comparisons in one place. Combine the existing reports into a browser bundle or a single HTML file, then switch between comparisons to explore motif statistics and aggregate profiles.

Example command

review-multi-comparisons \
  --inputs project/comparisons \
  --output-dir project/reports/review_multi_comparisons

Primary inputs

  • --inputs — report files or directories containing differential reports.
  • --output-dir — destination for the complete static bundle.

Main outputs

{bundle} is the --output-dir:

Path Meaning
{bundle}/index.html Report entry page, served together with the other bundle files.
{bundle}/app.js, {bundle}/plot_controls.js, and {bundle}/styles.css Local application code, shared plot behavior, and styling.
{bundle}/data/metadata.json Comparison index and payload checksums.
{bundle}/data/reports/{comparison}.json.gz Compact data for one comparison.
{bundle}/data/profiles/ Aggregate-profile shards loaded on demand.
{bundle}/data/logos/ Motif logo assets.

Project mode defaults to {project}/reports/review_multi_comparisons/index.html. Keep the entire directory together when copying or publishing a bundle.

Open the bundle

The bundle loads its data through a web server. To view it locally, run:

python -m http.server 8000 --bind 127.0.0.1 \
  --directory project/reports/review_multi_comparisons

Open http://127.0.0.1:8000/ in your browser. Stop the server with Ctrl+C when you finish. To open a report directly without a server, use the standalone output below.

To choose the bundle's opening view, add --default-comparison with the two condition or region labels, --default-aggregate-motifs with the motif names or IDs in display order, and --default-aggregate-plots with the number of panels. The selected labels and motifs must exist in the input reports. Use --documentation-url to add a documentation link.

Standalone output

Use --output-html instead of --output-dir to create one HTML file that you can open directly or share. It includes volcano plots, ranked motifs, logos, and SVG exports. Aggregate controls appear when the input reports contain profiles. Use --labels to give each input report a distinct name in the Comparison list, especially when reports compare the same condition pair.

In the ranked-motif plot, switch between differential footprint score and -log10(p-value) to change the ranking view. Use the legend to interpret the color scale and direction of change. In the volcano plot, use Label TFs to label selected TF names, motif IDs, or output prefixes, separated by commas. Select (none) in the highlight control to clear highlighting. SVG exports retain the selected comparison label and plot settings.

review-multi-comparisons --inputs baseline/report.html dose1/report.html dose2/report.html \
  --labels Baseline "Dose 1" "Dose 2" --output-html review.html

Complete options

usage: review-multi-comparisons [-h] [--inputs INPUTS [INPUTS ...]]
                                [--labels [LABELS ...]]
                                [--output-dir OUTPUT_DIR | --output-html OUTPUT_HTML]
                                [--outdir OUTDIR] [--layout {custom,project}]
                                [--default-comparison <group1> <group2>]
                                [--default-aggregate-motifs <motif> [<motif> ...]]
                                [--default-aggregate-plots <int>]
                                [--documentation-url <url>]
                                [--fill-missing-aggregate-profiles]
                                [--recompute-missing-aggregate-profiles]
                                [--aggregate-flank AGGREGATE_FLANK]
                                [--cores CORES] [--title TITLE]

Combine diff-footprints reports into a static browser bundle or one self-
contained HTML file.

options:
  -h, --help            show this help message and exit
  --inputs INPUTS [INPUTS ...]
                        diff-footprints HTML files or directories containing
                        diff_footprints_*.html files; directories are searched
                        recursively.
  --labels [LABELS ...]
                        Optional labels, one per resolved input HTML.
  --output-dir OUTPUT_DIR
                        Output directory for index.html, JavaScript, CSS, and
                        compact static data files.
  --output-html OUTPUT_HTML
                        One self-contained HTML report; aggregate profiles are
                        optional.
  --outdir OUTDIR       Project directory used with --layout project.
  --layout {custom,project}
                        Use fp-tools standard project output layout under
                        --outdir (default: project when only --outdir is
                        provided).
  --default-comparison <group1> <group2>
                        Region or condition pair initially shown in the static
                        browser
  --default-aggregate-motifs <motif> [<motif> ...]
                        Ordered motif IDs, names, or output prefixes initially
                        shown
  --default-aggregate-plots <int>
                        Number of aggregate profiles initially shown (default:
                        4; maximum: 12)
  --documentation-url <url>
                        Optional link back to the documentation site
  --fill-missing-aggregate-profiles
                        Fill missing motif aggregate panels from profiles
                        embedded elsewhere in the combined review payload.
  --recompute-missing-aggregate-profiles
                        Recompute still-missing motif aggregate panels from
                        project sample bigWigs and match-motifs BEDs.
  --aggregate-flank AGGREGATE_FLANK
                        Flank used when recomputing missing aggregate
                        profiles, or 'auto' to match the existing report axis
                        (default: auto).
  --cores CORES         Worker processes for --recompute-missing-aggregate-
                        profiles (default: all available cores).
  --title TITLE

run-yaml-workflow

Run one or more fp-tools commands from a saved YAML configuration. Use this to repeat a GUI run or apply the same settings to several samples or comparisons.

Example command

run-yaml-workflow --config workflow.yml --dry-run
run-yaml-workflow --config workflow.yml --run-root project/yaml_runs

Primary inputs

  • --config — YAML configuration exported by the GUI or written directly.
  • --run-root — optional directory for job status and logs; it does not replace each command's analysis output directory.
  • --dry-run — print the commands without starting the analysis.

Replace workflow.yml with your saved configuration. The first command prints the jobs without running them; check their inputs and output paths before running the second command. Jobs run sequentially, and each job's output appears live in the terminal as well as in its saved logs. Commands that support a core budget use all available cores when cores is omitted or set to null. An explicit number is retained. A job-level cores: null also overrides a limit in defaults, allowing the execution machine to choose its available cores.

Main outputs

  • The analysis files documented for each command named in the YAML.
  • Standard output containing the expanded command lines when --dry-run is used.
  • {run_root}/{job_id}/config.yml and command.txt — saved settings and exact command for each job.
  • {run_root}/{job_id}/status.json, stdout.log, and stderr.log — completion state and captured command output.
  • {run_root}/batch_index.tsv — one-row-per-job batch status index.

Relative paths in the YAML are interpreted from the directory where you launch the command. Run from the same directory each time, or use absolute paths. If neither --run-root nor a YAML run_root is supplied, logs go to a new fp-tools-batch-{timestamp} folder in that directory.

Complete options

usage: run-yaml-workflow [-h] --config CONFIG [--run-root RUN_ROOT]
                         [--only [ONLY ...]] [--dry-run] [--list-jobs]
                         [--fail-fast]

Run fp-tools jobs from a YAML config file.

options:
  -h, --help           show this help message and exit
  --config CONFIG      Path to YAML config.
  --run-root RUN_ROOT  Optional directory for run metadata/logs.
  --only [ONLY ...]    Optional tool filter, e.g. diff-footprints.
  --dry-run            Print expanded commands without running.
  --list-jobs          List expanded jobs and exit.
  --fail-fast          Stop at first failed job.

fp-tools-gui

Launch the browser interface for configuring and running fp-tools commands. The Windows and Apple Silicon desktop downloads present the same interface in a native fp-tools application window.

Bulk GUI workflows start from coordinate-sorted BAM/BAI files and matching peak BED files. The GUI does not perform FASTQ-to-BAM preprocessing. Missing inputs and unsupported options are reported before a run starts.

The GUI is available through the Python package, the complete container, and the self-contained desktop downloads on the release page.

Example command

fp-tools-gui --host 127.0.0.1 --port 8891 --run-dir project/gui_runs

Primary inputs

  • --host — interface on which the GUI listens (default: 127.0.0.1).
  • --port — fixed browser port.
  • --run-dir — directory for GUI-managed configurations and runs.

Start an analysis

  1. Select the workflow or command from the sidebar.
  2. Enter your input paths and output directory, then select Update page config.
  3. Review the displayed command and resolve any validation errors before starting the run.
  4. Open Run History to check progress, read logs, and find the output files.

Use the Config page to save the settings as YAML or load a previous run's configuration.

Main outputs

  • {run_dir}/{timestamp}_{label}/config.yml — saved YAML settings for the run.
  • {run_dir}/{timestamp}_{label}/status.json, launcher_stdout.log, and launcher_stderr.log — overall run status and logs.
  • {run_dir}/{timestamp}_{label}/{job_id}/status.json, command.txt, stdout.log, and stderr.log — each job's status, exact command, and analysis logs.
  • The analysis files documented by the selected command, written to the output directory you chose.

A saved YAML can also be run from the command line with run-yaml-workflow --config {run_dir}/{timestamp}_{label}/config.yml.

Local computer

Open the desktop executable to use the native application window. When using the Python package, run fp-tools-gui; a browser opens after the server is ready. If it does not, open the local URL printed in the terminal.

Remote Linux server

Start fp-tools on the server. --no-browser prevents it from opening a browser on the server, and the default host setting limits access to that server:

fp-tools-gui --no-browser --port 8891

On your computer, create an SSH tunnel and keep that terminal open:

ssh -N -L 8891:127.0.0.1:8891 USER@SERVER

Open http://127.0.0.1:8891 on your computer to use the server's GUI.

Complete options

usage: fp-tools-gui [-h] [--host HOST] [--port PORT] [--run-dir RUN_DIR]
                    [--no-browser]

Launch the fp-tools browser interface.

options:
  -h, --help         show this help message and exit
  --host HOST        Bind address (default: 127.0.0.1).
  --port PORT        Optional fixed port (default: first free port from 8891).
  --run-dir RUN_DIR  Directory for GUI-managed runs.
  --no-browser       Do not open a local browser automatically.

fp-tools-runtime

Check or install the external programs used by prepare-atac and discover-motifs. fp-tools manages these programs separately from your other software. Linux supports read preparation and motif discovery; macOS and Windows support motif discovery.

Example command

fp-tools-runtime status

Primary inputs

The status action takes no input files. To install the programs used for motif discovery before your first run:

fp-tools-runtime install meme

On Linux, use fp-tools-runtime install core for the default prepare-atac workflow, or fp-tools-runtime install homer for its homer-atac profile.

Main outputs

The command reports each runtime component, platform, installation state, and cache location. Commands that need a managed component install it on first use, so manual installation is optional. If an installation is damaged, run fp-tools-runtime repair meme (or substitute the affected component), then check fp-tools-runtime status again.

Complete options

usage: fp-tools-runtime [-h] {status,install,repair} ...

Inspect, install, or repair the managed fp-tools runtime.

positional arguments:
  {status,install,repair}
    status              Report managed runtime availability and installation
                        state.
    install             Install a runtime component.
    repair              Repair a runtime component.

options:
  -h, --help            show this help message and exit

discover-motifs

Find recurring DNA sequence patterns in candidate footprint regions, without starting from a known motif list. Provide candidate intervals and a reference genome, or use sequences you have already extracted into a FASTA file.

Example command

discover-motifs --candidates project/samples/sample/footprints/sample_candidate_footprints.bed --genome hg38.fa.gz \
  --flank 75 --method streme --known-motif-db jaspar2026_vertebrates --outdir project/de_novo/sample --execute

Primary inputs

  • --candidates — candidate-footprint BED intervals.
  • --genome — reference FASTA matching the candidate coordinates; required with --candidates.
  • --flank — bases included on each side of a candidate center.
  • --method — discovery method; the example uses STREME.
  • --known-motif-db — optional known-motif database for Tomtom matching.
  • --outdir — directory for candidate FASTA files and discovery results.
  • --execute — run discovery immediately using the managed MEME Suite runtime.

In the example, --flank 75 extracts up to 75 bases on each side of each candidate center. Without --execute, the command writes the FASTA and a command script for inspection. It does not run discovery or create result summaries.

Main outputs

{outdir} is the selected discovery directory:

Path Meaning
{outdir}/candidate_sequences.fa Reference sequences extracted around candidate footprint intervals.
{outdir}/run_motif_discovery.sh Commands for discovery, optional known-motif matching, and summary reports.
{outdir}/streme/streme.txt, meme/meme.txt, or dreme/dreme.txt Discovered motif models from the selected method when --execute is used.
{outdir}/tomtom/tomtom.tsv Optional similarity matches to the selected known-motif database.
{outdir}/motif_summary.tsv and motif_summary.html Motif tables written after successful execution.

Choose your sequence input

Create candidate intervals with call-footprints --output-bed. If you already have a sequence FASTA, pass it with --fasta instead of --candidates and --genome. In that case the command uses your existing sequences directly.

Complete options

usage: discover-motifs [-h] (--fasta FASTA | --candidates CANDIDATES)
                       [--genome GENOME] [--flank FLANK] --outdir OUTDIR
                       [--script SCRIPT] [--method {meme,dreme,streme}]
                       [--known-motifs KNOWN_MOTIFS [KNOWN_MOTIFS ...]]
                       [--known-motif-db KNOWN_MOTIF_DB] [--list-motif-dbs]
                       [--extra-args ...] [--execute]
                       [--runtime {auto,managed,system,container}]

Prepare or run a de novo motif discovery command plan.

options:
  -h, --help            show this help message and exit
  --fasta FASTA         Existing candidate FASTA.
  --candidates CANDIDATES
                        Candidate BED from call-footprints --output-bed or
                        another BED-like source.
  --genome GENOME       Genome FASTA, required when --candidates is used.
  --flank FLANK         If >0 with --candidates, export +/- flank bp around
                        each candidate center.
  --outdir OUTDIR       External motif discovery output directory.
  --script SCRIPT       Output shell script path. Defaults to
                        <outdir>/run_motif_discovery.sh.
  --method {meme,dreme,streme}
  --known-motifs KNOWN_MOTIFS [KNOWN_MOTIFS ...]
                        Optional known motif database file(s) for Tomtom
                        comparison.
  --known-motif-db KNOWN_MOTIF_DB
                        Optional built-in motif database for Tomtom
                        comparison.
  --list-motif-dbs      List available built-in motif databases and exit.
  --extra-args ...      Additional arguments appended to MEME/DREME/STREME.
  --execute             Run the generated script immediately.
  --runtime {auto,managed,system,container}
                        External-tool runtime: auto/managed provisions the
                        pinned fp-tools runtime, system uses PATH, and
                        container uses the complete image (default: auto).

summarize-motifs

Turn completed motif-discovery results into a table of motif sequences, significance values, and optional matches to known motifs. Add an HTML report to review the results in a browser.

Example command

summarize-motifs --meme-txt project/de_novo/sample/streme/streme.txt \
  --tomtom-tsv project/de_novo/sample/tomtom/tomtom.tsv \
  --out-tsv project/de_novo/sample/motif_summary.tsv --out-html project/de_novo/sample/motif_summary.html

Primary inputs

  • --meme-txt — discovery output such as meme.txt, streme.txt, or dreme.txt.
  • --tomtom-tsv — optional Tomtom known-motif matches.
  • --out-tsv — compact output table for discovered motifs and matches.
  • --out-html — optional browser report.

Omit --tomtom-tsv if you did not run known-motif matching. This command reads existing results; it does not rerun motif discovery.

Main outputs

  • The --out-tsv file — tab-separated discovered motif IDs, consensus sequences, significance values, and known-database matches when available.
  • The --out-html file — optional portable HTML table containing the same summary and consensus-sequence logos when available.

Open motif_summary.html in a browser or import motif_summary.tsv into a spreadsheet. Known-motif matches help identify candidate TF families; they do not establish which TF is bound.

Complete options

usage: summarize-motifs [-h] [--meme-txt MEME_TXT] [--tomtom-tsv TOMTOM_TSV]
                        --out-tsv OUT_TSV [--out-html OUT_HTML]
                        [--title TITLE]

Summarize MEME/Tomtom outputs into TSV and HTML reports.

options:
  -h, --help            show this help message and exit
  --meme-txt MEME_TXT   MEME text output, usually meme.txt.
  --tomtom-tsv TOMTOM_TSV
                        Tomtom TSV output, usually tomtom.tsv.
  --out-tsv OUT_TSV     Output motif summary TSV.
  --out-html OUT_HTML   Optional output HTML report.
  --title TITLE

pseudobulk-fragments

Combine single-cell ATAC-seq fragments into groups such as cell types or donor–cell-type pairs. Each group becomes a pseudobulk sample for downstream analysis.

Example command

pseudobulk-fragments --fragments pbmc_fragments.tsv.gz --annotations cell_annotations.tsv --group-by cell_type \
  --genome-sizes hg38.chrom.sizes --write-cutsite-bigwigs --outdir project/pseudobulk/fragments

Primary inputs

  • --fragments — TSV or TSV.GZ with chromosome, start, end, and barcode in its first four columns; a fifth column can contain fragment counts.
  • --annotations — TSV or CSV with barcode and the column named by --group-by (for example, cell_type).
  • --group-by — annotation column used to define pseudobulk groups; use donor,cell_type to keep donors separate within each cell type.
  • --genome-sizes — two-column chromosome-name and length file matching the fragments; required for the example's signal tracks.
  • --write-cutsite-bigwigs — write cut-site bigWigs for retained groups.
  • --outdir — directory for grouped fragments, tracks, and QC outputs.

Main outputs

{group} is the annotation value converted to a filename-safe name. Paths below are relative to {outdir}:

Path Meaning
{group}.fragments.tsv or {group}.fragments.tsv.gz Fragments assigned to the group; compressed/indexed form is controlled by the command options.
{group}.fragments.tsv.gz.tbi Optional Tabix index for random genomic access.
{group}.cutsites.cpm.bw Optional CPM-normalized cut-site signal bigWig written by --write-cutsite-bigwigs.
{group}.pseudo_pairs.sorted.bam and .bai Optional pseudo-paired alignment written with --write-pseudo-bams; use atac-correct --read_shift 0 0 on these files.
pseudobulk_manifest.tsv Per-group paths, cell/fragment counts, and filter status.
fp_tools_manifest.yml Machine-readable run settings and retained groups.
pseudobulk_downstream_commands.sh Optional generated downstream command examples.

Match cell barcodes

By default, barcode matching ignores a trailing suffix such as -1. Add --no-strip-barcode-suffix when suffixes distinguish cells in your dataset. Use --barcode-column if your annotation barcode column has a different name.

Complete options

usage: pseudobulk-fragments [-h] --fragments FRAGMENTS --annotations
                            ANNOTATIONS --group-by GROUP_BY
                            [--barcode-column BARCODE_COLUMN]
                            [--no-strip-barcode-suffix]
                            [--include-chroms INCLUDE_CHROMS]
                            [--exclude-chroms EXCLUDE_CHROMS]
                            [--min-cells MIN_CELLS]
                            [--min-fragments MIN_FRAGMENTS] --outdir OUTDIR
                            [--compress-output] [--index-output]
                            [--write-cutsite-bigwigs] [--write-pseudo-bams]
                            [--no-cpm-normalize] [--write-downstream-commands]
                            [--genome-sizes GENOME_SIZES] [--cores CORES]

Group single-cell ATAC fragments into pseudobulk fragment files.

options:
  -h, --help            show this help message and exit
  --fragments FRAGMENTS
                        10x-style fragments TSV/TSV.GZ with barcode in column
                        4.
  --annotations ANNOTATIONS
                        Cell annotation TSV or CSV.
  --group-by GROUP_BY   Comma-separated annotation columns to group by, e.g.
                        donor,cell_type.
  --barcode-column BARCODE_COLUMN
                        Annotation barcode column (default: barcode).
  --no-strip-barcode-suffix
                        Require exact barcode matches instead of matching
                        AAAC-1 to AAAC.
  --include-chroms INCLUDE_CHROMS
                        Comma-separated chromosomes to keep, e.g.
                        chr1,chr2,chrX.
  --exclude-chroms EXCLUDE_CHROMS
                        Comma-separated chromosomes to skip, e.g. chrM,chrY.
  --min-cells MIN_CELLS
                        Minimum cells for passes_filters (default: 1).
  --min-fragments MIN_FRAGMENTS
                        Minimum fragments for passes_filters (default: 1).
  --outdir OUTDIR       Output directory.
  --compress-output     Write grouped fragments as .tsv.gz files.
  --index-output        BGZF-compress and tabix-index grouped fragments for
                        random access.
  --write-cutsite-bigwigs
                        Write one sparse cut-site bigWig per kept pseudobulk
                        group.
  --write-pseudo-bams   Write sorted pseudo-paired BAMs for kept groups; use
                        atac-correct --read_shift 0 0 on these BAMs.
  --no-cpm-normalize    Write raw cut counts instead of CPM-normalized bigWig
                        values.
  --write-downstream-commands
                        Write a shell script for BED/BAM/bigWig generation
                        from kept pseudobulk groups.
  --genome-sizes GENOME_SIZES
                        Two-column chromosome sizes file used by generated
                        bedtools/UCSC commands and cut-site bigWigs.
  --cores CORES         Cores for compression, bigWig writing, and generated
                        samtools commands (default: all available cores).

find-signature-fp

Score selected motif sites in individual cells and plot the results on a UMAP and in heatmaps. The command pools signal from nearby cells to reduce sparsity, so the scores describe relative footprint signatures rather than independent binding calls for every cell.

Example command

find-signature-fp --annotations cell_annotations.tsv --fragments pbmc_fragments.tsv.gz --h5ad genomic_bin_counts.h5ad \
  --all-motif-diff-dir project/pseudobulk/diff_footprints \
  --all-motif-results project/pseudobulk/diff_footprints/pseudobulk_diff_footprints_results.txt \
  --outdir project/pseudobulk/signature_fp

Primary inputs

  • --annotations — TSV or CSV with required barcode, cell_type, snap_cell_type, umap_1, and umap_2 columns. The GUI and command check these columns before analysis starts.
  • --fragments — single-cell fragment file with chromosome, start, end, and barcode in its first four columns; the command creates a missing Tabix index by default.
  • --h5ad — AnnData file with matching cell names, genomic-bin counts, and a boolean selected column in var. Bin names must use chromosome:start-end; the default bin size is 500 bases.
  • --all-motif-diff-dir — completed diff-footprints directory containing motif-site BED files.
  • --all-motif-results — motif-level results from that directory; supply it together with --all-motif-diff-dir.
  • --outdir — directory for per-cell scores, heatmaps, and UMAP figures.

Use cell_type for broad labels and snap_cell_type for detailed labels; they can be the same if you have one annotation level. AnnData cell names must match the annotation barcodes. Smoothing uses obsm['X_spectral'] if available, otherwise the annotation UMAP coordinates. The companion activity scores need the genomic-bin counts, so an embedding-only AnnData file is not sufficient.

Main outputs

Under {outdir} the default names include:

Path Meaning
knn_footprint_signature_scores.tsv Per-cell KNN-smoothed footprint protection scores for selected TFs.
knn_footprint_orientation_summary.tsv Direction/orientation checks used to make marker scores comparable.
chromvar_like_motif_activity_scores.tsv Companion accessibility-derived motif activity scores.
knn_footprint_signature_umap.svg Per-marker footprint-signature UMAP panels.
per_cell_footprint_signature_heatmap.svg Selected-marker per-cell heatmap.
single_cell_footprinting_summary.svg Combined heatmap and representative UMAP summary.
all_motif_per_cell_footprint_signature_heatmap.tsv Optional all-motif score matrix and metadata when all-motif inputs are supplied.

Open the summary SVG first, then use the score tables to inspect individual cells. Additional top-motif plots and all-TF review PDFs are produced when all-motif inputs are supplied.

Choose marker TFs

Choose TFs with --markers TF1,TF2. The default markers are STAT6,FOSB,CEBPA,IRF8,RELA,ZNF683,NR4A1,SMAD3; each selected TF must have motif sites in your inputs. In the GUI, enter one marker per line. YAML accepts either a list such as markers: [STAT6, CEBPA, ZNF683] or the comma-separated form markers: STAT6,CEBPA,ZNF683. Saved GUI configurations run through run-yaml-workflow without conversion.

For selected-marker reports only, use --tf-site-dir and omit --all-motif-diff-dir and --all-motif-results from the example. This alternative directory must contain files named {TF}.motif_hits.bed or {TF}.motif_peaks.bed.

Complete options

usage: find-signature-fp [-h] --annotations ANNOTATIONS --fragments FRAGMENTS
                         --h5ad H5AD [--tf-site-dir TF_SITE_DIR] --outdir
                         OUTDIR [--markers MARKERS]
                         [--max-sites-per-tf MAX_SITES_PER_TF] [--knn KNN]
                         [--flank FLANK]
                         [--center-half-width CENTER_HALF_WIDTH]
                         [--flank-inner FLANK_INNER]
                         [--flank-outer FLANK_OUTER] [--bin-size BIN_SIZE]
                         [--marker-groups MARKER_GROUPS]
                         [--all-motif-diff-dir ALL_MOTIF_DIFF_DIR]
                         [--all-motif-results ALL_MOTIF_RESULTS]
                         [--all-motif-score-table ALL_MOTIF_SCORE_TABLE]
                         [--marker-score-table MARKER_SCORE_TABLE]
                         [--all-motif-batch-size ALL_MOTIF_BATCH_SIZE]
                         [--max-sites-per-motif MAX_SITES_PER_MOTIF]
                         [--max-motifs MAX_MOTIFS]
                         [--top-motif-signatures-per-cell-type TOP_MOTIF_SIGNATURES_PER_CELL_TYPE]
                         [--top-motif-min-specificity TOP_MOTIF_MIN_SPECIFICITY]
                         [--summary-output-prefix SUMMARY_OUTPUT_PREFIX]
                         [--all-tf-review-prefix ALL_TF_REVIEW_PREFIX]
                         [--all-tf-review-panels-per-page ALL_TF_REVIEW_PANELS_PER_PAGE]
                         [--skip-all-tf-review-pdfs]
                         [--no-create-fragment-index]

Generate per-cell footprint-signature heatmaps and UMAP reports.

options:
  -h, --help            show this help message and exit
  --annotations ANNOTATIONS
                        Cell annotation TSV/CSV requiring barcode, cell_type,
                        snap_cell_type, umap_1, and umap_2 columns.
  --fragments FRAGMENTS
                        10x-style fragments TSV/TSV.GZ used to count cut sites
                        around motif centers.
  --h5ad H5AD           AnnData with matching cell barcodes, genomic-bin
                        counts, and boolean var['selected']; optional
                        embeddings support nearest-neighbor smoothing.
  --tf-site-dir TF_SITE_DIR
                        Optional directory containing marker motif-site BED
                        files named by TF. When omitted, marker sites are
                        taken from --all-motif-diff-dir and --all-motif-
                        results.
  --outdir OUTDIR       Output directory for signature score tables, heatmaps,
                        and UMAP reports.
  --markers MARKERS     Comma-separated marker TFs to score and plot (default:
                        STAT6,FOSB,CEBPA,IRF8,RELA,ZNF683,NR4A1,SMAD3).
  --max-sites-per-tf MAX_SITES_PER_TF
                        Maximum marker motif sites per TF for selected-marker
                        UMAP scoring (default: 1500).
  --knn KNN             Number of nearest neighbors used to smooth per-cell
                        cut-site profiles (default: 75).
  --flank FLANK         Motif-centered half-window in bp for fragment counting
                        (default: 100).
  --center-half-width CENTER_HALF_WIDTH
                        Half-width in bp of the protected center window
                        (default: 10).
  --flank-inner FLANK_INNER
                        Inner flank distance from motif center in bp (default:
                        25).
  --flank-outer FLANK_OUTER
                        Outer flank distance from motif center in bp (default:
                        100).
  --bin-size BIN_SIZE   Bin size for the companion chromVAR-like motif
                        activity score (default: 500).
  --marker-groups MARKER_GROUPS
                        Comma-separated TF:cell_type pairs used to orient KNN
                        marker scores for UMAP review.
  --all-motif-diff-dir ALL_MOTIF_DIFF_DIR
                        Optional differential-footprint output directory
                        containing */beds/*_all.bed files for all-motif per-
                        cell heatmap scoring.
  --all-motif-results ALL_MOTIF_RESULTS
                        Differential-footprint results table used to order and
                        annotate all-motif heatmap rows.
  --all-motif-score-table ALL_MOTIF_SCORE_TABLE
                        Existing all-motif per-cell heatmap TSV to redraw
                        all/top heatmaps without rescoring fragments.
  --marker-score-table MARKER_SCORE_TABLE
                        Existing KNN marker score table used to orient
                        selected marker rows in top heatmaps and summary
                        UMAPs.
  --all-motif-batch-size ALL_MOTIF_BATCH_SIZE
                        Number of motif signatures to score per batch for the
                        all-motif heatmap.
  --max-sites-per-motif MAX_SITES_PER_MOTIF
                        Maximum motif instances per motif for all-motif
                        heatmap scoring; use 0 for all sites.
  --max-motifs MAX_MOTIFS
                        Optional all-motif smoke-test limit.
  --top-motif-signatures-per-cell-type TOP_MOTIF_SIGNATURES_PER_CELL_TYPE
                        Top cell-type-specific all-motif signatures to keep
                        per broad cell type (default: 40).
  --top-motif-min-specificity TOP_MOTIF_MIN_SPECIFICITY
                        Minimum dominant-vs-next cell-type mean z-score
                        difference for top all-motif heatmap rows (default:
                        0.5).
  --summary-output-prefix SUMMARY_OUTPUT_PREFIX
                        Output prefix for the combined heatmap and UMAP
                        summary SVG when all-motif heatmap data are available.
  --all-tf-review-prefix ALL_TF_REVIEW_PREFIX
                        Output prefix for three multi-page all-TF signature
                        review PDFs grouped by dominant broad cell type.
  --all-tf-review-panels-per-page ALL_TF_REVIEW_PANELS_PER_PAGE
                        Number of TF signature UMAP panels per all-TF review
                        PDF page (default: 12).
  --skip-all-tf-review-pdfs
                        Do not write the three all-TF signature review PDFs.
  --no-create-fragment-index
                        Do not create a tabix index for the fragment file when
                        it is missing.

sc-footprinting

Analyze single-cell ATAC-seq by combining cells into pseudobulk groups, calling footprints and motif differences between groups, then mapping selected footprint signatures back to individual cells.

Example command

sc-footprinting --fragments pbmc_fragments.tsv.gz --annotations cell_annotations.tsv --h5ad genomic_bin_counts.h5ad \
  --group-by cell_type --genome-sizes hg38.chrom.sizes --genome hg38.fa.gz --peaks merged_peaks.bed \
  --motif-db jaspar2026_vertebrates --outdir project/pseudobulk

For custom motifs alone, replace --motif-db jaspar2026_vertebrates with --motifs custom.jaspar. Supply both options only to combine the motif sets.

Primary inputs

The workflow uses all available cores by default. Grouping progress and child command output appear in the terminal; child output is also saved in the run's logs directory. In the graphical user interface (GUI), leave the Cores field blank for automatic selection, or enter a limit.

  • --fragments — TSV or TSV.GZ with chromosome, start, end, and cell barcode in its first four columns.
  • --annotations — cell annotation TSV or CSV with required barcode, cell_type, snap_cell_type, umap_1, and umap_2, plus any additional grouping columns. The GUI and command check these columns before analysis starts.
  • --h5ad — AnnData file containing the same cells and genomic-bin counts used for the companion motif-activity scores. See the requirements below.
  • --group-by — annotation column used to define pseudobulk groups.
  • --genome-sizes — two-column chromosome-name and length file used to write grouped signal tracks.
  • --genome — reference genome FASTA matching the fragments and peak coordinates.
  • --peaks — accessible-region BED file.
  • --motif-db — optional built-in motif database name.
  • --motifs — optional custom motif files. Custom motifs alone use only those files; explicitly supply --motif-db as well to combine them. When neither option is supplied, the command uses jaspar2026_vertebrates.
  • --outdir — directory for pseudobulk tracks, motif results, and reports.

The per-cell reports need annotation barcodes that match the AnnData cell names. Use cell_type for the broad cell labels and snap_cell_type for detailed labels; these can be the same if you have only one annotation level. UMAP coordinates go in umap_1 and umap_2.

The AnnData file needs genomic-bin names such as chr1:0-500, a count matrix, and a boolean selected column in its feature annotations (var). The companion activity calculation uses 500-base bins. Nearest-neighbor smoothing uses obsm['X_spectral'] if present, otherwise the annotation UMAP coordinates. An embedding-only AnnData file is not sufficient.

Main outputs

Paths below are relative to --outdir; {group} is a retained pseudobulk group:

Path Meaning
pseudobulk/{group}.fragments.tsv.gz and .tbi Indexed fragments for each retained cell group.
pseudobulk/{group}.cutsites.cpm.bw Group cut-site signal bigWig.
pseudobulk/{group}.pseudo_pairs.sorted.bam and .bai Pseudo-paired alignment used for bias correction.
atacorrect/{group}/{group}_corrected.bw Bias-corrected cut-site signal per group.
footprints/{group}_footprints.bw Footprint score signal per group.
diff_footprints/pseudobulk_diff_footprints_results.txt Optional motif-level group comparison results.
plots/single_cell_footprinting/ Per-cell score tables, heatmaps, and UMAP figures from find-signature-fp.
pseudobulk_footprint_manifest.tsv Group paths and workflow completion state.
pseudobulk_footprint_commands.sh Exact generated commands for reproducibility.
logs/{stage}.stdout.log and {stage}.stderr.log Captured output for each stage.

Check the manifest for completed groups, then review the motif comparison and per-cell heatmaps. Per-cell scores combine signal from neighboring cells to reduce sparsity; interpret them as relative signatures, not independent binding calls for every cell. Use --single-cell-signature-markers to choose the TFs shown in the marker reports.

Complete options

usage: sc-footprinting [-h] --fragments FRAGMENTS --annotations ANNOTATIONS
                       --group-by GROUP_BY --outdir OUTDIR
                       [--genome-sizes GENOME_SIZES] --genome GENOME --peaks
                       PEAKS [--blacklist BLACKLIST]
                       [--barcode-column BARCODE_COLUMN]
                       [--no-strip-barcode-suffix]
                       [--include-chroms INCLUDE_CHROMS]
                       [--exclude-chroms EXCLUDE_CHROMS] [--groups GROUPS]
                       [--min-cells MIN_CELLS] [--min-fragments MIN_FRAGMENTS]
                       [--no-cpm-normalize] [--top-n TOP_N]
                       [--read-shift FWD REV] [--motifs [MOTIFS ...]]
                       [--motif-db MOTIF_DB] [--list-motif-dbs]
                       [--peak-header PEAK_HEADER] [--diff-prefix DIFF_PREFIX]
                       [--diff-normalization {condition-quantile,sample-quantile,none}]
                       [--diff-plot-aggregate {sig,all,top,off}]
                       [--skip-excel | --no-skip-excel]
                       [--tf-site-dir TF_SITE_DIR]
                       [--site-summary SITE_SUMMARY] [--tfs TFS]
                       [--plot-flank PLOT_FLANK] [--plot-script PLOT_SCRIPT]
                       --h5ad SINGLE_CELL_SIGNATURE_H5AD
                       [--single-cell-signature-outdir SINGLE_CELL_SIGNATURE_OUTDIR]
                       [--single-cell-signature-markers SINGLE_CELL_SIGNATURE_MARKERS]
                       [--single-cell-signature-fig-prefix SINGLE_CELL_SIGNATURE_FIG_PREFIX]
                       [--single-cell-signature-all-motif-score-table SINGLE_CELL_SIGNATURE_ALL_MOTIF_SCORE_TABLE]
                       [--single-cell-signature-marker-score-table SINGLE_CELL_SIGNATURE_MARKER_SCORE_TABLE]
                       [--single-cell-signature-top-per-cell-type SINGLE_CELL_SIGNATURE_TOP_PER_CELL_TYPE]
                       [--single-cell-signature-top-min-specificity SINGLE_CELL_SIGNATURE_TOP_MIN_SPECIFICITY]
                       [--single-cell-signature-knn SINGLE_CELL_SIGNATURE_KNN]
                       [--single-cell-signature-max-sites-per-motif SINGLE_CELL_SIGNATURE_MAX_SITES_PER_MOTIF]
                       [--single-cell-signature-max-motifs SINGLE_CELL_SIGNATURE_MAX_MOTIFS]
                       [--cores CORES] [--resume] [--force] [--dry-run]
                       [--fail-fast]

Run the complete pseudobulk and per-cell footprint workflow from single-cell
fragments.

options:
  -h, --help            show this help message and exit
  --fragments FRAGMENTS
                        10x-style fragments TSV/TSV.GZ with barcode in column
                        4.
  --annotations ANNOTATIONS
                        Cell annotation TSV or CSV.
  --group-by GROUP_BY   Comma-separated annotation columns to group by.
  --outdir OUTDIR       Output directory for the full pseudobulk footprint
                        workflow.
  --genome-sizes GENOME_SIZES
                        Two-column chromosome sizes file used for fragment-
                        derived cut-site bigWigs.
  --genome GENOME       Genome FASTA for atac-correct.
  --peaks PEAKS         Peak BED used for atac-correct and footprint scoring.
  --blacklist BLACKLIST
                        Optional blacklist BED for atac-correct.
  --barcode-column BARCODE_COLUMN
                        Annotation barcode column (default: barcode).
  --no-strip-barcode-suffix
                        Require exact barcode matches instead of matching
                        AAAC-1 to AAAC.
  --include-chroms INCLUDE_CHROMS
                        Comma-separated chromosomes to keep.
  --exclude-chroms EXCLUDE_CHROMS
                        Comma-separated chromosomes to skip.
  --groups GROUPS       Comma-separated pseudobulk groups to process after
                        grouping; default processes all retained groups.
  --min-cells MIN_CELLS
                        Minimum cells for passes_filters (default: 1).
  --min-fragments MIN_FRAGMENTS
                        Minimum fragments/reads for passes_filters (default:
                        1).
  --no-cpm-normalize    Write raw cut counts instead of CPM-normalized cut-
                        site bigWigs for fragment input.
  --top-n TOP_N         Optional top N candidate footprints per group.
  --read-shift FWD REV  Override the atac-correct read shift for fragment cut
                        sites (default: 0 0).
  --motifs [MOTIFS ...]
                        Optional motif file(s); when provided, run motif-aware
                        diff-footprints on pseudobulk footprint tracks.
  --motif-db MOTIF_DB   Built-in motif database; defaults to
                        jaspar2026_vertebrates only when neither --motif-db
                        nor --motifs is supplied. Explicitly supply both to
                        combine them.
  --list-motif-dbs      List available built-in motif databases and exit.
  --peak-header PEAK_HEADER
                        Optional peak-header file passed to diff-footprints.
  --diff-prefix DIFF_PREFIX
                        Prefix for optional motif-aware diff-footprints
                        outputs.
  --diff-normalization {condition-quantile,sample-quantile,none}
                        Normalization mode for optional motif-aware diff-
                        footprints outputs (default: none).
  --diff-plot-aggregate {sig,all,top,off}
                        Aggregate plot selection for optional motif-aware
                        diff-footprints HTML/PDF outputs.
  --skip-excel, --no-skip-excel
                        Skip Excel files for optional diff-footprints outputs
                        (default: on).
  --tf-site-dir TF_SITE_DIR
                        Optional motif-centered BED directory to plot
                        corrected footprint aggregates.
  --site-summary SITE_SUMMARY
                        Optional motif-centered site summary TSV for plotting.
  --tfs TFS             Comma-separated TFs or 'auto' for plotting (default:
                        auto).
  --plot-flank PLOT_FLANK
                        Flank for optional aggregate plots (default: 100).
  --plot-script PLOT_SCRIPT
                        Plotting script path for optional aggregate plots.
  --h5ad SINGLE_CELL_SIGNATURE_H5AD, --single-cell-signature-h5ad SINGLE_CELL_SIGNATURE_H5AD
                        AnnData with matching cell barcodes, genomic-bin
                        counts, and boolean var['selected']; an embedding
                        alone is insufficient.
  --single-cell-signature-outdir SINGLE_CELL_SIGNATURE_OUTDIR
                        Output directory for optional per-cell signature
                        reports (default:
                        <outdir>/plots/single_cell_footprinting).
  --single-cell-signature-markers SINGLE_CELL_SIGNATURE_MARKERS
                        Comma-separated marker TFs for optional per-cell
                        signature UMAPs (default:
                        STAT6,FOSB,CEBPA,IRF8,RELA,ZNF683,NR4A1,SMAD3).
  --single-cell-signature-fig-prefix SINGLE_CELL_SIGNATURE_FIG_PREFIX
                        Output prefix for the combined single-cell footprint-
                        signature SVG (default: single_cell_footprinting).
  --single-cell-signature-all-motif-score-table SINGLE_CELL_SIGNATURE_ALL_MOTIF_SCORE_TABLE
                        Existing all-motif per-cell signature TSV; skips
                        rescoring all motif sites for the signature heatmap.
  --single-cell-signature-marker-score-table SINGLE_CELL_SIGNATURE_MARKER_SCORE_TABLE
                        Existing KNN marker score TSV used for marker rows and
                        UMAP plots.
  --single-cell-signature-top-per-cell-type SINGLE_CELL_SIGNATURE_TOP_PER_CELL_TYPE
                        Top all-motif signatures to keep per cell type in the
                        signature heatmap (default: 40).
  --single-cell-signature-top-min-specificity SINGLE_CELL_SIGNATURE_TOP_MIN_SPECIFICITY
                        Minimum dominant-vs-next cell-type z-score difference
                        for top heatmap rows (default: 0.5).
  --single-cell-signature-knn SINGLE_CELL_SIGNATURE_KNN
                        KNN size for optional per-cell footprint-signature
                        smoothing (default: 75).
  --single-cell-signature-max-sites-per-motif SINGLE_CELL_SIGNATURE_MAX_SITES_PER_MOTIF
                        Maximum motif instances per motif for optional all-
                        motif per-cell heatmap scoring; use 0 for all sites
                        (default: 200).
  --single-cell-signature-max-motifs SINGLE_CELL_SIGNATURE_MAX_MOTIFS
                        Optional smoke-test limit for all-motif per-cell
                        heatmap scoring.
  --cores CORES         Optional core limit for grouping, correction, and
                        footprint scoring (default: all available cores).
  --resume              Skip atac-correct/call-footprints steps whose expected
                        outputs already exist.
  --force               Run atac-correct/call-footprints even if outputs
                        already exist.
  --dry-run             Write manifests and commands without running atac-
                        correct, call-footprints, motif detection, or plots.
  --fail-fast           Stop after the first failed group command.