Command reference¶
Choose a command below to see when to use it, how to prepare its inputs, an example run, and the files it writes. The complete options follow each guide and are also available with <command> --help.
| Command | Purpose |
|---|---|
prepare-atac |
Linux CLI or Linux container only. Download public ATAC-seq reads or use local FASTQ files, then trim, align, filter, call peaks, calculate alignment coverage, and write QC files. The GUI and native macOS/Windows installations start from coordinate-sorted BAM/BAI and matching peak BED files. |
bulk-footprinting |
Run a complete bulk ATAC-seq analysis from aligned reads and peak regions to footprint scores, motif comparisons, and interactive reports. |
atac-correct |
Correct ATAC-seq cut-site signal for Tn5 sequence bias. Start with coordinate-sorted BAM files, adjacent BAI indexes, and matching peak BED files. Use the corrected signal as the input to call-footprints. |
call-footprints |
Calculate a footprint score at each position in accessible regions using the corrected bigWigs from atac-correct. The score measures local cut-site depletion relative to the surrounding signal; higher scores indicate stronger footprint evidence. |
match-motifs |
Find motif matches in accessible regions and measure the footprint score at each match. This produces a motif summary for each sample and classifies sites as predicted bound or unbound. |
diff-footprints |
Find motifs whose footprint scores differ between conditions. Start with the per-sample results from match-motifs. You can also compare two sets of genomic regions measured in the same samples. |
normalize-bigwig |
Put corrected cut-site tracks on a comparable signal scale using the same background regions for every sample. This is an optional step after atac-correct; use the resulting tracks for downstream scoring or plotting when you need this normalization. |
plot-aggregate |
Plot the average cut-site signal around motif sites to inspect the shape of a footprint. Start with corrected bigWigs from atac-correct and motif results from match-motifs. Outputs can be static figures or an interactive HTML report. |
review-multi-comparisons |
Review several diff-footprints comparisons in one place. Combine the existing reports into a browser bundle or a single HTML file, then switch between comparisons to explore motif statistics and aggregate profiles. |
run-yaml-workflow |
Run one or more fp-tools commands from a saved YAML configuration. Use this to repeat a GUI run or apply the same settings to several samples or comparisons. |
fp-tools-gui |
Launch the browser interface for configuring and running fp-tools commands. The Windows and Apple Silicon desktop downloads present the same interface in a native fp-tools application window. |
fp-tools-runtime |
Check or install the external programs used by prepare-atac and discover-motifs. fp-tools manages these programs separately from your other software. Linux supports read preparation and motif discovery; macOS and Windows support motif discovery. |
discover-motifs |
Find recurring DNA sequence patterns in candidate footprint regions, without starting from a known motif list. Provide candidate intervals and a reference genome, or use sequences you have already extracted into a FASTA file. |
summarize-motifs |
Turn completed motif-discovery results into a table of motif sequences, significance values, and optional matches to known motifs. Add an HTML report to review the results in a browser. |
pseudobulk-fragments |
Combine single-cell ATAC-seq fragments into groups such as cell types or donor–cell-type pairs. Each group becomes a pseudobulk sample for downstream analysis. |
find-signature-fp |
Score selected motif sites in individual cells and plot the results on a UMAP and in heatmaps. The command pools signal from nearby cells to reduce sparsity, so the scores describe relative footprint signatures rather than independent binding calls for every cell. |
sc-footprinting |
Analyze single-cell ATAC-seq by combining cells into pseudobulk groups, calling footprints and motif differences between groups, then mapping selected footprint signatures back to individual cells. |
prepare-atac¶
Linux CLI or Linux container only. Download public ATAC-seq reads or use local FASTQ files, then trim, align, filter, call peaks, calculate alignment coverage, and write QC files. The GUI and native macOS/Windows installations start from coordinate-sorted BAM/BAI and matching peak BED files.
Example command
prepare-atac --samples metadata.tsv --genome hg38 --outdir project
Primary inputs
The workflow uses all available cores by default, sharing its core budget across concurrent samples. Tool diagnostics appear live in the terminal and remain in the sample logs; alignment data continue to go to their output files.
--samples— TSV or CSV sample sheet. Providesample,condition, andfastq_1paths or URLs; addfastq_2for paired-end reads. A public sequencing run can instead be supplied inrun_accession.--genome— managedhg38ormm10reference label, or a custom label used with explicit reference options.--outdir— project directory represented by{project}below.
For paired-end local files, metadata.tsv can contain:
sample condition fastq_1 fastq_2
control_1 control reads/control_1_R1.fastq.gz reads/control_1_R2.fastq.gz
treated_1 treated reads/treated_1_R1.fastq.gz reads/treated_1_R2.fastq.gz
Repeated rows with the same sample, condition, and replicate combine
technical sequencing runs. Different sample values sharing a condition are
biological replicates.
Main outputs
For each {sample}, the default modern profile writes:
| Path | Meaning |
|---|---|
{project}/samples/{sample}/alignment/{sample}.filtered.bam |
Coordinate-sorted, filtered ATAC-seq alignment used downstream. |
{project}/samples/{sample}/alignment/{sample}.filtered.bam.bai |
Samtools index for the filtered BAM. |
{project}/samples/{sample}/peaks/{sample}.narrowPeak |
MACS3 narrow-peak calls before project-level merging. |
{project}/samples/{sample}/tracks/{sample}.rp10m.bw |
Sequencing-depth-normalized alignment coverage bigWig for viewing read coverage. |
{project}/samples/{sample}/qc/{sample}.fastp.html |
Interactive Fastp read-trimming QC report. |
{project}/samples/{sample}/qc/{sample}.fastp.json |
Machine-readable Fastp metrics. |
{project}/samples/{sample}/qc/flagstat.tsv |
Samtools alignment and filtering counts. |
{project}/samples/{sample}/qc/fragment_lengths.tsv |
Fragment-length distribution used to inspect ATAC-seq periodicity. |
{project}/samples/{sample}/qc/metrics.json |
Consolidated per-sample QC metrics. |
{project}/samples/{sample}/qc/commands.log |
External commands used for that sample. |
Project-level files include:
| Path | Meaning |
|---|---|
{project}/peaks/merged_peaks.bed |
Union of sample peak intervals. |
{project}/peaks/merged_peaks_filtered.bed |
Analysis peak set after excluded chromosomes are removed. |
{project}/metadata/resolved_runs.tsv |
Resolved local/downloaded FASTQ files and run grouping. |
{project}/metadata/samples.tsv |
sample, condition, bam, and peaks table ready for bulk-footprinting. |
{project}/reports/qc_summary.tsv |
Cross-sample QC summary. |
Reference options¶
With --genome hg38 or --genome mm10, the command downloads and verifies the
matching reference and blacklist as needed. Use --reference-dir to choose
where references are cached.
For a custom genome label, provide --fasta and --macs-genome-size. Supply
--bowtie2-index if you already have an index; otherwise the command builds one
from the FASTA. --blacklist replaces the managed hg38/mm10 blacklist, while
--no-blacklist disables it. Custom FASTA inputs use no blacklist unless you
provide one.
Complete options
usage: prepare-atac [-h] [--samples SAMPLES] [--genome GENOME]
[--outdir OUTDIR] [--config CONFIG]
[--profile {modern,homer-atac}] [--id-column ID_COLUMN]
[--sample-column SAMPLE_COLUMN]
[--condition-column CONDITION_COLUMN]
[--include INCLUDE [INCLUDE ...]]
[--reference-dir REFERENCE_DIR] [--fasta FASTA]
[--bowtie2-index BOWTIE2_INDEX]
[--blacklist BLACKLIST | --no-blacklist] [--tss TSS]
[--macs-genome-size MACS_GENOME_SIZE] [--cores CORES]
[--max-parallel-samples MAX_PARALLEL_SAMPLES]
[--memory-gb MEMORY_GB] [--keep-intermediates]
[--no-resume] [--fail-fast] [--dry-run] [--doctor]
[--write-default-config PATH]
[--runtime {auto,managed,system,container}]
Download single- or paired-end ATAC-seq reads and prepare filtered BAM, peak
BED, sequencing-depth-normalized alignment coverage bigWig, and QC files.
options:
-h, --help show this help message and exit
--samples SAMPLES TSV/CSV sample metadata with an ID/run_accession or
fastq_1 column.
--genome GENOME Named genome (hg38/mm10) or custom genome label.
--outdir OUTDIR Output project directory.
--config CONFIG Optional preprocessing YAML; CLI values override
packaged defaults.
--profile {modern,homer-atac}
Processing method: modern uses fastp, samtools, and
MACS3; homer-atac uses Trim Galore, Picard, and HOMER
(default: modern).
--id-column ID_COLUMN
Explicit accession column name.
--sample-column SAMPLE_COLUMN
Explicit sample-name column.
--condition-column CONDITION_COLUMN
Explicit condition column.
--include INCLUDE [INCLUDE ...]
Only process these sample names or accessions.
--reference-dir REFERENCE_DIR
Reference cache root (default: ~/.cache/fp-
tools/references).
--fasta FASTA Custom reference FASTA.
--bowtie2-index BOWTIE2_INDEX
Existing Bowtie2 index prefix.
--blacklist BLACKLIST
Custom blacklist BED.
--no-blacklist Disable the managed blacklist for hg38 or mm10.
--tss TSS Optional TSS BED for enrichment QC.
--macs-genome-size MACS_GENOME_SIZE
MACS3 genome size or hs/mm shorthand for custom
genomes.
--cores CORES Optional total core limit (default: all available
cores).
--max-parallel-samples MAX_PARALLEL_SAMPLES
Maximum samples processed concurrently (default:
config value).
--memory-gb MEMORY_GB
Enforced total memory budget in GiB; reserves 8 GiB
for the host.
--keep-intermediates Keep trimmed FASTQs and intermediate alignment files
under each sample .work directory.
--no-resume Recompute completed samples even when fingerprints
match.
--fail-fast Stop after the first failed sample.
--dry-run Validate and list the planned samples without
downloading or processing.
--doctor Report external preprocessing dependencies and exit.
--write-default-config PATH
Write the fully documented default YAML and exit.
--runtime {auto,managed,system,container}
External-tool runtime: auto/managed provisions the
pinned fp-tools runtime, system uses PATH, and
container uses the complete image (default: auto).
bulk-footprinting¶
Run a complete bulk ATAC-seq analysis from aligned reads and peak regions to footprint scores, motif comparisons, and interactive reports.
The bulk workflow guide provides minimal sample and comparison tables for a two-condition analysis.
Example command
bulk-footprinting --sample-table samples.tsv --comparison-table comparisons.tsv --genome hg38 \
--outdir project
Primary inputs
The workflow uses all available cores by default. It shows each stage and its command output live in the terminal, while keeping the log files listed below. The complete options include an optional core limit.
In the graphical user interface (GUI), leave the Cores field blank to select all available cores when the workflow runs. Enter a number only to limit it. Saved configurations retain automatic selection across machines.
--sample-table— TSV withsample,condition,bam, andpeakscolumns. Each BAM must be coordinate-sorted and have a matching BAI index.--comparison-table— TSV withcomparison,cond1, andcond2columns. Use condition names from the sample table.--genome— managedhg38ormm10assembly, or a reference FASTA matching the BAM and peak coordinates.--outdir— project output directory.
Use the same genome assembly and chromosome names for every BAM and BED file.
Main outputs
{project} is the --outdir, {sample} comes from the sample table, and
{comparison} comes from the comparison table:
| Path | Meaning |
|---|---|
{project}/samples/{sample}/atac_correct/{sample}_corrected.bw |
Bias-corrected cut-site signal. |
{project}/samples/{sample}/footprints/{sample}_footprints.bw |
Footprint score signal. |
{project}/samples/{sample}/match_motifs/motif_matches_results.txt |
Per-sample motif summary and binding calls. |
{project}/comparisons/{comparison}/diff_footprints_results.txt |
Motif-level differential statistics. |
{project}/comparisons/{comparison}/diff_footprints_{cond1}_{cond2}.html |
Portable interactive comparison report. |
{project}/reports/review_multi_comparisons/index.html |
Static browser combining every requested comparison. |
{project}/reports/review_multi_comparisons.html |
Aggregate-free portable review written when standalone HTML review mode is selected. |
{project}/logs/bulk_footprinting/bulk_footprinting_commands.sh |
Exact commands generated for the workflow stages. |
{project}/logs/bulk_footprinting/{stage}.stdout.log and {stage}.stderr.log |
Stage-specific logs for troubleshooting. |
Start by opening the comparison HTML report. Use the combined review to compare motif results across all requested comparisons, and inspect aggregate profiles alongside the statistics.
Add --dry-run to check the inputs and inspect the commands before starting.
For a combined review, include each condition pair only once. Reversing the
conditions still counts as the same pair. Repeated pairs are rejected before
reference downloads or analysis, with the comparison IDs and table line numbers
shown in the error. To run separate comparisons with the same condition labels
but different sample subsets, use --review-format none.
Reference and motif options¶
Choosing hg38 or mm10 downloads and verifies the matching reference and
blacklist as needed. Use --reference-dir to choose where they are cached.
--blacklist replaces a managed assembly's blacklist, while
--no-blacklist disables it. Custom FASTA inputs never infer a blacklist.
Choose a packaged motif database with --motif-db, provide custom files with
--motifs, or combine both options. With neither option, the workflow uses
jaspar2026_vertebrates. Run bulk-footprinting --list-motif-dbs to list the
packaged databases.
If you have FASTQ files on Linux, run
prepare-atac first, then provide its generated
metadata/samples.tsv to this command.
Complete options
usage: bulk-footprinting [-h] --sample-table SAMPLE_TABLE --comparison-table
COMPARISON_TABLE --genome GENOME
[--reference-dir REFERENCE_DIR]
[--blacklist BLACKLIST | --no-blacklist]
[--motifs [MOTIFS ...]] [--motif-db MOTIF_DB]
[--list-motif-dbs]
[--normalization {none,condition-quantile,sample-quantile}]
[--plot-aggregate {sig,all,top,off}]
[--review-format {auto,bundle,standalone,none}]
--outdir OUTDIR [--cores CORES] [--resume] [--force]
[--dry-run] [--fail-fast]
[--runtime {auto,managed,system,container}]
Run the complete bulk ATAC-seq footprinting workflow for explicit comparisons.
options:
-h, --help show this help message and exit
--sample-table SAMPLE_TABLE
TSV with sample, condition, BAM, and peak BED columns.
--comparison-table COMPARISON_TABLE
TSV with comparison, cond1, and cond2 columns.
--genome GENOME Managed assembly (hg38 or mm10) or a reference FASTA
matching the BAM and peak coordinates.
--reference-dir REFERENCE_DIR
Managed reference cache root (default: ~/.cache/fp-
tools/references).
--blacklist BLACKLIST
Blacklist BED overriding the managed assembly
blacklist.
--no-blacklist Disable the automatic hg38/mm10 blacklist.
--motifs [MOTIFS ...]
Optional motif files.
--motif-db MOTIF_DB Built-in motif database to use alone or combine with
--motifs (default when neither is provided:
jaspar2026_vertebrates).
--list-motif-dbs List available built-in motif databases and exit.
--normalization {none,condition-quantile,sample-quantile}
Differential-stage normalization (default: none).
--plot-aggregate {sig,all,top,off}
Aggregate profiles generated by diff-footprints
(default: all).
--review-format {auto,bundle,standalone,none}
Combined-review output; auto uses standalone when
aggregation is off and bundle otherwise (default:
auto).
--outdir OUTDIR Project output directory.
--cores CORES Optional total worker core limit (default: all
available cores).
--resume Skip stages whose expected outputs are complete.
--force Rerun stages even when outputs already exist.
--dry-run Validate inputs and print the commands without running
them.
--fail-fast Stop at the first failed stage.
--runtime {auto,managed,system,container}
External-tool runtime: auto/managed provisions the
pinned fp-tools runtime, system uses PATH, and
container uses the complete image (default: auto).
atac-correct¶
Correct ATAC-seq cut-site signal for Tn5 sequence bias. Start with
coordinate-sorted BAM files, adjacent BAI indexes, and matching peak BED files.
Use the corrected signal as the input to call-footprints.
Example command
atac-correct \
--sample-table project/metadata/samples.tsv \
--genome hg38.fa.gz \
--blacklist hg38.blacklist.bed \
--outdir project
Primary inputs
--sample-table— tab-separated file with one row per sample and the columnssample,condition,bam, andpeaks. Give each sample a unique name and use the same condition label for biological replicates.--genome— reference FASTA whose chromosome names and assembly match every BAM and peak BED.--blacklist— optional assembly-matched BED of regions to exclude from bias estimation and corrected output.--outdir— project directory represented by{project}below.
Use the sample-table example to prepare your inputs. Replace the FASTA and blacklist paths with your own files. A custom FASTA does not automatically select a blacklist.
Main outputs
For each {sample} in the table, this command writes:
| Path | Meaning |
|---|---|
{project}/samples/{sample}/atac_correct/{sample}_corrected.bw |
Bias-corrected cut-site signal. Positive positions have more observed cuts than expected; negative positions have fewer. |
{project}/samples/{sample}/atac_correct/{sample}_atacorrect.pdf |
Diagnostic plots comparing learned Tn5 sequence bias before and after correction. Omitted with --skip-qc. |
{project}/samples/{sample}/atac_correct/{sample}_AtacBias.pickle |
Saved bias model for reuse. Downstream commands use the corrected bigWig and do not need this file. |
With --write-tracks all, the same directory also contains:
| Path | Meaning |
|---|---|
{sample}_uncorrected.bw |
Observed base-resolution cut-site signal after the configured forward/reverse read shifts and sequencing-depth normalization. |
{sample}_bias.bw |
Tn5 sequence-bias score predicted from the reference sequence. |
{sample}_expected.bw |
Expected cut-site signal after the sequence-bias score is scaled to local observed cuts. |
Project-level peak outputs are {project}/peaks/merged_peaks.bed and
{project}/peaks/merged_peaks_filtered.bed. A direct single-BAM run writes the
same {prefix}_*.bw, {prefix}_atacorrect.pdf, and
{prefix}_AtacBias.pickle patterns directly under {outdir}.
Open the QC PDF to inspect the sequence-bias diagnostics, then use the corrected bigWig for footprint scoring. A negative corrected value means fewer cuts than the bias model expected; it does not by itself identify a bound TF.
Complete options
usage: atac-correct [-h] [--bams [<bam> ...]] [--fragments [<fragments.tsv.gz> ...]]
[-g <fasta>] [-p [<bed> ...]] [--regions-in <bed>]
[--regions-out <bed>] [--blacklist <bed>] [--extend <int>]
[--split-strands] [--norm-off] [--write-tracks [<track> ...]]
[--track-off [<track> ...]] [--skip-qc]
[--scale-corrected {auto,none,q95}] [--scale-background <bed>]
[--scale-corrected-bigwigs [<bigwig> ...]]
[--scale-target {median,mean}] [--scale-chrom-sizes <chrom.sizes>]
[--merged-peaks-out <bed>] [--drop-chroms [<chrom> ...]]
[--k_flank <int>] [--read_shift <int> <int>] [--bg_shift <int>]
[--window <int>] [--score_mat <mat>] [--bias-pkl <obj>]
[--prefix <prefix>] [--sample-names [<name> ...]]
[--sample-table <tsv>] [--layout {custom,project}]
[--sample-output-root <directory>] [--outdir <directory>]
[--cores <int>] [--sample-workers <int>] [--split <int>]
[--verbosity <int>]
__________________________________________________________________________________________
fp-tools atac-correct
__________________________________________________________________________________________
atac-correct corrects ATAC-seq cutsite signal for Tn5 sequence bias.
Usage:
atac-correct --bams <reads.bam> [<more_reads.bam> ...] --genome <genome.fa> --peaks
<merged_peaks.bed> [<sample_peaks.bed> ...]
Output files:
- <outdir>/<sample>/<sample>_corrected.bw for multi-BAM runs
- optional auxiliary tracks with --write-tracks
- <outdir>/<prefix>_atacorrect.pdf
------------------------------------------------------------------------------------------
Required arguments:
--bams [<bam> ...] One or more .bam files containing reads to be corrected
--fragments [<fragments.tsv.gz> ...]
One or more 10x-style fragment files; uses cut sites
directly without creating pseudo-BAMs
-g <fasta>, --genome <fasta> A .fasta-file containing whole genomic sequence
-p [<bed> ...], --peaks [<bed> ...]
One shared merged peak BED, or multiple per-sample peak
BEDs to merge internally
Optional arguments:
--regions-in <bed> Input regions for estimating bias (default: regions not
in peaks.bed)
--regions-out <bed> Output regions (default: peaks.bed)
--blacklist <bed> Blacklisted regions in .bed-format (default: None)
--extend <int> Extend output regions with basepairs
upstream/downstream (default: 100)
--split-strands Write out tracks per strand
--norm-off Switches off normalization based on number of reads
--write-tracks [<track> ...] Cut-site signal bigWigs to write (default: corrected;
use all for corrected, observed/uncorrected, sequence-
bias, and expected signals)
--track-off [<track> ...] Compatibility option to switch off individual bigWig
tracks after --write-tracks is resolved
--skip-qc Skip atac-correct diagnostic PDF and pre/post bias
verification counts. Corrected bigWig output is
unchanged.
--scale-corrected {auto,none,q95}
Optionally q95-scale corrected bigWigs after
correction. In auto mode this runs only when --scale-
corrected-bigwigs has more than one track (default:
auto)
--scale-background <bed> Shared BED regions used to estimate q95 scaling for
--scale-corrected
--scale-corrected-bigwigs [<bigwig> ...]
Corrected bigWigs to q95-scale together. Include the
current sample's corrected bigWig or omit to scale only
the current output
--scale-target {median,mean} Across-sample q95 target for --scale-corrected
(default: median)
--scale-chrom-sizes <chrom.sizes>
Optional chromosome sizes file for scaled bigWig output
validation
--merged-peaks-out <bed> Path for internally merged peak BED when multiple
--peaks files are supplied (default:
<outdir>/merged_peaks.bed)
--drop-chroms [<chrom> ...] Drop any chromosomes in the list from the correction.
The default is to drop the mitochondrial chromosome.
Default: ['chrM', 'chrMT', 'M', 'MT', 'Mito']
Advanced options (defaults are suitable for most analyses):
--k_flank <int> Flank +/- of cutsite to estimate bias from (default:
12)
--read_shift <int> <int> Read shift for forward and reverse reads (default: 4
-5)
--bg_shift <int> Read shift for estimation of background frequencies
(default: 100)
--window <int> Window size for calculating expected signal (default:
100)
--score_mat <mat> Type of matrix to use for bias estimation (PWM/DWM)
(default: DWM)
--bias-pkl <obj> Path to a pre-calculated AtacBias.pkl-object, as output
from a previous atac-correct run (default: None). Can
be used to bypass the internal bias estimation.
Run arguments:
--prefix <prefix> Prefix for output files in single-BAM runs (default:
BAM filename stem)
--sample-names [<name> ...] Sample labels for --bams (default: BAM filename stems)
--sample-table <tsv> Project sample table with sample, condition, bam, and
peaks columns
--layout {custom,project} Use fp-tools standard project output layout under
--outdir (default: project when --sample-table is
provided)
--sample-output-root <directory>
Sample output root; writes each sample under
<root>/<sample>/atac_correct, typically
<project>/samples
--outdir <directory> Output directory for files (default: current working
directory)
--cores <int> Number of cores to use for computation (default: all
available cores)
--sample-workers <int> Number of samples to process concurrently for multi-BAM
runs (default: auto when --cores is set)
--split <int> Split of multiprocessing jobs (default: 100)
--verbosity <int> Level of output logging (0: silent, 1: errors/warnings,
2: info, 3: stats, 4: debug, 5: spam) (default: 3)
call-footprints¶
Calculate a footprint score at each position in accessible regions using the
corrected bigWigs from atac-correct. The score measures local cut-site
depletion relative to the surrounding signal; higher scores indicate stronger
footprint evidence.
Example command
call-footprints \
--signals A_corrected.bw B_corrected.bw \
--sample-names A B \
--regions merged_peaks.bed \
--sample-output-root project/samples
Primary inputs
--signals— one bias-corrected cut-site signal bigWig per sample.--sample-names— labels in the same order as--signals.--regions— BED intervals in which scores are calculated; normally the project merged, filtered peaks.--sample-output-root— root represented by{sample_root}below.
Replace A_corrected.bw and B_corrected.bw with your corrected tracks. In a
project, these are under project/samples/{sample}/atac_correct/; use
project/peaks/merged_peaks_filtered.bed for the regions.
Main outputs
| Path | Meaning |
|---|---|
{sample_root}/{sample}/footprints/{sample}_footprints.bw |
Base-resolution footprint score bigWig used by match-motifs and diff-footprints. |
{sample_root}/{sample}/footprints/{sample}_candidate_footprints.bed |
Optional local score maxima for de novo motif discovery; written with --call-candidates. |
user-selected *.npz |
Optional compressed scale-by-position score arrays when --score multiscale is used with an NPZ output option. |
In direct mode, --output result.bw writes exactly result.bw; multiple
signals written through --outdir {outdir} use
{outdir}/{signal_stem}_footprints.bw.
Pass the footprint score tracks to match-motifs to summarize evidence at
known motif sites. Keep the corrected cut-site tracks for aggregate plots,
which show the signal shape around those sites.
Complete options
usage: call-footprints [-h] [-s <bigwig>] [--signals [<bigwig> ...]] [-o <bigwig>]
[--outputs [<bigwig> ...]] [-r <bed>] [--score <score>]
[--absolute] [--extend <int>] [--smooth <int>]
[--min-limit <float>] [--max-limit <float>] [--scales [<int> ...]]
[--multiscale-summary <method>] [--output-multiscale-npz <npz>]
[--output-multiscale-npzs [<npz> ...]] [--output-bed <bed>]
[--output-beds [<bed> ...]] [--output-bed-dir <directory>]
[--call-candidates] [--top-n <int>] [--min-score <float>]
[--call-width <bp>] [--min-distance <bp>] [--fp-min <int>]
[--fp-max <int>] [--flank-min <int>] [--flank-max <int>]
[--footprint-kernel {fast,reference}] [--window <int>]
[--sample-names [<name> ...]] [--sample-table <tsv>]
[--layout {custom,project}] [--sample-output-root <directory>]
[--outdir <directory>] [--cores <int>] [--sample-workers <int>]
[--split <int>] [--verbosity <int>]
__________________________________________________________________________________________
fp-tools call-footprints
__________________________________________________________________________________________
call-footprints calculates footprint, sum, mean, or pass-through scores from one or more
bigWig signals and can optionally call ranked footprint candidate intervals.
Usage: call-footprints --signals <cutsites.bw> [<more_cutsites.bw> ...] --regions
<regions.bed> --outdir <output_dir>
or: call-footprints --signal <cutsites.bw> --regions <regions.bed> --output <output.bw>
Output:
- footprint score bigWig(s)
- optional candidate BED from --output-bed/--output-beds for de novo motif discovery
------------------------------------------------------------------------------------------
Required arguments:
-s <bigwig>, --signal <bigwig> A .bw file of ATAC-seq cutsite signal
--signals [<bigwig> ...] One or more .bw files of ATAC-seq cutsite signal
-o <bigwig>, --output <bigwig> Full path to output bigwig
--outputs [<bigwig> ...] Output bigWig path per --signals input
-r <bed>, --regions <bed> Genomic regions to run footprinting within
Optional arguments:
--score <score> Type of scoring to perform on cutsites
(footprint/sum/mean/none/multiscale) (default:
footprint)
--absolute Convert bigwig signal to absolute values before
calculating score
--extend <int> Extend input regions with bp (default: 100)
--smooth <int> Smooth output signal by mean in <bp> windows
(default: no smoothing)
--min-limit <float> Limit input bigwig value range (default: no lower
limit)
--max-limit <float> Limit input bigwig value range (default: no upper
limit)
Parameters for score == multiscale:
--scales [<int> ...] Window sizes for multiscale depletion scoring
(default: 8 16 24 32 64 100 147)
--multiscale-summary <method> How to collapse scale-specific scores into the
output bigWig (default: max)
--output-multiscale-npz <npz> Optional compressed NumPy sidecar with per-region
scale-by-position multiscale scores (only for
--score multiscale)
--output-multiscale-npzs [<npz> ...] Output multiscale NPZ sidecar per --signals input
Optional footprint candidate BED calling:
--output-bed <bed> Optional BED-like file of genomic coordinates for
footprint peaks used by de novo motif discovery
--output-beds [<bed> ...] Candidate BED path per --signals input
--output-bed-dir <directory> Directory for candidate BED files derived from
--signals names
--call-candidates In project/sample-output-root mode, also write
candidate footprint BEDs for de novo motif
discovery
--top-n <int> Keep only the top N footprint calls by score
(default: keep all)
--min-score <float> Minimum footprint score for candidate BED calls
(default: no threshold)
--call-width <bp> Width of candidate BED intervals centered on local
maxima (default: 50)
--min-distance <bp> Minimum distance between retained local footprint
centers within a region (default: 20)
Parameters for score == footprint:
--fp-min <int> Minimum footprint width (default: 20)
--fp-max <int> Maximum footprint width (default: 50)
--flank-min <int> Minimum range of flanking regions (default: 10)
--flank-max <int> Maximum range of flanking regions (default: 30)
--footprint-kernel {fast,reference} Footprint scoring kernel (default: fast; use
reference for the original scalar implementation)
Parameters for score == sum:
--window <int> The window for calculation of sum (default: 100)
Run arguments:
--sample-names [<name> ...] Sample labels for --signals when using project
layout
--sample-table <tsv> Project sample table with sample, condition, bam,
and peaks columns
--layout {custom,project} Use fp-tools standard project output layout under
--outdir (default: project when --sample-table is
provided)
--sample-output-root <directory> Sample output root; writes each sample under
<root>/<sample>/footprints, typically
<project>/samples
--outdir <directory> Output directory used with --signals when
--outputs is not supplied
--cores <int> Number of cores to use for computation (default:
all available cores)
--sample-workers <int> Number of input signals to process concurrently
for multi-signal runs (default: auto when --cores
is set)
--split <int> Split of multiprocessing jobs (default: 100)
--verbosity <int> Level of output logging (0: silent, 1:
errors/warnings, 2: info, 3: stats, 4: debug, 5:
spam) (default: 3)
match-motifs¶
Find motif matches in accessible regions and measure the footprint score at each match. This produces a motif summary for each sample and classifies sites as predicted bound or unbound.
Example command
match-motifs \
--signals A_footprints.bw B_footprints.bw \
--sample-names A B \
--genome hg38.fa.gz \
--peaks merged_peaks.bed \
--motif-db jaspar2026_vertebrates \
--sample-output-root project/samples
Primary inputs
--signals— one footprint score bigWig per sample.--sample-names— sample labels in the same order as--signals.--genome— assembly-matched reference FASTA used to scan motif sequences.--peaks— accessible-region BED searched for motif instances.--motif-db— packaged motif collection; the example uses JASPAR 2026 vertebrates.--sample-output-root— root represented by{sample_root}below.
Use the score tracks from call-footprints. In a project, they are under
project/samples/{sample}/footprints/; use the same reference FASTA and
filtered peaks as the earlier steps.
Main outputs
For each {sample}, the example writes to
{sample_root}/{sample}/match_motifs/:
| Path | Meaning |
|---|---|
motif_matches_results.txt |
Tab-separated motif summary with site counts and per-sample mean scores. |
motif_matches_distances.txt |
Motif-similarity distances used for motif clustering. |
cache/motif_sites.tsv.gz |
Compact scanned motif-site cache reusable by differential analysis. |
cache/background_scores.tsv.gz |
Compact background-score cache. |
{motif}/beds/{motif}_{sample}_all.bed |
All scanned instances for one motif. |
{motif}/beds/{motif}_{sample}_bound.bed |
Instances classified as bound in the sample. |
{motif}/beds/{motif}_{sample}_unbound.bed |
Instances classified as unbound in the sample. |
{motif} follows the selected --naming convention, such as
CTCF_MA0139.2. --motif-outputs summary omits the per-motif BED files but
keeps the summary and reusable caches.
The default shared scan in the example writes tab-separated summaries. An
independent sample scan can also write motif_matches_results.xlsx unless
--skip-excel is used. A joint analysis with repeated condition labels can
write motif_matches_replicate_motif_score_matrix.tsv; separate per-sample
folders do not contain that joint matrix.
Start with motif_matches_results.txt to review site counts and mean scores.
Bound and unbound are model-based classifications, not direct measurements of
TF binding. Related TFs can recognize similar motifs, so a motif name alone
does not distinguish every TF in a family.
To use your own motifs, replace --motif-db jaspar2026_vertebrates with
--motifs your_motifs.meme. Custom motifs alone do not add the default
database; JASPAR 2026 vertebrates is used when neither option is supplied.
Complete options
usage: match-motifs [-h] [--signals [<bigwig> ...]] [--peaks <bed>] [--genome <fasta>]
[--motifs [<motifs> ...]] [--motif-db <name>] [--list-motif-dbs]
[--sample-names [<name> ...]] [--cond-names [<name> ...]]
[--sample-table <tsv>] [--layout {custom,project}]
[--match-scan-mode {auto,shared,per-sample}]
[--sample-output-root <directory>] [--peak-header <file>]
[--naming <string>] [--motif-pvalue <float>] [--bound-pvalue <float>]
[--cluster-threshold <float>] [--pseudo <float>] [--skip-excel]
[--output-peaks <bed>] [--norm-off]
[--normalization {condition-quantile,sample-quantile,none}]
[--aggregate-signals [<bigwig> ...]]
[--plot-aggregate {sig,all,top,off}] [--plot-aggregate-top-n <int>]
[--plot-aggregate-motifs <motif> [<motif> ...]]
[--default-aggregate-plots <int>]
[--aggregate-pvalue-threshold <float>] [--aggregate-flank <bp>]
[--aggregate-normalization {match,none,sample-quantile,size-factor}]
[--aggregate-site-set {all,bound}]
[--motif-outputs {auto,summary,full}] [--report-label <text>]
[--outdir <directory>] [--prefix <prefix>] [--cores <int>]
[--sample-workers <int>] [--split <int>] [--debug] [--verbosity <int>]
__________________________________________________________________________________________
fp-tools match-motifs
__________________________________________________________________________________________
match-motifs scans motifs in open chromatin regions for one or more footprint score tracks
and infers sample-specific bound and unbound motif sites.
Usage:
match-motifs --signals <footprints.bw> [<more_footprints.bw> ...] --genome <genome.fasta>
--peaks <peaks.bed> [--motif-db jaspar2026_vertebrates | --motifs <motifs.txt>]
Output files:
- <outdir>/<prefix>_results.{txt,xlsx}
- <outdir>/<prefix>_distances.txt
- <outdir>/cache/* compact reuse caches
- <outdir>/<TF>/beds/*_all.bed, *_bound.bed, and *_unbound.bed by default
- optional <outdir>/<TF>/<TF>_overview.{txt,xlsx} with --motif-outputs full
------------------------------------------------------------------------------------------
Required arguments:
--signals [<bigwig> ...] One or more footprint score bigWigs (.bigwig format)
--peaks <bed> Peaks.bed containing open chromatin regions
--genome <fasta> Genome .fasta file
Optional arguments:
--motifs [<motifs> ...] Motif file(s) in pfm/jaspar/meme/transfac format; if
omitted, the built-in JASPAR 2026 vertebrates set is
used
--motif-db <name> Built-in motif database to use or add to --motifs
(default when --motifs is omitted:
jaspar2026_vertebrates)
--list-motif-dbs List available built-in motif databases and exit
--sample-names [<name> ...] Sample labels for --signals (default: prefix of each
--signals file)
--cond-names [<name> ...] Optional condition labels for --signals (default:
prefix of each --signals file)
--sample-table <tsv> Project sample table with sample and condition columns
plus bam/peaks for upstream steps
--layout {custom,project} Use fp-tools standard project output layout under
--outdir (default: project when --sample-table is
provided)
--match-scan-mode {auto,shared,per-sample}
Project match-motifs scan mode. auto uses one shared
motif scan for multi-sample project runs; per-sample
preserves independent sample scans.
--sample-output-root <directory>
Sample output root; writes one match_motifs folder
under <root>/<sample> for each input signal, typically
<project>/samples
--peak-header <file> File containing the header of --peaks separated by
whitespace or newlines (default: peak columns are named
"_additional_<count>")
--naming <string> Naming convention for TF output files ('id', 'name',
'name_id', 'id_name') (default: 'name_id')
--motif-pvalue <float> Set p-value threshold for motif scanning (default:
1e-4)
--bound-pvalue <float> Set p-value threshold for bound/unbound split (default:
0.001)
--cluster-threshold <float> Set the clustering threshold. Motifs below this
threshold will be assigned to one cluster (default:
0.5)
--pseudo <float> Pseudocount for calculating log2fcs (default: estimated
from data)
--skip-excel Skip creation of Excel files to speed up large motif
analyses
--output-peaks <bed> Gives the possibility to set the output peak set
differently than the input --peaks. This will limit all
analysis to the regions in --output-peaks. NOTE:
--peaks must still be set to the full peak set!
--norm-off Turn off normalization of footprint scores
--normalization {condition-quantile,sample-quantile,none}
Signal normalization mode (default: none; --norm-off
maps to none)
--aggregate-signals [<bigwig> ...]
Corrected cut-site bigWigs used for embedded aggregate
profiles
--plot-aggregate {sig,all,top,off}
Embed aggregate profiles in HTML reports for
significant, all, top-N, or no motifs (default: sig)
--plot-aggregate-top-n <int> Maximum number of motifs to aggregate when --plot-
aggregate sig/top or fallback selection is used
(default: 20)
--plot-aggregate-motifs <motif> [<motif> ...]
Ordered motif IDs, names, or output prefixes to embed
as aggregate profiles; overrides automatic aggregate-
motif selection
--default-aggregate-plots <int> Number of aggregate profiles initially displayed in
interactive reports (default: 4; maximum: 12)
--aggregate-pvalue-threshold <float>
P-value threshold for --plot-aggregate sig (default:
0.05)
--aggregate-flank <bp> Flank around motif centers for embedded aggregate
profiles (default: 100)
--aggregate-normalization {match,none,sample-quantile,size-factor}
Normalization for embedded aggregate profiles (default:
match --normalization)
--aggregate-site-set {all,bound}
Motif-site BEDs used for embedded aggregate profiles:
all motif hits or sample-specific bound sites (default:
all)
--motif-outputs {auto,summary,full}
Per-motif output mode. For match-motifs, auto writes
compact caches plus per-motif BED files; summary writes
only main result/report tables and caches; full writes
per-motif BED and overview files synchronously. For
diff-footprints, auto writes full motif outputs only
when aggregate reports need them.
--report-label <text> Optional method label shown under the report subtitle
in interactive HTML reports
--prefix <prefix> Prefix for overview files in --outdir folder (default:
motif_matches)
Run arguments:
--outdir <directory> Output directory to place motif tables, BED files, and
plots in (default: motif_matches_output)
--cores <int> Number of cores to use for computation (default: all
available cores)
--sample-workers <int> Number of input samples to process concurrently in
project/sample-output-root mode (default: auto when
--cores is set)
--split <int> Split of multiprocessing jobs (default: 100)
--debug Creates an additional '_debug.pdf'-file with debug
plots
--verbosity <int> Level of output logging (0: silent, 1: errors/warnings,
2: info, 3: stats, 4: debug, 5: spam) (default: 3)
diff-footprints¶
Find motifs whose footprint scores differ between conditions. Start with the
per-sample results from match-motifs. You can also compare two sets of
genomic regions measured in the same samples.
Example command
diff-footprints \
--sample-table project/metadata/samples.tsv \
--comparison-table project/metadata/comparisons.tsv \
--genome hg38.fa.gz \
--peaks project/peaks/merged_peaks_filtered.bed \
--motif-db jaspar2026_vertebrates \
--outdir project
Primary inputs
--sample-table— TSV withsampleandconditioncolumns. Reuse the table from the earlier steps; the command finds results under{project}/samples/{sample}/.--comparison-table— TSV withcomparison,cond1, andcond2columns. Each row names one comparison and its two condition labels.--genome— the reference FASTA used for motif matching.--peaks— accessible-region BED file.--motif-db— built-in motif database name.--outdir— project directory for statistics, figures, and HTML reports.
Condition labels must match the sample table. Include each biological replicate as its own sample row. See the sample and comparison tables for a minimal example.
Main outputs
Each comparison is written below
{project}/comparisons/{comparison}/, where {comparison} is taken from the
comparison table and {prefix} defaults to diff_footprints:
| Path | Meaning |
|---|---|
{prefix}_results.txt |
Tab-separated motif-level differential footprint statistics; change direction is cond1 - cond2. |
{prefix}_results.xlsx |
Excel copy in direct runs unless --skip-excel is used. Project comparison-table runs, including the example above, omit it. |
{prefix}_distances.txt |
Motif distances used for clustering related motifs. |
{prefix}_{cond1}_{cond2}.html |
Portable interactive report with volcano, motif, and embedded aggregate-profile views. |
{prefix}_replicate_report.tsv |
Long-form per-replicate diagnostic data when replicate reporting is active. |
{prefix}_replicate_summary.tsv |
Motif-level replicate agreement summary. |
{prefix}_replicate_report.png |
Replicate diagnostic figure. |
{prefix}_figures.pdf and {prefix}_clusters.pdf |
Optional static summaries written with --static-plots. |
{motif}/beds/{motif}_{condition}_bound.bed |
Motif instances classified as bound for a condition when full motif outputs are required. |
Open the HTML report to explore motifs, then use the text result table for
further analysis. Positive changes favor cond1; negative changes favor
cond2. Check replicate agreement and the aggregate cut-site profiles along
with the statistics when interpreting a difference.
Compare region sets¶
Region-set analyses use the same result/report patterns and add confidence intervals, motif prevalence, region counts, per-replicate effects, and matching balance to the result tables.
For a region-set comparison, use --comparison-axis regions, provide two or
more BED files with --regions, and name them with --region-labels. An
optional --region-strata-column preserves accessibility or other matching
strata during resampling. One sample uses a stratified label-permutation test;
two or more biological replicates use a paired empirical-Bayes model.
Use --plot-aggregate-motifs to choose an ordered aggregate panel without
limiting the motifs tested, and --default-aggregate-plots to set its initial
size.
Complete options
usage: diff-footprints [-h] [--signals [<bigwig> ...]] [--peaks <bed>] [--genome <fasta>]
[--motifs [<motifs> ...]] [--motif-db <name>] [--list-motif-dbs]
[--sample-names [<name> ...]] [--cond-names [<name> ...]]
[--comparison-axis {conditions,regions}] [--regions [<bed> ...]]
[--region-labels [<name> ...]] [--region-strata-column <int>]
[--region-permutations <int>] [--region-bootstrap <int>]
[--min-regions-per-set <int>] [--random-seed <int>]
[--sample-table <tsv>] [--comparison-table <tsv>]
[--layout {custom,project}] [--sample-dirs [<directory> ...]]
[--project-dir <directory>] [--peak-header <file>]
[--naming <string>] [--motif-pvalue <float>]
[--bound-pvalue <float>] [--cluster-threshold <float>]
[--pseudo <float>] [--time-series] [--time-course] [--skip-excel]
[--output-peaks <bed>] [--norm-off]
[--normalization {condition-quantile,sample-quantile,none}]
[--replicate-report {auto,on,off}] [--replicate-map <tsv>]
[--replicate-report-out <tsv>] [--replicate-summary-out <tsv>]
[--replicate-figure-out <figure>]
[--aggregate-signals [<bigwig> ...]]
[--plot-aggregate {sig,all,top,off}] [--plot-aggregate-top-n <int>]
[--plot-aggregate-motifs <motif> [<motif> ...]]
[--default-aggregate-plots <int>]
[--aggregate-pvalue-threshold <float>] [--aggregate-flank <bp>]
[--aggregate-normalization {match,none,sample-quantile,size-factor}]
[--aggregate-site-set {all,bound}] [--reuse-existing-results]
[--motif-outputs {auto,summary,full}] [--static-plots]
[--per-motif-plots] [--skew-report] [--report-label <text>]
[--outdir <directory>] [--prefix <prefix>] [--cores <int>]
[--split <int>] [--debug] [--verbosity <int>]
__________________________________________________________________________________________
fp-tools diff-footprints
__________________________________________________________________________________________
diff-footprints takes motifs, footprint signals, and genome sequence as input to compare
motif-associated footprint evidence across biological conditions or user-defined region
sets. Region-set mode gives every region equal weight and supports matching strata and
paired biological replicates.
Usage:
diff-footprints --signals <bigwig1> (<bigwig2> (...)) --genome <genome.fasta> --peaks
<peaks.bed> [--motif-db jaspar2026_vertebrates | --motifs <motifs.txt>]
diff-footprints --comparison-axis regions --signals <bigwig1> [<replicate2> ...] --regions
<set1.bed> <set2.bed> --genome <genome.fasta> [--region-labels <set1> <set2>]
Output files:
- <outdir>/<prefix>_results.{txt,xlsx}
- <outdir>/<prefix>_distances.txt
- <outdir>/<prefix>_<condition1>_<condition2>.html
- optional <outdir>/<prefix>_figures.pdf with --static-plots
- optional <outdir>/<prefix>_clusters.pdf with --static-plots
- optional <outdir>/<TF>/plots/<TF>_log2fcs.pdf with --per-motif-plots
- optional <outdir>/<TF>/<TF>_overview.{txt,xlsx} with --motif-outputs full
- <outdir>/<TF>/beds/<TF>_<condition>_bound.bed (per motif-condition pair)
- <outdir>/<TF>/beds/<TF>_<condition>_unbound.bed (per motif-condition pair)
------------------------------------------------------------------------------------------
Required arguments:
--signals [<bigwig> ...] Signal per condition (.bigwig format)
--peaks <bed> Peaks.bed containing open chromatin regions across all
conditions (not used with --comparison-axis regions)
--genome <fasta> Genome .fasta file
Optional arguments:
--motifs [<motifs> ...] Motif file(s) in pfm/jaspar/meme/transfac format; if
omitted, the built-in JASPAR 2026 vertebrates set is
used
--motif-db <name> Built-in motif database to use or add to --motifs
(default when --motifs is omitted:
jaspar2026_vertebrates)
--list-motif-dbs List available built-in motif databases and exit
--sample-names [<name> ...] Sample labels for --signals; distinct from --cond-names
and used for per-sample score columns, replicate
reports, and aggregate profiles (default: prefix of
each --signals file)
--cond-names [<name> ...] Condition labels for --signals; repeat names to define
biological replicates (default: prefix of each
--signals file)
--comparison-axis {conditions,regions}
Compare biological conditions or region sets measured
in the same sample(s) (default: conditions)
--regions [<bed> ...] Two or more non-overlapping BED files for --comparison-
axis regions
--region-labels [<name> ...] Labels for --regions (default: BED filename stems)
--region-strata-column <int> Optional 1-based BED column containing matching strata
--region-permutations <int> Within-stratum label permutations for a one-sample
region comparison (default: 100000)
--region-bootstrap <int> Within-set, within-stratum bootstrap samples for
confidence intervals (default: 1000)
--min-regions-per-set <int> Minimum motif-containing regions required in each set
(default: 10)
--random-seed <int> Random seed for region-set resampling (default: 1)
--sample-table <tsv> Project sample table with sample and condition columns
plus bam/peaks for upstream steps
--comparison-table <tsv> Project comparison table with comparison, cond1, and
cond2 columns
--layout {custom,project} Use fp-tools standard project output layout under
--outdir (default: project when --sample-table is
provided)
--sample-dirs [<directory> ...] Sample output folders containing match_motifs/ outputs
to reuse for differential analysis
--project-dir <directory> Parent folder containing sample output folders for
folder-based differential analysis, typically
<project>/samples
--peak-header <file> File containing the header of --peaks separated by
whitespace or newlines (default: peak columns are named
"_additional_<count>")
--naming <string> Naming convention for TF output files ('id', 'name',
'name_id', 'id_name') (default: 'name_id')
--motif-pvalue <float> Set p-value threshold for motif scanning (default:
1e-4)
--bound-pvalue <float> Set p-value threshold for bound/unbound split (default:
0.001)
--cluster-threshold <float> Set the clustering threshold. Motifs below this
threshold will be assigned to one cluster (default:
0.5)
--pseudo <float> Pseudocount for calculating log2fcs (default: estimated
from data)
--time-series Will only compare signals1<->signals2<->signals3 (...)
in order of input, and skip all-against-all comparison.
--time-course Alias for --time-series; compare adjacent ordered
conditions only.
--skip-excel Skip creation of Excel files to speed up large motif
analyses
--output-peaks <bed> Gives the possibility to set the output peak set
differently than the input --peaks. This will limit all
analysis to the regions in --output-peaks. NOTE:
--peaks must still be set to the full peak set!
--norm-off Turn off normalization of footprint scores across
conditions
--normalization {condition-quantile,sample-quantile,none}
Cross-sample normalization mode (default: none; --norm-
off maps to none)
--replicate-report {auto,on,off}
Write replicate-aware differential-footprint
diagnostics (default: auto for repeated condition names
or --replicate-map)
--replicate-map <tsv> Optional TSV with condition/replicate or
condition/n_replicates columns
--replicate-report-out <tsv> Output long-form replicate diagnostic TSV (default:
<outdir>/<prefix>_replicate_report.tsv)
--replicate-summary-out <tsv> Output replicate diagnostic summary TSV (default:
<outdir>/<prefix>_replicate_summary.tsv)
--replicate-figure-out <figure> Output replicate diagnostic figure (default:
<outdir>/<prefix>_replicate_report.png)
--aggregate-signals [<bigwig> ...]
Corrected cut-site bigWigs used for embedded aggregate
profiles
--plot-aggregate {sig,all,top,off}
Embed aggregate profiles in HTML reports for
significant, all, top-N, or no motifs (default: sig)
--plot-aggregate-top-n <int> Maximum number of motifs to aggregate when --plot-
aggregate sig/top or fallback selection is used
(default: 20)
--plot-aggregate-motifs <motif> [<motif> ...]
Ordered motif IDs, names, or output prefixes to embed
as aggregate profiles; overrides automatic aggregate-
motif selection
--default-aggregate-plots <int> Number of aggregate profiles initially displayed in
interactive reports (default: 4; maximum: 12)
--aggregate-pvalue-threshold <float>
P-value threshold for --plot-aggregate sig (default:
0.05)
--aggregate-flank <bp> Flank around motif centers for embedded aggregate
profiles (default: 100)
--aggregate-normalization {match,none,sample-quantile,size-factor}
Normalization for embedded aggregate profiles (default:
match --normalization)
--aggregate-site-set {all,bound}
Motif-site BEDs used for embedded aggregate profiles:
all motif hits or condition-specific bound sites
(default: all)
--reuse-existing-results Regenerate final diff-footprints reports from existing
<prefix>_results.txt and per-motif BEDs without
rescanning motifs
--motif-outputs {auto,summary,full}
Per-motif output mode. For match-motifs, auto writes
compact caches plus per-motif BED files; summary writes
only main result/report tables and caches; full writes
per-motif BED and overview files synchronously. For
diff-footprints, auto writes full motif outputs only
when aggregate reports need them.
--static-plots Also write static volcano and cluster PDF summaries. By
default diff-footprints writes the interactive HTML
report without these PDFs.
--per-motif-plots Also write one diagnostic log2 fold-change PDF per
motif. Disabled by default to keep diff-footprints
fast.
--skew-report Also write the optional skew/shift PDF report. Disabled
by default.
--report-label <text> Optional method label shown under the report subtitle
in interactive HTML reports
--prefix <prefix> Prefix for overview files in --outdir folder (default:
diff_footprints)
Run arguments:
--outdir <directory> Output directory to place motif tables, BED files, and
plots in (default: diff_footprints_output)
--cores <int> Number of cores to use for computation (default: all
available cores)
--split <int> Split of multiprocessing jobs (default: 100)
--debug Creates an additional '_debug.pdf'-file with debug
plots
--verbosity <int> Level of output logging (0: silent, 1: errors/warnings,
2: info, 3: stats, 4: debug, 5: spam) (default: 3)
normalize-bigwig¶
Put corrected cut-site tracks on a comparable signal scale using the same
background regions for every sample. This is an optional step after
atac-correct; use the resulting tracks for downstream scoring or plotting
when you need this normalization.
Example command
normalize-bigwig \
--sample-table project/metadata/samples.tsv \
--background project/peaks/merged_peaks_filtered.bed \
--outdir project \
--method background-scale \
--stat q95 \
--target median
Primary inputs
--sample-table— TSV withsampleandconditioncolumns; reuse the table fromatac-correct. Tracks are read from{project}/samples/{sample}/atac_correct/.--background— shared BED intervals used to calculate comparable background statistics.--outdir— project directory represented by{project}below.--method— transformation;background-scalemultiplies each signal by a shared-target scale factor.--stat— within-sample background statistic; the example uses the 95th percentile.--target— across-sample target for the selected statistic; the example uses the median.
Main outputs
| Path | Meaning |
|---|---|
{project}/samples/{sample}/normalize/{sample}_corrected_q95_scaled.bw |
Q95-scaled bias-corrected cut-site signal bigWig for one sample. |
{project}/normalize_bigwig_qc.tsv |
Background statistics, selected statistic, target, and scale factor for every sample. |
{project}/normalize_bigwig_manifest.tsv |
Sample-to-input/output signal mapping for downstream use. |
Check the QC table to see how much each track was scaled. The manifest lists the output paths to pass to the next command. Scaling changes signal magnitude; it does not by itself provide evidence of TF binding.
In custom layout, default outputs use
{outdir}/{input_stem}.background_scale_{stat}.bw, plus the two QC tables in
{outdir}. background-zscore instead writes standardized signal and uses a
method-specific filename suffix.
Complete options
usage: normalize-bigwig [-h] [--bigwigs BIGWIGS [BIGWIGS ...]]
[--background BACKGROUND] [--outdir OUTDIR]
[--sample-names [SAMPLE_NAMES ...]]
[--sample-table SAMPLE_TABLE]
[--layout {custom,project}]
[--sample-output-root SAMPLE_OUTPUT_ROOT]
[--method {background-scale,background-zscore,none}]
[--stat STAT] [--target {median,mean}]
[--chrom-sizes CHROM_SIZES] [--workers WORKERS]
Normalize input signal bigWigs using robust statistics from shared background
BED regions. For corrected cut-site bigWigs, the recommended method is
background-scale.
options:
-h, --help show this help message and exit
--bigwigs BIGWIGS [BIGWIGS ...]
Input bigWig files to normalize together.
--background BACKGROUND
Shared background BED used to estimate sample
statistics.
--outdir OUTDIR Output directory for normalized bigWig QC tables and
default outputs.
--sample-names [SAMPLE_NAMES ...]
Sample labels for --bigwigs when using project layout.
--sample-table SAMPLE_TABLE
Project sample table with sample, condition, bam, and
peaks columns.
--layout {custom,project}
Use fp-tools standard project output layout under
--outdir (default: project when --sample-table is
provided).
--sample-output-root SAMPLE_OUTPUT_ROOT
Sample output root; writes each sample under
<root>/<sample>/normalize, typically
<project>/samples.
--method {background-scale,background-zscore,none}
Normalization method (default: background-scale).
--stat STAT Background statistic used by background-scale
(default: q90). Use median, iqr, or quantiles such as
q90, q95, q97.5, or q99.
--target {median,mean}
Across-sample target statistic for background-scale
(default: median).
--chrom-sizes CHROM_SIZES
Optional chromosome sizes file for output
validation/header.
--workers WORKERS Number of input signal bigWigs to normalize
concurrently (default: all available cores, capped by
input count).
plot-aggregate¶
Plot the average cut-site signal around motif sites to inspect the shape of a
footprint. Start with corrected bigWigs from atac-correct and motif results
from match-motifs. Outputs can be static figures or an interactive HTML
report.
Example command
plot-aggregate \
--sample-table project/metadata/samples.tsv \
--motifs SPIB CEBPB \
--site-set bound \
--outdir project
Primary inputs
--sample-table— TSV withsampleandconditioncolumns; the command reads corrected tracks and motif results from{project}/samples/{sample}/.--motifs— motif names or identifiers to plot.--site-set— sites to average; the example uses sites classified as bound in each sample.--outdir— project directory containing motif results and receiving plots.
Replace SPIB CEBPB with motifs present in your motif results.
Main outputs
{project}/reports/plot_aggregate.html— default project-layout interactive aggregate report with motif-centered signal profiles.- the exact
--outputpath — static PDF/PNG/SVG or interactive HTML in custom layout. - the exact
--output-txtpath — optional per-position aggregate values. - the exact
--output_aggregated_signals,--output_aggregated_scores, and--output_aggregated_statspaths — optional source tables when requested. - the exact
--outputpath in--motif-gridmode — multipage motif-by-comparison PDF built from a review bundle.
When both signal types are available, use footprint score bigWigs for motif statistics and bias-corrected cut-site signal bigWigs for observed aggregate profiles; label the chosen signal explicitly in figure captions.
Look for central cut-site depletion relative to the flanking signal. Also
compare the number of sites and the sample-to-sample consistency. Bound-site
plots can use different sites in each sample; choose --site-set all to
inspect all matched motif sites instead.
For your own site sets, supply BED files with --TFBS. Use --regions to
restrict or compare regions of interest.
Export a grid from a review report¶
plot-aggregate \
--input-html project/reports/review_multi_comparisons/index.html \
--motif-grid \
--output project/reports/motif_aggregate_grid.pdf
Complete options
usage: plot-aggregate [-h] [--TFBS [<bed> ...]] [--signals [<bigwig> ...]]
[--match-dir [<directory> ...]] [--sample-dirs [<directory> ...]]
[--sample-table <tsv>] [--layout {custom,project}]
[--manifest <tsv>] [--input-html [<html> ...]]
[--regions [<bed> ...]] [--whitelist [<bed> ...]]
[--blacklist [<bed> ...]] [--output] [--outdir <directory>]
[--output-txt] [--output-csv] [--output_aggregated_signals]
[--output_aggregated_scores] [--multiscale-npz <npz>]
[--output-multiscale-aggregate] [--title] [--format {auto,pdf,html}]
[--flank] [--motifs [<motif> ...]] [--site-set {bound,all,unbound}]
[--top-n <int>] [--default-layout {1x1,1x2,2x2,2x3}]
[--hide-summary] [--TFBS-labels [...]] [--signal-labels [...]]
[--cond-names [<name> ...]] [--region-labels [...]]
[--control-label <label>] [--grid <rows>x<cols>] [--share-y]
[--normalize]
[--normalization {none,condition-quantile,sample-quantile}]
[--normalization-comparison-output] [--output_aggregated_stats]
[--show-replicate-sd] [--negate] [--smooth <int>] [--log-transform]
[--plot-boundaries] [--signal-on-x] [--remove-outliers <float>]
[--motif-grid] [--rows-per-page ROWS_PER_PAGE]
[--order-htmls [ORDER_HTMLS ...]] [--fill-missing-profiles]
[--recompute-missing-profiles] [--repeat-column-labels {none,row}]
[--cores <int>] [--verbosity <int>]
__________________________________________________________________________________________
fp-tools plot-aggregate
__________________________________________________________________________________________
Input / output arguments:
--TFBS [<bed> ...] TFBS sites (*required)
--signals [<bigwig> ...] Signals in bigwig format (*required)
--match-dir [<directory> ...] match-motifs output directory or directories to
use as the motif-site source
--sample-dirs [<directory> ...] Alias for --match-dir in HTML mode; sample or
differential output directories containing motif
BEDs
--sample-table <tsv> Project sample table with sample and condition
columns
--layout {custom,project} Use fp-tools standard project output layout under
--outdir (default: project when --sample-table is
provided)
--manifest <tsv> TSV with sample, signal, and match_dir/sample_dir
columns for HTML mode
--input-html [<html> ...] Existing aggregate or diff-footprints HTML
payload(s) to merge in HTML mode
--regions [<bed> ...] Regions to overlap with TFBS (optional)
--whitelist [<bed> ...] Only plot sites overlapping whitelist (optional)
--blacklist [<bed> ...] Exclude sites overlapping blacklist (optional)
--output Path to output plot (default: fp-
tools_aggregate.pdf)
--outdir <directory> Project directory used with --layout project
--output-txt Path to output file for aggregates in .txt-format
(default: None)
--output-csv Path to aggregated signal CSV output (default:
None)
--output_aggregated_signals Path to CSV file for per-base aggregated signals
(default: None)
--output_aggregated_scores Path to CSV file for aggregated footprint-score
table (default: None)
--multiscale-npz <npz> Optional call-footprints --output-multiscale-npz
sidecar to render as a scale-by-position aggregate
figure
--output-multiscale-aggregate Path for the optional multiscale aggregate figure
(default: <output stem>_multiscale.<output ext>)
Plot arguments:
--title Title of plot (default: "Aggregated signals")
--format {auto,pdf,html} Output format for --output. auto uses the output
file extension (default: auto)
--flank Flanking base pairs (+/-) to show in plot (counted
from middle of the TFBS) (default: 60)
--motifs [<motif> ...] Motif prefixes, names, or IDs to plot from
--match-dir
--site-set {bound,all,unbound} Motif-site BED set to use from --match-dir
(default: bound)
--top-n <int> Number of motifs to plot from --match-dir when
--motifs is omitted (default: 12)
--default-layout {1x1,1x2,2x2,2x3} Initial HTML subplot layout (default: 2x2)
--hide-summary Hide the TF site-count summary sidebar in HTML
mode
--TFBS-labels [ ...] Labels used for each TFBS file (default: prefix of
each --TFBS)
--signal-labels [ ...] Labels used for each signal file (default: prefix
of each --signals)
--cond-names [<name> ...] Condition names for --signals; repeated names are
averaged as replicates
--region-labels [ ...] Labels used for each regions file (default: prefix
of each --regions)
--control-label <label> Overlay each non-control signal against this
control signal label (must match one of --signal-
labels)
--grid <rows>x<cols> Explicit grid layout for subplots, e.g. 2x5 or
3x4. Panels fill in order of the input signal
files.
--share-y Share y-axis range across plots
(none/signals/sites/both). Use "--share-y signals"
if bigwig signals have similar ranges. Use "--
share-y sites" if sites per bigwig are comparable,
but bigwigs themselves aren't comparable (default:
none)
--normalize Normalize the aggregate signal(s) to be between
0-1 (default: the true range of values is shown)
--normalization {none,condition-quantile,sample-quantile}
diff-footprints-compatible quantile normalization
before aggregate plotting (default: none)
--normalization-comparison-output Optional paired raw-vs-normalized aggregate figure
--output_aggregated_stats Path to CSV file for aggregate mean/SD/stat
summaries (default: None)
--show-replicate-sd Draw replicate SD ribbons when --cond-names
contains repeated condition names
--negate Negate overlap with regions
--smooth <int> Smooth output signal by taking the mean of
<smooth> bp windows (default: 1; no smoothing)
--log-transform Log transform the signals before aggregation
--plot-boundaries Plot TFBS boundaries (Note: estimated from first
region in each --TFBS)
--signal-on-x Show signals on x-axis and TFBSs on y-axis
(default: signal is on y-axis)
--remove-outliers <float> Value between 0-1 indicating the percentile of
regions to include, e.g. 0.99 to remove the sites
with 1% highest values (default: 1)
--motif-grid Create a multi-page motif-by-comparison PDF from
one review-multi-comparisons report
--rows-per-page ROWS_PER_PAGE Motif rows per page in --motif-grid mode (default:
16)
--order-htmls [ORDER_HTMLS ...] Optional review reports used to define one shared
motif order in --motif-grid mode
--fill-missing-profiles Fill missing motif profiles from profiles embedded
elsewhere in the review report
--recompute-missing-profiles Recompute missing motif profiles from project
bigWigs and motif BEDs
--repeat-column-labels {none,row} Repeat comparison labels in every motif row
(default: none)
Run arguments:
--cores <int> Worker processes for recomputing missing motif
profiles
--verbosity <int> Level of output logging (0: silent, 1:
errors/warnings, 2: info, 3: stats, 4: debug, 5:
spam) (default: 3)
review-multi-comparisons¶
Review several diff-footprints comparisons in one place. Combine the
existing reports into a browser bundle or a single HTML file, then switch
between comparisons to explore motif statistics and aggregate profiles.
Example command
review-multi-comparisons \
--inputs project/comparisons \
--output-dir project/reports/review_multi_comparisons
Primary inputs
--inputs— report files or directories containing differential reports.--output-dir— destination for the complete static bundle.
Main outputs
{bundle} is the --output-dir:
| Path | Meaning |
|---|---|
{bundle}/index.html |
Report entry page, served together with the other bundle files. |
{bundle}/app.js, {bundle}/plot_controls.js, and {bundle}/styles.css |
Local application code, shared plot behavior, and styling. |
{bundle}/data/metadata.json |
Comparison index and payload checksums. |
{bundle}/data/reports/{comparison}.json.gz |
Compact data for one comparison. |
{bundle}/data/profiles/ |
Aggregate-profile shards loaded on demand. |
{bundle}/data/logos/ |
Motif logo assets. |
Project mode defaults to
{project}/reports/review_multi_comparisons/index.html. Keep the entire
directory together when copying or publishing a bundle.
Open the bundle¶
The bundle loads its data through a web server. To view it locally, run:
python -m http.server 8000 --bind 127.0.0.1 \
--directory project/reports/review_multi_comparisons
Open http://127.0.0.1:8000/ in your browser. Stop the
server with Ctrl+C when you finish. To open a report directly without a
server, use the standalone output below.
To choose the bundle's opening view, add --default-comparison with the two
condition or region labels, --default-aggregate-motifs with the motif names
or IDs in display order, and --default-aggregate-plots with the number of
panels. The selected labels and motifs must exist in the input reports.
Use --documentation-url to add a documentation link.
Standalone output¶
Use --output-html instead of --output-dir to create one HTML file that you
can open directly or share. It includes volcano plots, ranked motifs, logos,
and SVG exports. Aggregate controls appear when the input reports contain
profiles. Use --labels to give each input report a distinct name in the
Comparison list, especially when reports compare the same condition pair.
In the ranked-motif plot, switch between differential footprint score and
-log10(p-value) to change the ranking view. Use the legend to interpret the
color scale and direction of change. In the volcano plot, use Label TFs to label
selected TF names, motif IDs, or output prefixes, separated by commas. Select
(none) in the highlight control to clear highlighting. SVG exports retain
the selected comparison label and plot settings.
review-multi-comparisons --inputs baseline/report.html dose1/report.html dose2/report.html \
--labels Baseline "Dose 1" "Dose 2" --output-html review.html
Complete options
usage: review-multi-comparisons [-h] [--inputs INPUTS [INPUTS ...]]
[--labels [LABELS ...]]
[--output-dir OUTPUT_DIR | --output-html OUTPUT_HTML]
[--outdir OUTDIR] [--layout {custom,project}]
[--default-comparison <group1> <group2>]
[--default-aggregate-motifs <motif> [<motif> ...]]
[--default-aggregate-plots <int>]
[--documentation-url <url>]
[--fill-missing-aggregate-profiles]
[--recompute-missing-aggregate-profiles]
[--aggregate-flank AGGREGATE_FLANK]
[--cores CORES] [--title TITLE]
Combine diff-footprints reports into a static browser bundle or one self-
contained HTML file.
options:
-h, --help show this help message and exit
--inputs INPUTS [INPUTS ...]
diff-footprints HTML files or directories containing
diff_footprints_*.html files; directories are searched
recursively.
--labels [LABELS ...]
Optional labels, one per resolved input HTML.
--output-dir OUTPUT_DIR
Output directory for index.html, JavaScript, CSS, and
compact static data files.
--output-html OUTPUT_HTML
One self-contained HTML report; aggregate profiles are
optional.
--outdir OUTDIR Project directory used with --layout project.
--layout {custom,project}
Use fp-tools standard project output layout under
--outdir (default: project when only --outdir is
provided).
--default-comparison <group1> <group2>
Region or condition pair initially shown in the static
browser
--default-aggregate-motifs <motif> [<motif> ...]
Ordered motif IDs, names, or output prefixes initially
shown
--default-aggregate-plots <int>
Number of aggregate profiles initially shown (default:
4; maximum: 12)
--documentation-url <url>
Optional link back to the documentation site
--fill-missing-aggregate-profiles
Fill missing motif aggregate panels from profiles
embedded elsewhere in the combined review payload.
--recompute-missing-aggregate-profiles
Recompute still-missing motif aggregate panels from
project sample bigWigs and match-motifs BEDs.
--aggregate-flank AGGREGATE_FLANK
Flank used when recomputing missing aggregate
profiles, or 'auto' to match the existing report axis
(default: auto).
--cores CORES Worker processes for --recompute-missing-aggregate-
profiles (default: all available cores).
--title TITLE
run-yaml-workflow¶
Run one or more fp-tools commands from a saved YAML configuration. Use this to repeat a GUI run or apply the same settings to several samples or comparisons.
Example command
run-yaml-workflow --config workflow.yml --dry-run
run-yaml-workflow --config workflow.yml --run-root project/yaml_runs
Primary inputs
--config— YAML configuration exported by the GUI or written directly.--run-root— optional directory for job status and logs; it does not replace each command's analysis output directory.--dry-run— print the commands without starting the analysis.
Replace workflow.yml with your saved configuration. The first command prints
the jobs without running them; check their inputs and output paths before
running the second command. Jobs run sequentially, and each job's output appears
live in the terminal as well as in its saved logs. Commands that support a core
budget use all available cores when cores is omitted or set to null. An
explicit number is retained. A job-level cores: null also overrides a limit
in defaults, allowing the execution machine to choose its available cores.
Main outputs
- The analysis files documented for each command named in the YAML.
- Standard output containing the expanded command lines when
--dry-runis used. {run_root}/{job_id}/config.ymlandcommand.txt— saved settings and exact command for each job.{run_root}/{job_id}/status.json,stdout.log, andstderr.log— completion state and captured command output.{run_root}/batch_index.tsv— one-row-per-job batch status index.
Relative paths in the YAML are interpreted from the directory where you launch
the command. Run from the same directory each time, or use absolute paths. If
neither --run-root nor a YAML run_root is supplied, logs go to a new
fp-tools-batch-{timestamp} folder in that directory.
Complete options
usage: run-yaml-workflow [-h] --config CONFIG [--run-root RUN_ROOT]
[--only [ONLY ...]] [--dry-run] [--list-jobs]
[--fail-fast]
Run fp-tools jobs from a YAML config file.
options:
-h, --help show this help message and exit
--config CONFIG Path to YAML config.
--run-root RUN_ROOT Optional directory for run metadata/logs.
--only [ONLY ...] Optional tool filter, e.g. diff-footprints.
--dry-run Print expanded commands without running.
--list-jobs List expanded jobs and exit.
--fail-fast Stop at first failed job.
fp-tools-gui¶
Launch the browser interface for configuring and running fp-tools commands. The Windows and Apple Silicon desktop downloads present the same interface in a native fp-tools application window.
Bulk GUI workflows start from coordinate-sorted BAM/BAI files and matching peak BED files. The GUI does not perform FASTQ-to-BAM preprocessing. Missing inputs and unsupported options are reported before a run starts.
The GUI is available through the Python package, the complete container, and the self-contained desktop downloads on the release page.
Example command
fp-tools-gui --host 127.0.0.1 --port 8891 --run-dir project/gui_runs
Primary inputs
--host— interface on which the GUI listens (default:127.0.0.1).--port— fixed browser port.--run-dir— directory for GUI-managed configurations and runs.
Start an analysis¶
- Select the workflow or command from the sidebar.
- Enter your input paths and output directory, then select Update page config.
- Review the displayed command and resolve any validation errors before starting the run.
- Open Run History to check progress, read logs, and find the output files.
Use the Config page to save the settings as YAML or load a previous run's configuration.
Main outputs
{run_dir}/{timestamp}_{label}/config.yml— saved YAML settings for the run.{run_dir}/{timestamp}_{label}/status.json,launcher_stdout.log, andlauncher_stderr.log— overall run status and logs.{run_dir}/{timestamp}_{label}/{job_id}/status.json,command.txt,stdout.log, andstderr.log— each job's status, exact command, and analysis logs.- The analysis files documented by the selected command, written to the output directory you chose.
A saved YAML can also be run from the command line with
run-yaml-workflow --config {run_dir}/{timestamp}_{label}/config.yml.
Local computer¶
Open the desktop executable to use the native application window. When using
the Python package, run fp-tools-gui; a browser opens after the server is
ready. If it does not, open the local URL printed in the terminal.
Remote Linux server¶
Start fp-tools on the server. --no-browser prevents it from opening a browser
on the server, and the default host setting limits access to that server:
fp-tools-gui --no-browser --port 8891
On your computer, create an SSH tunnel and keep that terminal open:
ssh -N -L 8891:127.0.0.1:8891 USER@SERVER
Open http://127.0.0.1:8891 on your computer to use the server's GUI.
Complete options
usage: fp-tools-gui [-h] [--host HOST] [--port PORT] [--run-dir RUN_DIR]
[--no-browser]
Launch the fp-tools browser interface.
options:
-h, --help show this help message and exit
--host HOST Bind address (default: 127.0.0.1).
--port PORT Optional fixed port (default: first free port from 8891).
--run-dir RUN_DIR Directory for GUI-managed runs.
--no-browser Do not open a local browser automatically.
fp-tools-runtime¶
Check or install the external programs used by prepare-atac and
discover-motifs. fp-tools manages these programs separately from your other
software. Linux supports read preparation and motif discovery; macOS and
Windows support motif discovery.
Example command
fp-tools-runtime status
Primary inputs
The status action takes no input files. To install the programs used for motif
discovery before your first run:
fp-tools-runtime install meme
On Linux, use fp-tools-runtime install core for the default prepare-atac
workflow, or fp-tools-runtime install homer for its homer-atac profile.
Main outputs
The command reports each runtime component, platform, installation state, and
cache location. Commands that need a managed component install it on first use,
so manual installation is optional. If an installation is damaged, run
fp-tools-runtime repair meme (or substitute the affected component), then
check fp-tools-runtime status again.
Complete options
usage: fp-tools-runtime [-h] {status,install,repair} ...
Inspect, install, or repair the managed fp-tools runtime.
positional arguments:
{status,install,repair}
status Report managed runtime availability and installation
state.
install Install a runtime component.
repair Repair a runtime component.
options:
-h, --help show this help message and exit
discover-motifs¶
Find recurring DNA sequence patterns in candidate footprint regions, without starting from a known motif list. Provide candidate intervals and a reference genome, or use sequences you have already extracted into a FASTA file.
Example command
discover-motifs --candidates project/samples/sample/footprints/sample_candidate_footprints.bed --genome hg38.fa.gz \
--flank 75 --method streme --known-motif-db jaspar2026_vertebrates --outdir project/de_novo/sample --execute
Primary inputs
--candidates— candidate-footprint BED intervals.--genome— reference FASTA matching the candidate coordinates; required with--candidates.--flank— bases included on each side of a candidate center.--method— discovery method; the example uses STREME.--known-motif-db— optional known-motif database for Tomtom matching.--outdir— directory for candidate FASTA files and discovery results.--execute— run discovery immediately using the managed MEME Suite runtime.
In the example, --flank 75 extracts up to 75 bases on each side of each
candidate center.
Without --execute, the command writes the FASTA and a command script for
inspection. It does not run discovery or create result summaries.
Main outputs
{outdir} is the selected discovery directory:
| Path | Meaning |
|---|---|
{outdir}/candidate_sequences.fa |
Reference sequences extracted around candidate footprint intervals. |
{outdir}/run_motif_discovery.sh |
Commands for discovery, optional known-motif matching, and summary reports. |
{outdir}/streme/streme.txt, meme/meme.txt, or dreme/dreme.txt |
Discovered motif models from the selected method when --execute is used. |
{outdir}/tomtom/tomtom.tsv |
Optional similarity matches to the selected known-motif database. |
{outdir}/motif_summary.tsv and motif_summary.html |
Motif tables written after successful execution. |
Choose your sequence input¶
Create candidate intervals with call-footprints --output-bed. If you already
have a sequence FASTA, pass it with --fasta instead of --candidates and
--genome. In that case the command uses your existing sequences directly.
Complete options
usage: discover-motifs [-h] (--fasta FASTA | --candidates CANDIDATES)
[--genome GENOME] [--flank FLANK] --outdir OUTDIR
[--script SCRIPT] [--method {meme,dreme,streme}]
[--known-motifs KNOWN_MOTIFS [KNOWN_MOTIFS ...]]
[--known-motif-db KNOWN_MOTIF_DB] [--list-motif-dbs]
[--extra-args ...] [--execute]
[--runtime {auto,managed,system,container}]
Prepare or run a de novo motif discovery command plan.
options:
-h, --help show this help message and exit
--fasta FASTA Existing candidate FASTA.
--candidates CANDIDATES
Candidate BED from call-footprints --output-bed or
another BED-like source.
--genome GENOME Genome FASTA, required when --candidates is used.
--flank FLANK If >0 with --candidates, export +/- flank bp around
each candidate center.
--outdir OUTDIR External motif discovery output directory.
--script SCRIPT Output shell script path. Defaults to
<outdir>/run_motif_discovery.sh.
--method {meme,dreme,streme}
--known-motifs KNOWN_MOTIFS [KNOWN_MOTIFS ...]
Optional known motif database file(s) for Tomtom
comparison.
--known-motif-db KNOWN_MOTIF_DB
Optional built-in motif database for Tomtom
comparison.
--list-motif-dbs List available built-in motif databases and exit.
--extra-args ... Additional arguments appended to MEME/DREME/STREME.
--execute Run the generated script immediately.
--runtime {auto,managed,system,container}
External-tool runtime: auto/managed provisions the
pinned fp-tools runtime, system uses PATH, and
container uses the complete image (default: auto).
summarize-motifs¶
Turn completed motif-discovery results into a table of motif sequences, significance values, and optional matches to known motifs. Add an HTML report to review the results in a browser.
Example command
summarize-motifs --meme-txt project/de_novo/sample/streme/streme.txt \
--tomtom-tsv project/de_novo/sample/tomtom/tomtom.tsv \
--out-tsv project/de_novo/sample/motif_summary.tsv --out-html project/de_novo/sample/motif_summary.html
Primary inputs
--meme-txt— discovery output such asmeme.txt,streme.txt, ordreme.txt.--tomtom-tsv— optional Tomtom known-motif matches.--out-tsv— compact output table for discovered motifs and matches.--out-html— optional browser report.
Omit --tomtom-tsv if you did not run known-motif matching. This command reads
existing results; it does not rerun motif discovery.
Main outputs
- The
--out-tsvfile — tab-separated discovered motif IDs, consensus sequences, significance values, and known-database matches when available. - The
--out-htmlfile — optional portable HTML table containing the same summary and consensus-sequence logos when available.
Open motif_summary.html in a browser or import motif_summary.tsv into a
spreadsheet. Known-motif matches help identify candidate TF families; they do
not establish which TF is bound.
Complete options
usage: summarize-motifs [-h] [--meme-txt MEME_TXT] [--tomtom-tsv TOMTOM_TSV]
--out-tsv OUT_TSV [--out-html OUT_HTML]
[--title TITLE]
Summarize MEME/Tomtom outputs into TSV and HTML reports.
options:
-h, --help show this help message and exit
--meme-txt MEME_TXT MEME text output, usually meme.txt.
--tomtom-tsv TOMTOM_TSV
Tomtom TSV output, usually tomtom.tsv.
--out-tsv OUT_TSV Output motif summary TSV.
--out-html OUT_HTML Optional output HTML report.
--title TITLE
pseudobulk-fragments¶
Combine single-cell ATAC-seq fragments into groups such as cell types or donor–cell-type pairs. Each group becomes a pseudobulk sample for downstream analysis.
Example command
pseudobulk-fragments --fragments pbmc_fragments.tsv.gz --annotations cell_annotations.tsv --group-by cell_type \
--genome-sizes hg38.chrom.sizes --write-cutsite-bigwigs --outdir project/pseudobulk/fragments
Primary inputs
--fragments— TSV or TSV.GZ with chromosome, start, end, and barcode in its first four columns; a fifth column can contain fragment counts.--annotations— TSV or CSV withbarcodeand the column named by--group-by(for example,cell_type).--group-by— annotation column used to define pseudobulk groups; usedonor,cell_typeto keep donors separate within each cell type.--genome-sizes— two-column chromosome-name and length file matching the fragments; required for the example's signal tracks.--write-cutsite-bigwigs— write cut-site bigWigs for retained groups.--outdir— directory for grouped fragments, tracks, and QC outputs.
Main outputs
{group} is the annotation value converted to a filename-safe name. Paths below
are relative to {outdir}:
| Path | Meaning |
|---|---|
{group}.fragments.tsv or {group}.fragments.tsv.gz |
Fragments assigned to the group; compressed/indexed form is controlled by the command options. |
{group}.fragments.tsv.gz.tbi |
Optional Tabix index for random genomic access. |
{group}.cutsites.cpm.bw |
Optional CPM-normalized cut-site signal bigWig written by --write-cutsite-bigwigs. |
{group}.pseudo_pairs.sorted.bam and .bai |
Optional pseudo-paired alignment written with --write-pseudo-bams; use atac-correct --read_shift 0 0 on these files. |
pseudobulk_manifest.tsv |
Per-group paths, cell/fragment counts, and filter status. |
fp_tools_manifest.yml |
Machine-readable run settings and retained groups. |
pseudobulk_downstream_commands.sh |
Optional generated downstream command examples. |
Match cell barcodes¶
By default, barcode matching ignores a trailing suffix such as -1. Add
--no-strip-barcode-suffix when suffixes distinguish cells in your dataset.
Use --barcode-column if your annotation barcode column has a different name.
Complete options
usage: pseudobulk-fragments [-h] --fragments FRAGMENTS --annotations
ANNOTATIONS --group-by GROUP_BY
[--barcode-column BARCODE_COLUMN]
[--no-strip-barcode-suffix]
[--include-chroms INCLUDE_CHROMS]
[--exclude-chroms EXCLUDE_CHROMS]
[--min-cells MIN_CELLS]
[--min-fragments MIN_FRAGMENTS] --outdir OUTDIR
[--compress-output] [--index-output]
[--write-cutsite-bigwigs] [--write-pseudo-bams]
[--no-cpm-normalize] [--write-downstream-commands]
[--genome-sizes GENOME_SIZES] [--cores CORES]
Group single-cell ATAC fragments into pseudobulk fragment files.
options:
-h, --help show this help message and exit
--fragments FRAGMENTS
10x-style fragments TSV/TSV.GZ with barcode in column
4.
--annotations ANNOTATIONS
Cell annotation TSV or CSV.
--group-by GROUP_BY Comma-separated annotation columns to group by, e.g.
donor,cell_type.
--barcode-column BARCODE_COLUMN
Annotation barcode column (default: barcode).
--no-strip-barcode-suffix
Require exact barcode matches instead of matching
AAAC-1 to AAAC.
--include-chroms INCLUDE_CHROMS
Comma-separated chromosomes to keep, e.g.
chr1,chr2,chrX.
--exclude-chroms EXCLUDE_CHROMS
Comma-separated chromosomes to skip, e.g. chrM,chrY.
--min-cells MIN_CELLS
Minimum cells for passes_filters (default: 1).
--min-fragments MIN_FRAGMENTS
Minimum fragments for passes_filters (default: 1).
--outdir OUTDIR Output directory.
--compress-output Write grouped fragments as .tsv.gz files.
--index-output BGZF-compress and tabix-index grouped fragments for
random access.
--write-cutsite-bigwigs
Write one sparse cut-site bigWig per kept pseudobulk
group.
--write-pseudo-bams Write sorted pseudo-paired BAMs for kept groups; use
atac-correct --read_shift 0 0 on these BAMs.
--no-cpm-normalize Write raw cut counts instead of CPM-normalized bigWig
values.
--write-downstream-commands
Write a shell script for BED/BAM/bigWig generation
from kept pseudobulk groups.
--genome-sizes GENOME_SIZES
Two-column chromosome sizes file used by generated
bedtools/UCSC commands and cut-site bigWigs.
--cores CORES Cores for compression, bigWig writing, and generated
samtools commands (default: all available cores).
find-signature-fp¶
Score selected motif sites in individual cells and plot the results on a UMAP and in heatmaps. The command pools signal from nearby cells to reduce sparsity, so the scores describe relative footprint signatures rather than independent binding calls for every cell.
Example command
find-signature-fp --annotations cell_annotations.tsv --fragments pbmc_fragments.tsv.gz --h5ad genomic_bin_counts.h5ad \
--all-motif-diff-dir project/pseudobulk/diff_footprints \
--all-motif-results project/pseudobulk/diff_footprints/pseudobulk_diff_footprints_results.txt \
--outdir project/pseudobulk/signature_fp
Primary inputs
--annotations— TSV or CSV with requiredbarcode,cell_type,snap_cell_type,umap_1, andumap_2columns. The GUI and command check these columns before analysis starts.--fragments— single-cell fragment file with chromosome, start, end, and barcode in its first four columns; the command creates a missing Tabix index by default.--h5ad— AnnData file with matching cell names, genomic-bin counts, and a booleanselectedcolumn invar. Bin names must usechromosome:start-end; the default bin size is 500 bases.--all-motif-diff-dir— completeddiff-footprintsdirectory containing motif-site BED files.--all-motif-results— motif-level results from that directory; supply it together with--all-motif-diff-dir.--outdir— directory for per-cell scores, heatmaps, and UMAP figures.
Use cell_type for broad labels and snap_cell_type for detailed labels; they
can be the same if you have one annotation level. AnnData cell names must match
the annotation barcodes. Smoothing uses obsm['X_spectral'] if available,
otherwise the annotation UMAP coordinates. The companion activity scores need
the genomic-bin counts, so an embedding-only AnnData file is not sufficient.
Main outputs
Under {outdir} the default names include:
| Path | Meaning |
|---|---|
knn_footprint_signature_scores.tsv |
Per-cell KNN-smoothed footprint protection scores for selected TFs. |
knn_footprint_orientation_summary.tsv |
Direction/orientation checks used to make marker scores comparable. |
chromvar_like_motif_activity_scores.tsv |
Companion accessibility-derived motif activity scores. |
knn_footprint_signature_umap.svg |
Per-marker footprint-signature UMAP panels. |
per_cell_footprint_signature_heatmap.svg |
Selected-marker per-cell heatmap. |
single_cell_footprinting_summary.svg |
Combined heatmap and representative UMAP summary. |
all_motif_per_cell_footprint_signature_heatmap.tsv |
Optional all-motif score matrix and metadata when all-motif inputs are supplied. |
Open the summary SVG first, then use the score tables to inspect individual cells. Additional top-motif plots and all-TF review PDFs are produced when all-motif inputs are supplied.
Choose marker TFs¶
Choose TFs with --markers TF1,TF2. The default markers are
STAT6,FOSB,CEBPA,IRF8,RELA,ZNF683,NR4A1,SMAD3; each selected TF must have motif
sites in your inputs. In the GUI, enter one marker per line. YAML accepts either
a list such as markers: [STAT6, CEBPA, ZNF683] or the comma-separated form
markers: STAT6,CEBPA,ZNF683. Saved GUI configurations run through
run-yaml-workflow without conversion.
For selected-marker reports only, use --tf-site-dir
and omit --all-motif-diff-dir and --all-motif-results from the example.
This alternative directory must contain files named {TF}.motif_hits.bed or
{TF}.motif_peaks.bed.
Complete options
usage: find-signature-fp [-h] --annotations ANNOTATIONS --fragments FRAGMENTS
--h5ad H5AD [--tf-site-dir TF_SITE_DIR] --outdir
OUTDIR [--markers MARKERS]
[--max-sites-per-tf MAX_SITES_PER_TF] [--knn KNN]
[--flank FLANK]
[--center-half-width CENTER_HALF_WIDTH]
[--flank-inner FLANK_INNER]
[--flank-outer FLANK_OUTER] [--bin-size BIN_SIZE]
[--marker-groups MARKER_GROUPS]
[--all-motif-diff-dir ALL_MOTIF_DIFF_DIR]
[--all-motif-results ALL_MOTIF_RESULTS]
[--all-motif-score-table ALL_MOTIF_SCORE_TABLE]
[--marker-score-table MARKER_SCORE_TABLE]
[--all-motif-batch-size ALL_MOTIF_BATCH_SIZE]
[--max-sites-per-motif MAX_SITES_PER_MOTIF]
[--max-motifs MAX_MOTIFS]
[--top-motif-signatures-per-cell-type TOP_MOTIF_SIGNATURES_PER_CELL_TYPE]
[--top-motif-min-specificity TOP_MOTIF_MIN_SPECIFICITY]
[--summary-output-prefix SUMMARY_OUTPUT_PREFIX]
[--all-tf-review-prefix ALL_TF_REVIEW_PREFIX]
[--all-tf-review-panels-per-page ALL_TF_REVIEW_PANELS_PER_PAGE]
[--skip-all-tf-review-pdfs]
[--no-create-fragment-index]
Generate per-cell footprint-signature heatmaps and UMAP reports.
options:
-h, --help show this help message and exit
--annotations ANNOTATIONS
Cell annotation TSV/CSV requiring barcode, cell_type,
snap_cell_type, umap_1, and umap_2 columns.
--fragments FRAGMENTS
10x-style fragments TSV/TSV.GZ used to count cut sites
around motif centers.
--h5ad H5AD AnnData with matching cell barcodes, genomic-bin
counts, and boolean var['selected']; optional
embeddings support nearest-neighbor smoothing.
--tf-site-dir TF_SITE_DIR
Optional directory containing marker motif-site BED
files named by TF. When omitted, marker sites are
taken from --all-motif-diff-dir and --all-motif-
results.
--outdir OUTDIR Output directory for signature score tables, heatmaps,
and UMAP reports.
--markers MARKERS Comma-separated marker TFs to score and plot (default:
STAT6,FOSB,CEBPA,IRF8,RELA,ZNF683,NR4A1,SMAD3).
--max-sites-per-tf MAX_SITES_PER_TF
Maximum marker motif sites per TF for selected-marker
UMAP scoring (default: 1500).
--knn KNN Number of nearest neighbors used to smooth per-cell
cut-site profiles (default: 75).
--flank FLANK Motif-centered half-window in bp for fragment counting
(default: 100).
--center-half-width CENTER_HALF_WIDTH
Half-width in bp of the protected center window
(default: 10).
--flank-inner FLANK_INNER
Inner flank distance from motif center in bp (default:
25).
--flank-outer FLANK_OUTER
Outer flank distance from motif center in bp (default:
100).
--bin-size BIN_SIZE Bin size for the companion chromVAR-like motif
activity score (default: 500).
--marker-groups MARKER_GROUPS
Comma-separated TF:cell_type pairs used to orient KNN
marker scores for UMAP review.
--all-motif-diff-dir ALL_MOTIF_DIFF_DIR
Optional differential-footprint output directory
containing */beds/*_all.bed files for all-motif per-
cell heatmap scoring.
--all-motif-results ALL_MOTIF_RESULTS
Differential-footprint results table used to order and
annotate all-motif heatmap rows.
--all-motif-score-table ALL_MOTIF_SCORE_TABLE
Existing all-motif per-cell heatmap TSV to redraw
all/top heatmaps without rescoring fragments.
--marker-score-table MARKER_SCORE_TABLE
Existing KNN marker score table used to orient
selected marker rows in top heatmaps and summary
UMAPs.
--all-motif-batch-size ALL_MOTIF_BATCH_SIZE
Number of motif signatures to score per batch for the
all-motif heatmap.
--max-sites-per-motif MAX_SITES_PER_MOTIF
Maximum motif instances per motif for all-motif
heatmap scoring; use 0 for all sites.
--max-motifs MAX_MOTIFS
Optional all-motif smoke-test limit.
--top-motif-signatures-per-cell-type TOP_MOTIF_SIGNATURES_PER_CELL_TYPE
Top cell-type-specific all-motif signatures to keep
per broad cell type (default: 40).
--top-motif-min-specificity TOP_MOTIF_MIN_SPECIFICITY
Minimum dominant-vs-next cell-type mean z-score
difference for top all-motif heatmap rows (default:
0.5).
--summary-output-prefix SUMMARY_OUTPUT_PREFIX
Output prefix for the combined heatmap and UMAP
summary SVG when all-motif heatmap data are available.
--all-tf-review-prefix ALL_TF_REVIEW_PREFIX
Output prefix for three multi-page all-TF signature
review PDFs grouped by dominant broad cell type.
--all-tf-review-panels-per-page ALL_TF_REVIEW_PANELS_PER_PAGE
Number of TF signature UMAP panels per all-TF review
PDF page (default: 12).
--skip-all-tf-review-pdfs
Do not write the three all-TF signature review PDFs.
--no-create-fragment-index
Do not create a tabix index for the fragment file when
it is missing.
sc-footprinting¶
Analyze single-cell ATAC-seq by combining cells into pseudobulk groups, calling footprints and motif differences between groups, then mapping selected footprint signatures back to individual cells.
Example command
sc-footprinting --fragments pbmc_fragments.tsv.gz --annotations cell_annotations.tsv --h5ad genomic_bin_counts.h5ad \
--group-by cell_type --genome-sizes hg38.chrom.sizes --genome hg38.fa.gz --peaks merged_peaks.bed \
--motif-db jaspar2026_vertebrates --outdir project/pseudobulk
For custom motifs alone, replace --motif-db jaspar2026_vertebrates with
--motifs custom.jaspar. Supply both options only to combine the motif sets.
Primary inputs
The workflow uses all available cores by default. Grouping progress and child
command output appear in the terminal; child output is also saved in the run's
logs directory. In the graphical user interface (GUI), leave the Cores field
blank for automatic selection, or enter a limit.
--fragments— TSV or TSV.GZ with chromosome, start, end, and cell barcode in its first four columns.--annotations— cell annotation TSV or CSV with requiredbarcode,cell_type,snap_cell_type,umap_1, andumap_2, plus any additional grouping columns. The GUI and command check these columns before analysis starts.--h5ad— AnnData file containing the same cells and genomic-bin counts used for the companion motif-activity scores. See the requirements below.--group-by— annotation column used to define pseudobulk groups.--genome-sizes— two-column chromosome-name and length file used to write grouped signal tracks.--genome— reference genome FASTA matching the fragments and peak coordinates.--peaks— accessible-region BED file.--motif-db— optional built-in motif database name.--motifs— optional custom motif files. Custom motifs alone use only those files; explicitly supply--motif-dbas well to combine them. When neither option is supplied, the command usesjaspar2026_vertebrates.--outdir— directory for pseudobulk tracks, motif results, and reports.
The per-cell reports need annotation barcodes that match the AnnData cell names.
Use cell_type for the broad cell labels and snap_cell_type for detailed
labels; these can be the same if you have only one annotation level. UMAP
coordinates go in umap_1 and umap_2.
The AnnData file needs genomic-bin names such as chr1:0-500, a count matrix,
and a boolean selected column in its feature annotations (var). The
companion activity calculation uses 500-base bins. Nearest-neighbor smoothing
uses obsm['X_spectral'] if present, otherwise the annotation UMAP coordinates.
An embedding-only AnnData file is not sufficient.
Main outputs
Paths below are relative to --outdir; {group} is a retained pseudobulk group:
| Path | Meaning |
|---|---|
pseudobulk/{group}.fragments.tsv.gz and .tbi |
Indexed fragments for each retained cell group. |
pseudobulk/{group}.cutsites.cpm.bw |
Group cut-site signal bigWig. |
pseudobulk/{group}.pseudo_pairs.sorted.bam and .bai |
Pseudo-paired alignment used for bias correction. |
atacorrect/{group}/{group}_corrected.bw |
Bias-corrected cut-site signal per group. |
footprints/{group}_footprints.bw |
Footprint score signal per group. |
diff_footprints/pseudobulk_diff_footprints_results.txt |
Optional motif-level group comparison results. |
plots/single_cell_footprinting/ |
Per-cell score tables, heatmaps, and UMAP figures from find-signature-fp. |
pseudobulk_footprint_manifest.tsv |
Group paths and workflow completion state. |
pseudobulk_footprint_commands.sh |
Exact generated commands for reproducibility. |
logs/{stage}.stdout.log and {stage}.stderr.log |
Captured output for each stage. |
Check the manifest for completed groups, then review the motif comparison and
per-cell heatmaps. Per-cell scores combine signal from neighboring cells to
reduce sparsity; interpret them as relative signatures, not independent
binding calls for every cell. Use --single-cell-signature-markers to choose
the TFs shown in the marker reports.
Complete options
usage: sc-footprinting [-h] --fragments FRAGMENTS --annotations ANNOTATIONS
--group-by GROUP_BY --outdir OUTDIR
[--genome-sizes GENOME_SIZES] --genome GENOME --peaks
PEAKS [--blacklist BLACKLIST]
[--barcode-column BARCODE_COLUMN]
[--no-strip-barcode-suffix]
[--include-chroms INCLUDE_CHROMS]
[--exclude-chroms EXCLUDE_CHROMS] [--groups GROUPS]
[--min-cells MIN_CELLS] [--min-fragments MIN_FRAGMENTS]
[--no-cpm-normalize] [--top-n TOP_N]
[--read-shift FWD REV] [--motifs [MOTIFS ...]]
[--motif-db MOTIF_DB] [--list-motif-dbs]
[--peak-header PEAK_HEADER] [--diff-prefix DIFF_PREFIX]
[--diff-normalization {condition-quantile,sample-quantile,none}]
[--diff-plot-aggregate {sig,all,top,off}]
[--skip-excel | --no-skip-excel]
[--tf-site-dir TF_SITE_DIR]
[--site-summary SITE_SUMMARY] [--tfs TFS]
[--plot-flank PLOT_FLANK] [--plot-script PLOT_SCRIPT]
--h5ad SINGLE_CELL_SIGNATURE_H5AD
[--single-cell-signature-outdir SINGLE_CELL_SIGNATURE_OUTDIR]
[--single-cell-signature-markers SINGLE_CELL_SIGNATURE_MARKERS]
[--single-cell-signature-fig-prefix SINGLE_CELL_SIGNATURE_FIG_PREFIX]
[--single-cell-signature-all-motif-score-table SINGLE_CELL_SIGNATURE_ALL_MOTIF_SCORE_TABLE]
[--single-cell-signature-marker-score-table SINGLE_CELL_SIGNATURE_MARKER_SCORE_TABLE]
[--single-cell-signature-top-per-cell-type SINGLE_CELL_SIGNATURE_TOP_PER_CELL_TYPE]
[--single-cell-signature-top-min-specificity SINGLE_CELL_SIGNATURE_TOP_MIN_SPECIFICITY]
[--single-cell-signature-knn SINGLE_CELL_SIGNATURE_KNN]
[--single-cell-signature-max-sites-per-motif SINGLE_CELL_SIGNATURE_MAX_SITES_PER_MOTIF]
[--single-cell-signature-max-motifs SINGLE_CELL_SIGNATURE_MAX_MOTIFS]
[--cores CORES] [--resume] [--force] [--dry-run]
[--fail-fast]
Run the complete pseudobulk and per-cell footprint workflow from single-cell
fragments.
options:
-h, --help show this help message and exit
--fragments FRAGMENTS
10x-style fragments TSV/TSV.GZ with barcode in column
4.
--annotations ANNOTATIONS
Cell annotation TSV or CSV.
--group-by GROUP_BY Comma-separated annotation columns to group by.
--outdir OUTDIR Output directory for the full pseudobulk footprint
workflow.
--genome-sizes GENOME_SIZES
Two-column chromosome sizes file used for fragment-
derived cut-site bigWigs.
--genome GENOME Genome FASTA for atac-correct.
--peaks PEAKS Peak BED used for atac-correct and footprint scoring.
--blacklist BLACKLIST
Optional blacklist BED for atac-correct.
--barcode-column BARCODE_COLUMN
Annotation barcode column (default: barcode).
--no-strip-barcode-suffix
Require exact barcode matches instead of matching
AAAC-1 to AAAC.
--include-chroms INCLUDE_CHROMS
Comma-separated chromosomes to keep.
--exclude-chroms EXCLUDE_CHROMS
Comma-separated chromosomes to skip.
--groups GROUPS Comma-separated pseudobulk groups to process after
grouping; default processes all retained groups.
--min-cells MIN_CELLS
Minimum cells for passes_filters (default: 1).
--min-fragments MIN_FRAGMENTS
Minimum fragments/reads for passes_filters (default:
1).
--no-cpm-normalize Write raw cut counts instead of CPM-normalized cut-
site bigWigs for fragment input.
--top-n TOP_N Optional top N candidate footprints per group.
--read-shift FWD REV Override the atac-correct read shift for fragment cut
sites (default: 0 0).
--motifs [MOTIFS ...]
Optional motif file(s); when provided, run motif-aware
diff-footprints on pseudobulk footprint tracks.
--motif-db MOTIF_DB Built-in motif database; defaults to
jaspar2026_vertebrates only when neither --motif-db
nor --motifs is supplied. Explicitly supply both to
combine them.
--list-motif-dbs List available built-in motif databases and exit.
--peak-header PEAK_HEADER
Optional peak-header file passed to diff-footprints.
--diff-prefix DIFF_PREFIX
Prefix for optional motif-aware diff-footprints
outputs.
--diff-normalization {condition-quantile,sample-quantile,none}
Normalization mode for optional motif-aware diff-
footprints outputs (default: none).
--diff-plot-aggregate {sig,all,top,off}
Aggregate plot selection for optional motif-aware
diff-footprints HTML/PDF outputs.
--skip-excel, --no-skip-excel
Skip Excel files for optional diff-footprints outputs
(default: on).
--tf-site-dir TF_SITE_DIR
Optional motif-centered BED directory to plot
corrected footprint aggregates.
--site-summary SITE_SUMMARY
Optional motif-centered site summary TSV for plotting.
--tfs TFS Comma-separated TFs or 'auto' for plotting (default:
auto).
--plot-flank PLOT_FLANK
Flank for optional aggregate plots (default: 100).
--plot-script PLOT_SCRIPT
Plotting script path for optional aggregate plots.
--h5ad SINGLE_CELL_SIGNATURE_H5AD, --single-cell-signature-h5ad SINGLE_CELL_SIGNATURE_H5AD
AnnData with matching cell barcodes, genomic-bin
counts, and boolean var['selected']; an embedding
alone is insufficient.
--single-cell-signature-outdir SINGLE_CELL_SIGNATURE_OUTDIR
Output directory for optional per-cell signature
reports (default:
<outdir>/plots/single_cell_footprinting).
--single-cell-signature-markers SINGLE_CELL_SIGNATURE_MARKERS
Comma-separated marker TFs for optional per-cell
signature UMAPs (default:
STAT6,FOSB,CEBPA,IRF8,RELA,ZNF683,NR4A1,SMAD3).
--single-cell-signature-fig-prefix SINGLE_CELL_SIGNATURE_FIG_PREFIX
Output prefix for the combined single-cell footprint-
signature SVG (default: single_cell_footprinting).
--single-cell-signature-all-motif-score-table SINGLE_CELL_SIGNATURE_ALL_MOTIF_SCORE_TABLE
Existing all-motif per-cell signature TSV; skips
rescoring all motif sites for the signature heatmap.
--single-cell-signature-marker-score-table SINGLE_CELL_SIGNATURE_MARKER_SCORE_TABLE
Existing KNN marker score TSV used for marker rows and
UMAP plots.
--single-cell-signature-top-per-cell-type SINGLE_CELL_SIGNATURE_TOP_PER_CELL_TYPE
Top all-motif signatures to keep per cell type in the
signature heatmap (default: 40).
--single-cell-signature-top-min-specificity SINGLE_CELL_SIGNATURE_TOP_MIN_SPECIFICITY
Minimum dominant-vs-next cell-type z-score difference
for top heatmap rows (default: 0.5).
--single-cell-signature-knn SINGLE_CELL_SIGNATURE_KNN
KNN size for optional per-cell footprint-signature
smoothing (default: 75).
--single-cell-signature-max-sites-per-motif SINGLE_CELL_SIGNATURE_MAX_SITES_PER_MOTIF
Maximum motif instances per motif for optional all-
motif per-cell heatmap scoring; use 0 for all sites
(default: 200).
--single-cell-signature-max-motifs SINGLE_CELL_SIGNATURE_MAX_MOTIFS
Optional smoke-test limit for all-motif per-cell
heatmap scoring.
--cores CORES Optional core limit for grouping, correction, and
footprint scoring (default: all available cores).
--resume Skip atac-correct/call-footprints steps whose expected
outputs already exist.
--force Run atac-correct/call-footprints even if outputs
already exist.
--dry-run Write manifests and commands without running atac-
correct, call-footprints, motif detection, or plots.
--fail-fast Stop after the first failed group command.