Skip to content

Single-cell ATAC-seq workflow

sc-footprinting groups fragments into pseudobulk samples, runs the footprint analysis, and calculates per-cell footprint signatures. A pseudobulk sample combines fragments from cells that share an annotation, such as a cell type.

Main commands

Prepare your inputs

For your first run, use the complete PBMC example below. The general command and annotation rows in this section are templates for your own data. Run them from a folder containing the named input files.

Input Required content
fragments.tsv.gz Fragment records with chromosome, start, end, and barcode in the first four columns.
cell_annotations.tsv Cell barcodes, group labels, and UMAP coordinates, as shown below.
genomic_bin_counts.h5ad AnnData with genomic-bin counts and matching cell names; an embedding alone is insufficient.
hg38.chrom.sizes Two columns: chromosome name and length.
hg38.fa.gz Reference FASTA matching the fragments and peaks.
merged_peaks.bed Accessible regions in the same assembly and chromosome naming scheme.

The annotation table needs these columns:

barcode cell_type   snap_cell_type  umap_1  umap_2
AAACGAAAGAAACGCC-1  B_cell  B_cell  1.2 -0.8
AAACGAAAGAAAGCAG-1  T_cell  T_cell  -2.1    0.4

Use your actual barcodes and UMAP coordinates. cell_type contains broad labels and snap_cell_type contains detailed labels; they can be identical if you use one annotation level. --group-by cell_type combines cells with the same cell_type value.

The AnnData cell names must match the annotation barcodes. Its features must be genomic bins such as chr1:0-500, with a boolean var['selected'] column and a count matrix. The companion motif-activity calculation uses 500-base bins. The workflow uses obsm['X_spectral'] for nearest-neighbor smoothing when available, otherwise it uses the annotation UMAP coordinates.

Run the workflow

sc-footprinting \
  --fragments fragments.tsv.gz \
  --annotations cell_annotations.tsv \
  --h5ad genomic_bin_counts.h5ad \
  --group-by cell_type \
  --genome-sizes hg38.chrom.sizes \
  --genome hg38.fa.gz \
  --peaks merged_peaks.bed \
  --motif-db jaspar2026_vertebrates \
  --outdir project/pseudobulk

Review the results

Check project/pseudobulk/pseudobulk_footprint_manifest.tsv for completed groups. Per-cell score tables, heatmaps, and UMAPs are written under project/pseudobulk/plots/single_cell_footprinting/. The command guide lists the intermediate tracks and motif comparison files.

Per-cell signatures use neighboring cells to reduce sparse signal. Compare patterns across annotated groups without treating each score as an independent binding measurement. See the Single-cell output example for a visual guide.

Small PBMC example

Use fp-tools 0.2.9 or later for this tutorial on Windows.

This real peripheral blood mononuclear cell (PBMC) example contains 300 cells: 100 B cells, 100 monocytes, and 100 T or natural killer cells. It retains human hg38 chromosome 22, 101,637 genomic bins (5,669 selected), and 110,888 fragment records. Counts and cell coordinates come from the annotated 10x PBMC5k dataset distributed by SnapATAC2; coordinates were not recomputed for this subset. Accessible regions are the source's selected bins, not newly called peaks.

Download fp-tools-pbmc-chr22-demo-v1.zip and its .sha256 file from the v0.2.9 release. The ZIP is approximately 14.3 MB. Open a terminal in the folder containing your download. On macOS, run shasum -a 256 fp-tools-pbmc-chr22-demo-v1.zip; on Linux use sha256sum instead of shasum -a 256. In Windows PowerShell, run Get-FileHash .\fp-tools-pbmc-chr22-demo-v1.zip -Algorithm SHA256. Compare the complete hash with the .sha256 file, then extract the entire folder. It includes fragments and their index, annotations, count AnnData, the chromosome FASTA and index, chromosome sizes, a matched blacklist, selected regions, and workflow.yml. manifest.json records the source hashes and subset definition; SHA256SUMS.txt records every input hash.

Open a terminal inside the extracted fp-tools-pbmc-chr22-demo-v1 folder. With the Python package installed, run this supplied example:

run-yaml-workflow --config workflow.yml

For the Windows or macOS desktop app, open the GUI's Config page and load workflow.yml. Replace the relative input filenames in the editor with their full paths inside the extracted folder, and set outdir to a new results folder using its full path. Apply the configuration, resolve any validation errors, then start the run. Saved YAML remains runnable from the command line. The command uses all available cores. The first use also prepares the motif database.

Check results/pseudobulk_footprint_manifest.tsv for the three completed groups. Open the figures under results/plots/single_cell_footprinting/ and inspect the group-level signal tracks and motif comparisons alongside them. This small subset teaches the workflow; it does not reproduce the full PBMC example figures or validate binding, biological differences, or full-scale performance. There are no biological replicates in this example.

Prepare the full PBMC5k inputs

For a source installation, the repository includes prepare_10x_pbmc5k_scatac.py. This preparation step uses SnapATAC2 on Linux x86-64 or macOS. The downloadable tutorial does not require installing SnapATAC2. The recipe below was checked with Python 3.12 and the listed package versions. Run it from the root of a v0.2.8 source checkout:

python3.12 -m venv ../fp-tools-pbmc-prep
source ../fp-tools-pbmc-prep/bin/activate
python -m pip install \
  "snapatac2==2.9.0" "anndata==0.12.17" "pandas==2.3.3" \
  "matplotlib==3.10.9" "pysam==0.23.3" "PyYAML==6.0.3"
python benchmarks/scripts/prepare_10x_pbmc5k_scatac.py --chroms chr22

The script retrieves the matching fragment file and annotated AnnData through snapatac2.datasets.pbmc5k, preserving the source count matrix and boolean var['selected']. It writes barcode-matched annotations for all 4,437 cells (422 B cells, 1,702 monocytes, and 2,313 T or natural killer cells), selected-bin BED files, and a preparation summary. The --chroms setting limits the demo region list and chromosome-size table; it does not subset the source cells or count matrix. BED coordinates use zero-based starts and exclude the end. After preparing the files, reactivate your fp-tools environment to run the analysis, or select the prepared inputs in the desktop app.

The annotated AnnData comes from the SnapATAC2 PBMC5k archive; its SHA-256 is 592f1551c27d0cfe4d81e7febad624d6b7d3ebf977b0c3ea64e06b3f3d76f078. The matching 10x fragment file has SHA-256 5fe44c0f8f76ce1534c1ae418cf0707ca5ef712004eee77c3d98d2d4b35ceaec. Use an hg38 reference and matching blacklist when adapting the command above. For the exact 300-cell subset, the versioned build_pbmc_chr22_demo.py accepts these prepared files plus the reference and blacklist; its --help lists the required paths. It selects the first 100 sorted barcodes per broad group and writes the complete ZIP and checksums.