Single-cell ATAC-seq workflow¶
sc-footprinting groups fragments into pseudobulk samples, runs the footprint
analysis, and calculates per-cell footprint signatures. A pseudobulk sample
combines fragments from cells that share an annotation, such as a cell type.
Main commands¶
sc-footprintingruns the complete workflow.pseudobulk-fragmentsandfind-signature-fpremain available as focused utilities.
Prepare your inputs¶
For your first run, use the complete PBMC example below. The general command and annotation rows in this section are templates for your own data. Run them from a folder containing the named input files.
| Input | Required content |
|---|---|
fragments.tsv.gz |
Fragment records with chromosome, start, end, and barcode in the first four columns. |
cell_annotations.tsv |
Cell barcodes, group labels, and UMAP coordinates, as shown below. |
genomic_bin_counts.h5ad |
AnnData with genomic-bin counts and matching cell names; an embedding alone is insufficient. |
hg38.chrom.sizes |
Two columns: chromosome name and length. |
hg38.fa.gz |
Reference FASTA matching the fragments and peaks. |
merged_peaks.bed |
Accessible regions in the same assembly and chromosome naming scheme. |
The annotation table needs these columns:
barcode cell_type snap_cell_type umap_1 umap_2
AAACGAAAGAAACGCC-1 B_cell B_cell 1.2 -0.8
AAACGAAAGAAAGCAG-1 T_cell T_cell -2.1 0.4
Use your actual barcodes and UMAP coordinates. cell_type contains broad
labels and snap_cell_type contains detailed labels; they can be identical
if you use one annotation level. --group-by cell_type combines cells with
the same cell_type value.
The AnnData cell names must match the annotation barcodes. Its features must
be genomic bins such as chr1:0-500, with a boolean var['selected'] column
and a count matrix. The companion motif-activity calculation uses 500-base
bins. The workflow uses obsm['X_spectral'] for nearest-neighbor smoothing
when available, otherwise it uses the annotation UMAP coordinates.
Run the workflow¶
sc-footprinting \
--fragments fragments.tsv.gz \
--annotations cell_annotations.tsv \
--h5ad genomic_bin_counts.h5ad \
--group-by cell_type \
--genome-sizes hg38.chrom.sizes \
--genome hg38.fa.gz \
--peaks merged_peaks.bed \
--motif-db jaspar2026_vertebrates \
--outdir project/pseudobulk
Review the results¶
Check project/pseudobulk/pseudobulk_footprint_manifest.tsv for completed
groups. Per-cell score tables, heatmaps, and UMAPs are written under
project/pseudobulk/plots/single_cell_footprinting/. The
command guide lists the intermediate tracks
and motif comparison files.
Per-cell signatures use neighboring cells to reduce sparse signal. Compare patterns across annotated groups without treating each score as an independent binding measurement. See the Single-cell output example for a visual guide.
Small PBMC example¶
Use fp-tools 0.2.9 or later for this tutorial on Windows.
This real peripheral blood mononuclear cell (PBMC) example contains 300 cells: 100 B cells, 100 monocytes, and 100 T or natural killer cells. It retains human hg38 chromosome 22, 101,637 genomic bins (5,669 selected), and 110,888 fragment records. Counts and cell coordinates come from the annotated 10x PBMC5k dataset distributed by SnapATAC2; coordinates were not recomputed for this subset. Accessible regions are the source's selected bins, not newly called peaks.
Download fp-tools-pbmc-chr22-demo-v1.zip and its .sha256 file from the
v0.2.9 release.
The ZIP is approximately 14.3 MB. Open a terminal in the folder containing your
download. On macOS, run
shasum -a 256 fp-tools-pbmc-chr22-demo-v1.zip; on Linux use sha256sum instead
of shasum -a 256. In Windows PowerShell, run
Get-FileHash .\fp-tools-pbmc-chr22-demo-v1.zip -Algorithm SHA256.
Compare the complete hash with the .sha256 file, then extract the entire
folder. It includes fragments and their index, annotations,
count AnnData, the chromosome FASTA and index, chromosome sizes, a matched
blacklist, selected regions, and workflow.yml. manifest.json records the
source hashes and subset definition; SHA256SUMS.txt records every input hash.
Open a terminal inside the extracted fp-tools-pbmc-chr22-demo-v1 folder.
With the Python package installed, run this supplied example:
run-yaml-workflow --config workflow.yml
For the Windows or macOS desktop app, open the GUI's Config page and load
workflow.yml. Replace the relative input filenames in the editor with their
full paths inside the extracted folder, and set outdir to a new results
folder using its full path. Apply the configuration, resolve any validation
errors, then start the run. Saved YAML remains runnable from the command line.
The command uses all
available cores. The first use also prepares the motif database.
Check results/pseudobulk_footprint_manifest.tsv for the three completed groups.
Open the figures under results/plots/single_cell_footprinting/ and inspect the
group-level signal tracks and motif comparisons alongside them. This small
subset teaches the workflow; it does not reproduce the full PBMC example
figures or validate binding, biological differences, or full-scale performance.
There are no biological replicates in this example.
Prepare the full PBMC5k inputs¶
For a source installation, the repository includes
prepare_10x_pbmc5k_scatac.py.
This preparation step uses SnapATAC2 on Linux x86-64 or macOS. The downloadable
tutorial does not require installing SnapATAC2. The recipe below was checked
with Python 3.12 and the listed package versions. Run it from the root of a
v0.2.8 source checkout:
python3.12 -m venv ../fp-tools-pbmc-prep
source ../fp-tools-pbmc-prep/bin/activate
python -m pip install \
"snapatac2==2.9.0" "anndata==0.12.17" "pandas==2.3.3" \
"matplotlib==3.10.9" "pysam==0.23.3" "PyYAML==6.0.3"
python benchmarks/scripts/prepare_10x_pbmc5k_scatac.py --chroms chr22
The script retrieves the matching fragment file and annotated AnnData through
snapatac2.datasets.pbmc5k, preserving the source count matrix and boolean
var['selected']. It writes barcode-matched annotations for all 4,437 cells
(422 B cells, 1,702 monocytes, and 2,313 T or natural killer cells), selected-bin
BED files, and a preparation summary. The --chroms setting limits the demo
region list and chromosome-size table; it does not subset the source cells or
count matrix. BED coordinates use zero-based starts and exclude the end.
After preparing the files, reactivate your fp-tools environment to run the
analysis, or select the prepared inputs in the desktop app.
The annotated AnnData comes from
the SnapATAC2 PBMC5k archive; its SHA-256 is
592f1551c27d0cfe4d81e7febad624d6b7d3ebf977b0c3ea64e06b3f3d76f078.
The matching 10x fragment file
has SHA-256 5fe44c0f8f76ce1534c1ae418cf0707ca5ef712004eee77c3d98d2d4b35ceaec.
Use an hg38 reference and matching blacklist when adapting the command above.
For the exact 300-cell subset, the versioned
build_pbmc_chr22_demo.py
accepts these prepared files plus the reference and blacklist; its --help
lists the required paths. It selects the first 100 sorted barcodes per broad
group and writes the complete ZIP and checksums.