Skip to main content

Practice Data and Interactive Environments

You do not need your own MRI data to learn DTI preprocessing. Several publicly available datasets provide high-quality diffusion data that you can use to practice every step in this tutorial. This page lists recommended datasets, how to download them, and how to use the interactive Binder environment.

Quick Start: Stanford HARDI​

The Stanford HARDI dataset is small, well-documented, and frequently used in diffusion MRI tutorials. It is available directly through the DIPY Python library:

from dipy.data import fetch_stanford_hardi, read_stanford_hardi
fetch_stanford_hardi() # Downloads ~90 MB
img, gtab = read_stanford_hardi()

Details:

  • Single subject
  • 160 gradient directions at b=2000, 10 b=0 images
  • 1.5 mm isotropic resolution
  • ~90 MB download

This dataset is ideal for testing individual pipeline steps but does not include fieldmaps or a T1 structural, so you cannot practice TOPUP or skull stripping with it.

Full Pipeline Practice: OpenNeuro​

OpenNeuro hosts thousands of freely available neuroimaging datasets in BIDS format. These include the complete set of scans needed for the full pipeline.

DatasetDescriptionShellsFieldmaps?Size
ds000201Multi-shell diffusion, young adultsb=1000, 2000Yes~2 GB/subject
ds001021HCP-style multi-shellMultipleYes~3 GB/subject
ds000030UCLA Consortium, large sampleb=1000Varies~1 GB/subject

Download a single subject first to test your pipeline before processing the full dataset. For ds000201, download just sub-01 (~2 GB) — this is enough to practice all 14 pipeline steps.

Downloading from OpenNeuro​

# Option 1: OpenNeuro CLI (recommended for large downloads)
pip install openneuro-cli
openneuro download --dataset ds000201 --include sub-01 /path/to/output

# Option 2: AWS CLI (direct from S3)
aws s3 sync --no-sign-request \
s3://openneuro.org/ds000201/sub-01 \
/path/to/output/sub-01

# Option 3: DataLad (version-controlled download)
pip install datalad
datalad install https://github.com/OpenNeuroDatasets/ds000201.git
cd ds000201
datalad get sub-01/

Gold Standard: Human Connectome Project (HCP)​

The HCP provides the highest-quality multi-shell diffusion data available:

  • 3 shells: b=1000, 2000, 3000 s/mm2^2
  • 90 gradient directions per shell + 18 b=0 images
  • 1.25 mm isotropic resolution
  • Both AP and PA acquisitions for comprehensive distortion correction
  • ~4 GB per subject (diffusion data only)

Access: Registration required (free for academic use). Apply at ConnectomeDB.

Disk Space Requirements​

Plan your storage before downloading:

DatasetPer Subject10 Subjects
Stanford HARDI~90 MB~90 MB (single subject only)
OpenNeuro ds000201~2 GB raw~20 GB raw
HCP (diffusion only)~4 GB raw~40 GB raw
Processing space~5–10 GB/subject~50–100 GB

Preprocessing generates intermediate files at each stage. For multi-shell data, expect to use 5–10 GB per subject in addition to the raw data. BedpostX alone generates 2–5 GB of output. Plan for 3–5x your raw data size in total disk space.

Setting Up Your Own Practice Environment​

If you prefer to work locally:

  1. Install the tools: Follow the Environment Setup guide to install FSL, ANTs, MRtrix3, and dcm2niix
  2. Download a practice dataset: Use one of the OpenNeuro datasets above
  3. Organize your data: Create a project directory structure:
my_dti_project/
raw/ # Downloaded data (keep untouched)
sub-01/
nifti/ # Converted NIfTI files
ants/ # Skull stripping output
topup/ # TOPUP output
eddy/ # Eddy output
dtifit/ # Tensor fitting output
...
config/ # acqp.txt, index.txt
logs/ # Processing logs
  1. Create configuration files: Set up acqp.txt and index.txt for your data — see Configuration Files
  2. Follow the pipeline: Work through each step in the Pipeline section

Adapting to Your Own Data​

Once you are comfortable with practice data, applying the pipeline to your own research data requires changing:

What ChangesWhere to Look
File paths and namingEvery script — update base_dir, subj, file suffixes
acqp.txtMust match your scanner's phase encoding direction and readout time — see Configuration Files
index.txtMust have one entry per DWI volume — see Configuration Files
Shell selectionShell Extraction — depends on your b-values and planned analysis
Template choiceStep 2 and Step 10 — match your population

The preprocessing steps themselves are identical regardless of the data source. The commands, quality checks, and troubleshooting all apply.