Vision Data Curation (VDC) is a data-curation toolkit in the Birder ecosystem for cleaning, filtering and sampling large-scale image datasets. Built for computer vision researchers and practitioners who want higher-quality data with less manual effort.
At its core, VDC embraces a data-centric AI philosophy, aiming to enhance the quality and diversity of your datasets with minimal manual intervention. By providing a structured, iterative pipeline, VDC helps researchers and practitioners build higher-performing models by starting with better data.
VDC provides modular tools for dataset cleanup:
- Input validation - detect corrupt files, invalid formats, low resolution, or extreme aspect ratios
- Rotation correction - correct 90°/180°/270° orientation errors
- Example-based filtering - remove images similar to a set of unwanted examples
- Image quality filtering - remove images based on aesthetic score or NSFW classification
- Duplicate removal - identify and remove near-duplicate images from your dataset
- Hierarchical K-Means sampling - select diverse, representative subsets from large datasets
flowchart LR
A[Raw<br/>Dataset] --> V[Validation]
V --> R[Rotation]
R --> D[Dedup]
D --> E[Example<br/>Filter]
E --> Q[Quality Filter<br/>Aesthetic/NSFW]
Q --> S[Cluster-based<br/>Sampling]
S --> F[Curated<br/>Dataset]
U[Unwanted<br/>Examples] --> E
pip install vision-data-curationgit clone https://gitlab.com/birder/vision-data-curation.git
cd vision-data-curation
pip install -e .Developing directly from the project root allows for script and configuration execution as if fully installed.
Each step is a script under vdc.scripts.
Examples:
# Report corrupt/invalid images
python -m vdc.scripts.sanitize_images data/raw_images/
# Filter based on "Unwanted examples"
python -m vdc.scripts.filter_by_examples data/embeddings.csv --examples-embeddings-file bad_examples.csv- Run
python -m vdc.scriptsto see available scripts - Run
python -m vdc.scripts.<script> --helpfor options
Model Downloads: VDC automatically downloads the aesthetic and NSFW predictor heads to the models/ directory (defined by vdc.conf.settings.MODELS_DIR). Embedding models used with VDC, such as SSCD, PE and CLIP, are managed separately by Birder and should be downloaded with python -m birder.tools download-model <model> before running Birder prediction commands.
- Default settings live in vdc/conf/config.json
- A
config.jsonin your project root will take precedence (or pass--configto any script)
VDC provides accompanying R Markdown notebooks (found in the notebooks/ directory) to visually inspect and understand the impact of various curation steps, aiding in data-driven decision making.
While these notebooks use the .Rmd extension, they are entirely written in Python.
This choice was made for their text-based nature, which significantly improves version control by producing clean and meaningful Git diffs compared to traditional .ipynb files.
For a seamless interactive experience in a Jupyter environment, we recommend installing jupytext (pip install jupytext). With jupytext installed, you can open and run .Rmd files directly in Jupyter as if they were .ipynb notebooks.
For detailed walkthroughs, examples, and in-depth guides, please refer to our full documentation.
This project is licensed under the Apache-2.0 License - see the LICENSE file for details.