Training CLIP models on Data from Scientific Papers
-
Updated
Nov 9, 2023 - TeX
Training CLIP models on Data from Scientific Papers
Bulk-download image-text datasets into Backblaze B2 as reproducible WebDataset tar shards with img2dataset — no local staging disk. Sample app (FastAPI + Next.js) that streams shards to S3-compatible object storage, validates download yield, and streams them back for PyTorch/JAX training.
A simple toolkit to transform datasource generate by img2dataset from parquet file to Huggingface dataset.
Add a description, image, and links to the img2dataset topic page so that developers can more easily learn about it.
To associate your repository with the img2dataset topic, visit your repo's landing page and select "manage topics."