Sign inSign up

szymonmaszke/torchdata

By szymonmaszke

•Updated over 6 years ago

tf.data-like library for PyTorch (cache, map etc.) with additions

Image
0

3.2K

szymonmaszke/torchdata repository overview

VersionDocsTestsCoverageStylePyPIPythonPyTorchDockerRoadmap
VersionDocumentationTestsPyPIPythonPyTorchDockerRoadmap

torchdata⁠ is PyTorch⁠ oriented library focused on data processing and input pipelines in general.

It extends torch.utils.data.Dataset and equips it with functionalities known from tensorflow.data⁠ like map or cache (with some additions unavailable in aforementioned) . All of that with minimal interference (single call to super().__init__()) in original PyTorch's datasets.

⁠Functionalities overview:
  • map or apply arbitrary functions to dataset
  • cache allows you to cache data in memory or on disk (even partially, say first 20%)
  • Full torch.utils.data.IterableDataset and torch.utils.data.Dataset support
  • Easy to create custom methods of caching, choosing elements to cache, maps and datasets
  • Concrete and base classes designed for file reading and other general tasks

⁠Quick examples

  • Create image dataset, convert it to Tensors, cache and concatenate with smoothed labels:
# Imports assumed
# Example dataset return all 1 labels
class Labels(torchdata.Dataset):
    def __init__(self, length):
        self.length = length
        super().__init__()

    def __getitem__(self, _):
        return 1

    def __len__(self):
        return len(length)


# Convenience class based on torchdata.Dataset
class ImageDataset(torchdata.Files):
    def __getitem__(self, index):
        return Image.open(self.files[index])


images = (
    ImageDataset.from_folder("./data").map(torchvision.transforms.ToTensor()).cache()
)

smoothed_labels = Labels(len(images)).map(lambda label: label - 0.1)

# That's how you concatenate sample-wise
for image, label in images | smoothed_labels:
    pass
  • Cache first 1000 samples in memory, save the rest on disk in folder ./cache:
images = (
    ImageDataset.from_folder("./data").map(torchvision.transforms.ToTensor())
    # First 1000 samples in memory
    .cache(torchdata.modifiers.UpToIndex(torchdata.cachers.Memory(), 1000))
    # Sample from 1000 to the end saved with Pickle on disk
    .cache(torchdata.modifiers.FromIndex(torchdata.cachers.Pickle("./cache"), 1000))
    # You can define your own cachers, modifiers, see docs
)

To see what else you can do please check torchdata documentation⁠

⁠Installation

⁠pip⁠

⁠Latest release:
pip install --user torchdata
⁠Nightly:
pip install --user torchdata-nightly

⁠Docker⁠

CPU standalone and various versions of GPU enabled images are available at dockerhub⁠.

For CPU quickstart, issue:

docker pull szymonmaszke/torchdata:18.04

Nightly builds are also available, just prefix tag with nightly_. If you are going for GPU image make sure you have nvidia/docker⁠ installed and it's runtime set.

⁠Contributing

If you find any issue or you think some functionality may be useful to others and fits this library, please open new Issue⁠ or create Pull Request⁠.

To get an overview of something which one can done to help this project, see Roadmap⁠

Tag summary

Content type

Image

Digest

Size

1.2 GB

Last updated

over 6 years ago

docker pull szymonmaszke/torchdata:nightly_10.1-cudnn7-runtime-ubuntu18.04