Sign inSign up

lightly/train

By lightly

•Updated 2 months ago

Train models with self-supervised learning in a single command.

Image
Machine learning & AI
2

5.0K

lightly/train repository overview

LightlyTrain

⁠LightlyTrain - SOTA Pretraining, Fine-tuning and Distillation

Python Docker Documentation Discord

Train Better Models, Faster

LightlyTrain is the leading framework for transforming your data into state-of-the-art computer vision models. It covers the entire model development lifecycle from pretraining DINOv2/v3 vision foundation models on your unlabeled data to fine-tuning transformer and YOLO models on detection and segmentation tasks for edge deployment.

Struggling to get good results with pre-training? Talk to one of our experts Contact us⁠

Using LightlyTrain at work, in production, on the edge, or to build proprietary models? You likely need a Commercial License. Contact us⁠ to request a license for commercial use.

Also check out LightlyStudio⁠ to easily visualize your annotations and predictions.

⁠Installation

Install LightlyTrain on Python 3.8+ for Windows, Linux or MacOS with:

pip install lightly-train

⁠Workflows

Object Detection

Train LTDETR detection models with DINOv2, DINOv3, or EdgeCrafter ECViT backbones.

⁠Usage

Documentation Colab

import lightly_train

if __name__ == "__main__":
    # Train with our most recent LT-DETRv2 detector based on DINOv3 and EdgeCrafter.
    lightly_train.train_object_detection(
        out="out/my_experiment",
        model="ltdetrv2-s-coco",
        data={
            "path": "my_data_dir",
            "train": "images/train",
            "val": "images/val",
            "names": {
                0: "person",
                1: "bicycle",
                2: "car",
                # ...
            },
        },
    )

    # Load model and run inference
    model = lightly_train.load_model("out/my_experiment/exported_models/exported_best.pt")
    # Or use one of the models provided by LightlyTrain
    # model = lightly_train.load_model("ltdetrv2-s-coco")
    results = model.predict("image.jpg")
    results["labels"]   # Class labels, tensor of shape (num_boxes,)
    results["bboxes"]   # Bounding boxes in (xmin, ymin, xmax, ymax) absolute pixel
                        # coordinates of the original image. Tensor of shape (num_boxes, 4).
    results["scores"]   # Confidence scores, tensor of shape (num_boxes,)
Instance Segmentation

Train state-of-the-art instance segmentation models with DINOv3 backbones using the EoMT method from CVPR 2025.

⁠Usage

Documentation Colab

import lightly_train

if __name__ == "__main__":
    # Train an instance segmentation model with our new LTDETRv2 family
    lightly_train.train_instance_segmentation(
        out="out/my_experiment",
        model="ltdetrv2-seg-s-coco",
        data={
            "format": "yolo",           # either "yolo" or "coco"
            "path": "my_data_dir",
            "train": "images/train",
            "val": "images/val",
            "names": {
                0: "background",
                1: "vehicle",
                2: "pedestrian",
                # ...
            },
        },
    )

    model = lightly_train.load_model("out/my_experiment/exported_models/exported_best.pt")
    # Or use one of the models provided by LightlyTrain
    # model = lightly_train.load_model("ltdetrv2-seg-s-coco")
    results = model.predict("image.jpg")
    results["labels"]   # Class labels, tensor of shape (num_instances,)
    results["bboxes"]   # Bounding boxes in (xmin, ymin, xmax, ymax) absolute pixel
                        # coordinates of the original image. Tensor of shape (num_instances, 4).
    results["masks"]    # Binary masks, tensor of shape (num_instances, height, width).
                        # Height and width correspond to the original image size.
    results["scores"]   # Confidence scores, tensor of shape (num_instances,)
Semantic Segmentation

Train state-of-the-art semantic segmentation models with DINOv2 or DINOv3 backbones using the EoMT method from CVPR 2025.

⁠Usage

Documentation Colab

import lightly_train

if __name__ == "__main__":
    # Train a semantic segmentation model with a DINOv3 backbone
    lightly_train.train_semantic_segmentation(
        out="out/my_experiment",
        model="dinov3/vits16-eomt",
        data={
            "train": {
                "images": "my_data_dir/train/images",
                "masks": "my_data_dir/train/masks",
            },
            "val": {
                "images": "my_data_dir/val/images",
                "masks": "my_data_dir/val/masks",
            },
            "classes": {
                0: "background",
                1: "road",
                2: "building",
                # ...
            },
        },
    )

    # Load model and run inference
    model = lightly_train.load_model("out/my_experiment/exported_models/exported_best.pt")
    # Or use one of the models provided by LightlyTrain
    # model = lightly_train.load_model("dinov3/vits16-eomt")
    masks = model.predict("image.jpg")
    # Masks is a tensor of shape (height, width) with class labels as values.
    # It has the same height and width as the input image.
Panoptic Segmentation

Train state-of-the-art panoptic segmentation models with DINOv3 backbones using the EoMT method from CVPR 2025.

⁠Usage

Documentation Colab

import lightly_train

if __name__ == "__main__":
    # Train an panoptic segmentation model with a DINOv3 backbone
    lightly_train.train_panoptic_segmentation(
        out="out/my_experiment",
        model="dinov3/vitb16-eomt-panoptic-coco",
        data={
            "train": {
                "images": "images/train",
                "masks": "annotations/train",
                "annotations": "annotations/train.json",
            },
            "val": {
                "images": "images/val",
                "masks": "annotations/val",
                "annotations": "annotations/val.json",
            },
        },
    )

    model = lightly_train.load_model("out/my_experiment/exported_models/exported_best.pt")
    results = model.predict("image.jpg")
    results["masks"]    # Masks with (class_label, segment_id) for each pixel, tensor of
                        # shape (height, width, 2). Height and width correspond to the
                        # original image size.
    results["segment_ids"]    # Segment ids, tensor of shape (num_segments,).
    results["scores"]   # Confidence scores, tensor of shape (num_segments,)
Depth Estimation

Run monocular depth inference with Depth Anything V2 and V3 models. Training support will be released soon!

⁠Usage

Documentation Colab

import lightly_train

# Load a depth model provided by LightlyTrain
model = lightly_train.load_model("dinov2/dav3-relative-large")

# Predict a relative-depth map
depth = model.predict("image.jpg")
# depth is a tensor of shape (height, width) matching the input image.

Metric depth (in meters) and the full list of available models are covered in the documentation⁠.

Image Classification

Train multiclass or multilabel image classification models with any backbone.

⁠Usage

Documentation Colab

import lightly_train

if __name__ == "__main__":
    # Train an image classification model with a DINOv3 backbone
    lightly_train.train_image_classification(
        out="out/my_experiment",
        model="dinov3/vitt16",
        data={
            "train": "my_data_dir/train/",
            "val": "my_data_dir/val/",
            "classes": {
                0: "cat",
                1: "car",
                2: "dog",
                # ...
            },
        },
    )

    model = lightly_train.load_model("out/my_experiment/exported_models/exported_best.pt")
    results = model.predict("image.jpg", topk=1, threshold=0.5)
    results["labels"]   # Class labels, tensor of shape (topk,)
    results["scores"]   # Confidence scores, tensor of shape (topk,)
Distillation (DINOv2/v3)

Pretrain any model architecture with unlabeled data by distilling the knowledge from DINOv2 or DINOv3 foundation models into your model. On the COCO dataset, YOLOv8-s models pretrained with LightlyTrain achieve high performance across all tested label fractions. These improvements hold for other architectures like YOLOv11, RT-DETR, and Faster R-CNN. See our announcement post⁠ for more benchmarks and details.

⁠Usage

Documentation Google Colab

import lightly_train

if __name__ == "__main__":
    # Distill the knowledge from a DINOv3 teacher into a YOLOv8 model
    lightly_train.pretrain(
        out="out/my_experiment",
        data="my_data_dir",
        model="ultralytics/yolov8s",
        method="distillation",
        method_args={
            "teacher": "dinov3/vitb16",
        },
    )

    # Load model for fine-tuning
    model = YOLO("out/my_experiment/exported_models/exported_last.pt")
    model.train(data="coco8.yaml")
Pretraining (DINOv2 Foundation Models)

With LightlyTrain you can train your very own foundation model like DINOv2 on your data.

⁠Usage

Documentation

import lightly_train

if __name__ == "__main__":
    # Pretrain a DINOv2 vision foundation model
    lightly_train.pretrain(
        out="out/my_experiment",
        data="my_data_dir",
        model="dinov2/vitb14",
        method="dinov2",
    )
Autolabeling

LightlyTrain provides simple commands to autolabel your unlabeled data using DINOv2 or DINOv3 pretrained models. This allows you to efficiently boost performance of your smaller models by leveraging all your unlabeled images.

⁠Usage

Documentation

import lightly_train

if __name__ == "__main__":
    # Autolabel your data with a DINOv3 semantic segmentation model
    lightly_train.predict_semantic_segmentation(
        out="out/my_autolabeled_data",
        data="my_data_dir",
        model="dinov3/vitb16-eomt-coco",
        # Or use one of your own model checkpoints
        # model="out/my_experiment/exported_models/exported_best.pt",
    )

    # The autolabeled masks will be saved in this format:
    # out/my_autolabeled_data
    # ├── <image name>.png
    # ├── <image name>.png
    # └── …

⁠Features

⁠Models

LightlyTrain supports the following model and workflow combinations.

⁠Fine-tuning
ModelObject
Detection
Instance
Segmentation
Panoptic
Segmentation
Semantic
Segmentation
Image
Classification
DINOv3✅ 🔗⁠✅ 🔗⁠✅ 🔗⁠✅ 🔗⁠✅ 🔗⁠
DINOv2✅ 🔗⁠✅ 🔗⁠✅ 🔗⁠✅ 🔗⁠✅ 🔗⁠
EdgeCrafter✅ 🔗⁠✅ 🔗⁠
Any✅ 🔗⁠
⁠Distillation & Pretraining
ModelDistillationPretraining
DINOv3✅ 🔗⁠
DINOv2✅ 🔗⁠✅ 🔗⁠
Torchvision ResNet, ConvNext, ShuffleNetV2✅ 🔗⁠✅ 🔗⁠
TIMM models✅ 🔗⁠✅ 🔗⁠
Ultralytics YOLOv5–YOLO26, RT-DETR✅ 🔗⁠✅ 🔗⁠
RT-DETR, RT-DETRv2✅ 🔗⁠✅ 🔗⁠
RF-DETR✅ 🔗⁠✅ 🔗⁠
YOLOv12✅ 🔗⁠✅ 🔗⁠
Custom PyTorch Model✅ 🔗⁠✅ 🔗⁠

Contact us⁠ if you need support for additional models.

⁠Usage Events

LightlyTrain collects anonymous usage events to help us improve the product. We only track training method, model architecture, and system information (OS, GPU, CI, Container). To opt-out, set the environment variable: export LIGHTLY_TRAIN_EVENTS_DISABLED=1

⁠License

LightlyTrain offers flexible licensing options to suit your specific needs:

  • AGPL-3.0 License: Perfect for open-source projects, academic research, and community contributions. Share your innovations with the world while benefiting from community improvements.

  • Commercial License: Ideal for businesses and organizations that need proprietary development freedom. Enjoy all the benefits of LightlyTrain while keeping your code and models private. Includes model training and runtime license.

  • Free Community License: Available for students, researchers, startups in early stages, or anyone exploring or experimenting with LightlyTrain. Empower the next generation of innovators with full access to the world of pretraining.

⁠Commercial Pricing
PlanPriceEligibility
Startup$5,000 / year< $1M revenue or < 10 employees
Growth$10,000 / year< $10M revenue or < 100 employees
EnterpriseCustom> $10M revenue or > 100 employees

All commercial plans include a license for model training, edge deployment, and inference. Enterprise plans include priority support, a joint Slack channel, co-development engineering, and influence on the product roadmap.

Contact us⁠ to get started — we'll find the right option for your project!

⁠Contact

Website
Discord
GitHub
X
YouTube
LinkedIn

Tag summary

Content type

Image

Digest

sha256:0e5072608…

Size

3.8 GB

Last updated

2 months ago

docker pull lightly/train