diff --git a/README.md b/README.md index 72feb7592..47ff630dd 100644 --- a/README.md +++ b/README.md @@ -1,12 +1,12 @@ -[**Installation**](#installing-torchcodec) | [**Simple Example**](#using-torchcodec) | [**Detailed Example**](https://meta-pytorch.org/torchcodec/stable/generated_examples/) | [**Documentation**](https://meta-pytorch.org/torchcodec) | [**Contributing**](CONTRIBUTING.md) | [**License**](#license) +[**Installation**](#installing-torchcodec) | [**Documentation**](https://meta-pytorch.org/torchcodec) | [**Contributing**](CONTRIBUTING.md) | [**License**](#license) # TorchCodec -TorchCodec is a Python library for decoding video and audio data into PyTorch -tensors, on CPU and CUDA GPU. It also supports video and audio encoding on CPU! -It aims to be fast, easy to use, and well integrated -into the PyTorch ecosystem. If you want to use PyTorch to train ML models on -videos and audio, or run inference, TorchCodec is how you turn these into data. +TorchCodec is a PyTorch-native library for decoding and encoding media: videos, +audio, and images, on CPU and CUDA GPU. It aims to be fast, easy to +use, and well integrated into the PyTorch ecosystem. If you want to use PyTorch +to train ML models on videos, audio, or images, or run inference, TorchCodec is +how you turn these into tensors, and back. We achieve these capabilities through: @@ -16,7 +16,7 @@ We achieve these capabilities through: installed. FFmpeg is a mature library with broad coverage available on most systems. It is, however, not easy to use. TorchCodec abstracts FFmpeg's complexity to ensure it is used correctly and efficiently. (FFmpeg is - optional, and the image decoders don't need it: see [Installing + optional, and the image decoders and encoders don't need it: see [Installing TorchCodec](#installing-torchcodec).) * Returning data as PyTorch tensors, ready to be fed into PyTorch transforms or used directly to train models. @@ -70,7 +70,7 @@ the `VideoDecoder`: ffmpeg -f lavfi -i testsrc2=size=640x400:duration=10:rate=25 /tmp/output_video.mp4 ``` -#### Encoding +#### Video and Audio Encoding ```python from torchcodec.encoders import Encoder @@ -90,6 +90,22 @@ with encoder.open_file("output.mp4"): # ... ``` +#### Image Decoding and Encoding + +```python +from torchcodec.decoders import decode_image, decode_jpeg +from torchcodec.encoders import JpegEncoder + +# JPEG, PNG, WebP, GIF, AVIF and HEIC, with the format detected automatically. +image = decode_image("path/to/image.jpg") # uint8 tensor of shape [C, H, W] + +# Or use the format-specific decoders, e.g. to decode JPEGs on GPU: +image = decode_jpeg("path/to/image.jpg", device="cuda") + +# JPEG and PNG encoding. JPEGEncoder also supports CUDA encoding! +JpegEncoder(image).to_file("output.jpg") # also .to_tensor() and .to_file_like() +``` + ## Installing TorchCodec 1. Install FFmpeg, if it's not already installed. TorchCodec supports all major @@ -109,9 +125,9 @@ with encoder.open_file("output.mp4"): > **Note:** FFmpeg is an *optional* dependency. It is needed for video > and audio decoding and encoding (`VideoDecoder`, `AudioDecoder`, - > `VideoEncoder`, `AudioEncoder`, etc.). The image decoders - > (`decode_image`, `decode_jpeg`, `decode_png`, etc.) - > do **not** require FFmpeg, so if you only need image decoding you can skip + > `VideoEncoder`, `AudioEncoder`, etc.). The image decoders and encoders + > (`decode_image`, `decode_jpeg`, `JpegEncoder`, `PngEncoder`, etc.) + > do **not** require FFmpeg, so if you only need images you can skip > this step. 2. Install PyTorch and TorchCodec: @@ -128,15 +144,15 @@ with encoder.open_file("output.mp4"): ### CUDA support -CUDA-enabled wheels are installed by default on Linux. For Windows, you'll need -to pass `--index-url` as described below. +On CUDA GPUs, TorchCodec supports decoding and encoding of videos and jpeg +images. CUDA-enabled wheels are installed by default on Linux. For Windows, +you'll need to pass `--index-url` as described below. + -Make sure you have a GPU with NVDEC hardware that can decode the format you -want. Refer to Nvidia's GPU support matrix +For video, make sure you have a GPU with NVDEC and NVENC hardware that supports +the formats you want. Refer to Nvidia's GPU support matrix [here](https://developer.nvidia.com/video-encode-and-decode-gpu-support-matrix-new). -You will need the `libnvrtc` CUDA library, which is usually part of the CUDA -Toolkit. To select a specific CUDA Toolkit version, use `--index-url`. Make sure to install the corresponding PyTorch version as well (refer to the @@ -147,12 +163,16 @@ install the corresponding PyTorch version as well (refer to the pip install torch torchcodec --index-url=https://download.pytorch.org/whl/cu130 ``` -Make sure your FFmpeg has NVDEC support: +Make sure your FFmpeg has NVDEC and NVENC support: ```bash ffmpeg -decoders | grep -i nvidia # This should show a line like this: # V..... h264_cuvid Nvidia CUVID H264 decoder (codec h264) + +ffmpeg -encoders | grep -i nvidia +# This should show a line like this: +# V....D h264_nvenc NVIDIA NVENC H.264 encoder (codec h264) ``` To check that FFmpeg libraries work with NVDEC correctly you can decode a diff --git a/docs/source/conf.py b/docs/source/conf.py index 350cf4067..8041d570b 100644 --- a/docs/source/conf.py +++ b/docs/source/conf.py @@ -64,9 +64,9 @@ def __init__(self, src_dir): self.src_dir = src_dir def __call__(self, filename): - # We have two top-level galleries, one for decoding examples and one for - # encoding examples. We define the example order within each gallery - # individually. + # We have three top-level galleries: decoding examples, encoding + # examples, and migration guides. We define the example order within + # each gallery individually. if "examples/decoding" in self.src_dir: order = [ "basic_example.py", @@ -82,14 +82,18 @@ def __call__(self, filename): "transforms.py", "hdr_decoding.py", ] - else: - assert "examples/encoding" in self.src_dir + elif "examples/encoding" in self.src_dir: order = [ "image_encoding.py", "audio_encoding.py", "video_encoding.py", "multi_stream_encoding.py", ] + else: + assert "examples/migration" in self.src_dir + order = [ + "torchvision_migration.py", + ] try: return order.index(filename) diff --git a/docs/source/index.rst b/docs/source/index.rst index 1f6e1016c..a1addc33b 100644 --- a/docs/source/index.rst +++ b/docs/source/index.rst @@ -1,20 +1,21 @@ Welcome to the TorchCodec documentation! ======================================== -TorchCodec is a Python library for decoding video and audio data into PyTorch -tensors, on CPU and CUDA GPU. It also supports audio and video encoding! -It aims to be fast, easy to use, and well integrated into the PyTorch ecosystem. -If you want to use PyTorch to train ML models on videos and audio, TorchCodec is -how you turn these into data. +TorchCodec is a PyTorch-native library for decoding and encoding media: videos, +audio, and images, on CPU and CUDA GPU. It aims to be fast, easy to +use, and well integrated into the PyTorch ecosystem. If you want to use PyTorch +to train ML models on videos, audio, or images, TorchCodec is how you turn these +into tensors, and back. We achieve these capabilities through: * Pythonic APIs that mirror Python and PyTorch conventions. -* Relying on `FFmpeg `_ to do the decoding / encoding. - TorchCodec uses the version of FFmpeg you already have installed. FFmpeg is a - mature library with broad coverage available on most systems. It is, however, - not easy to use. TorchCodec abstracts FFmpeg's complexity to ensure it is - used correctly and efficiently. +* Relying on `FFmpeg `_ to do the video and audio + decoding / encoding. TorchCodec uses the version of FFmpeg you already have + installed. FFmpeg is a mature library with broad coverage available on most + systems. It is, however, not easy to use. TorchCodec abstracts FFmpeg's + complexity to ensure it is used correctly and efficiently. FFmpeg is optional: + the image decoders and encoders don't need it. * Returning data as PyTorch tensors, ready to be fed into PyTorch transforms or used directly to train models. @@ -140,6 +141,19 @@ Encoding How to encode audio samples into an audio file +Migrating from torchvision +^^^^^^^^^^^^^^^^^^^^^^^^^^ + +.. grid:: 3 + + .. grid-item-card:: :octicon:`file-code;1em` + Migrating from torchvision + :link: generated_examples/migration/torchvision_migration.html + :link-type: url + + How to port ``torchvision.io`` image decoding and encoding code + + .. toctree:: :maxdepth: 1 :hidden: diff --git a/examples/decoding/image_decoding.py b/examples/decoding/image_decoding.py index 66c4aebc3..63acac4da 100644 --- a/examples/decoding/image_decoding.py +++ b/examples/decoding/image_decoding.py @@ -21,6 +21,13 @@ - :func:`~torchcodec.decoders.decode_gif` - :func:`~torchcodec.decoders.decode_avif` - :func:`~torchcodec.decoders.decode_heic` + +.. note:: + + These decoders supersede the ones from ``torchvision.io``: they are more + robust and support more features. See + :ref:`sphx_glr_generated_examples_migration_torchvision_migration.py` for a + migration guide. """ # %% diff --git a/examples/encoding/image_encoding.py b/examples/encoding/image_encoding.py index 45d0e53ea..aed5cf1e3 100644 --- a/examples/encoding/image_encoding.py +++ b/examples/encoding/image_encoding.py @@ -12,6 +12,13 @@ In this example, we'll learn how to encode an image tensor to JPEG or PNG using the :class:`~torchcodec.encoders.JpegEncoder` and :class:`~torchcodec.encoders.PngEncoder` classes. + +.. note:: + + These encoders supersede the ones from ``torchvision.io``: they are more + robust and support more features. See + :ref:`sphx_glr_generated_examples_migration_torchvision_migration.py` for a + migration guide. """ # %% diff --git a/examples/migration/README.rst b/examples/migration/README.rst new file mode 100644 index 000000000..e59c35315 --- /dev/null +++ b/examples/migration/README.rst @@ -0,0 +1,2 @@ +Migrating from torchvision +-------------------------- diff --git a/examples/migration/torchvision_migration.py b/examples/migration/torchvision_migration.py new file mode 100644 index 000000000..ef055320e --- /dev/null +++ b/examples/migration/torchvision_migration.py @@ -0,0 +1,150 @@ +# Copyright (c) Meta Platforms, Inc. and affiliates. +# All rights reserved. +# +# This source code is licensed under the BSD-style license found in the +# LICENSE file in the root directory of this source tree. + +""" +======================================== +Migrating from TorchVision to TorchCodec +======================================== + +The image decoders and encoders of ``torchvision.io`` now live in torchcodec. +This is a short guide to porting your code over. Everything you could do with +``torchvision.io`` you can do with TorchCodec, usually with a very similar call. +And TorchCodec supports many more features. To learn more about the image +decoding and encoding features of TorchCodec, refer to the +:ref:`image decoding ` +and +:ref:`image encoding ` +tutorials. + +TL;DR +----- + +- ``decode_image(x)`` -> ``decode_image(x)``, but watch out for the + :ref:`changed defaults ` +- ``decode_jpeg(x, device="cuda")`` -> ``decode_jpeg(x, device="cuda")``, same + caveat +- ``read_file(path)`` -> not needed, pass ``path`` to the decoder +- ``encode_jpeg(img, quality)`` -> ``JpegEncoder(img).to_tensor(quality=...)`` +- ``write_jpeg(img, path, quality)`` -> ``JpegEncoder(img).to_file(path, quality=...)`` +- ``encode_png(img, level)`` -> ``PngEncoder(img).to_tensor(compression_level=...)`` +- ``write_png(img, path, level)`` -> ``PngEncoder(img).to_file(path, compression_level=...)`` +- ``write_file(path, encoded)`` -> not needed, use ``to_file`` + +The rest of this guide goes over these one by one. +""" + +# %% +# A bit of boilerplate first: let's make up some encoded image bytes to play +# with, by encoding a random image. +import torch + +from torchcodec.encoders import JpegEncoder, PngEncoder + +raw_image_bytes = JpegEncoder( + torch.randint(0, 256, (3, 256, 256), dtype=torch.uint8) +).to_tensor() + +# %% +# Decoding +# -------- +# +# ``torchvision.io.decode_image`` becomes +# :func:`torchcodec.decoders.decode_image`. Both accept raw encoded bytes, a +# tensor of encoded bytes, or a path to a file: +# +# .. code-block:: python +# +# # Before +# from torchvision.io import decode_image +# image = decode_image("image.jpg") +# +# # After +# from torchcodec.decoders import decode_image +# image = decode_image("image.jpg") +# +# The format-specific decoders map over one-to-one as well: +# ``decode_jpeg``, ``decode_png``, ``decode_webp``, ``decode_gif``, and +# torchcodec adds ``decode_avif`` and ``decode_heic`` without needing the +# separate ``torchvision-extra-decoders`` package. + +from torchcodec.decoders import decode_image + +image = decode_image(raw_image_bytes) +print(f"{image.shape = }, {image.dtype = }") + +# %% +# ``torchvision.io.read_file`` has no equivalent, and you don't need one: pass +# the path (a ``str`` or a ``pathlib.Path``) straight to the decoder. +# +# .. code-block:: python +# +# # Before +# from torchvision.io import decode_image, read_file +# image = decode_image(read_file("image.jpg")) +# +# # After +# from torchcodec.decoders import decode_image +# image = decode_image("image.jpg") + +# %% +# .. _decoding_defaults: +# +# A few decoding defaults changed +# ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ +# +# - ``mode`` now defaults to ``"RGB"`` instead of ``"UNCHANGED"``. If you were +# relying on the source's own channel layout, pass ``mode="UNCHANGED"``. +# - The output is always ``torch.uint8`` by default, even for 16-bit sources. +# To get torchvision's behaviour, where the dtype follows the source, pass +# ``output_dtype="auto"``. +# - The ``apply_exif_orientation`` parameter is gone: EXIF orientation is +# always applied. + +print(f"{decode_image(raw_image_bytes, mode='GRAY').shape = }") +print(f"{decode_image(raw_image_bytes, output_dtype=torch.uint16).dtype = }") + +# %% +# Encoding +# -------- +# +# The encoding functions became classes: instantiate an encoder with the image, +# then choose where the encoded bytes should go. +# +# .. code-block:: python +# +# # Before +# from torchvision.io import encode_jpeg, write_jpeg +# encoded = encode_jpeg(image, quality=80) # to a tensor +# write_jpeg(image, "image.jpg", quality=80) # to a file +# +# # After +# from torchcodec.encoders import JpegEncoder +# encoded = JpegEncoder(image).to_tensor(quality=80) # to a tensor +# JpegEncoder(image).to_file("image.jpg", quality=80) # to a file +# +# PNG works the same way with :class:`~torchcodec.encoders.PngEncoder` and +# ``compression_level``: + +print(f"{JpegEncoder(image).to_tensor(quality=80).shape = }") +print(f"{PngEncoder(image).to_tensor(compression_level=6).shape = }") + +# %% +# There is no batch equivalent to ``encode_jpeg(list_of_images)``: an encoder +# takes a single image, so encode a batch with a plain Python loop. You're not +# losing any speed: +# +# .. code-block:: python +# +# encoded = [JpegEncoder(image).to_tensor() for image in images] + +# %% +# Encoders also support a third destination that torchvision didn't have: a +# file-like object, i.e. anything with ``write`` and ``seek``. +import io + +buffer = io.BytesIO() +JpegEncoder(image).to_file_like(buffer) +print(f"{len(buffer.getvalue()) = }")