Sign inSign up

somewatson/ffmpeg-ampere-n1

By somewatson

•Updated 3 months ago

This image provides a build of FFmpeg optimized specifically for the ARM Neoverse-N1 architecture.

Image
0

1.2K

somewatson/ffmpeg-ampere-n1 repository overview

⁠Github

Repository: github.com/somewatson/ffmpeg-ampere-optimized⁠

⁠FFmpeg Ampere Neoverse-N1 Optimized Image

This image provides a build of FFmpeg optimized specifically for the ARM Neoverse-N1 architecture, following the Ampere Computing tuning guidelines.

Source: Ampere FFmpeg Tuning Guide⁠

⁠Build the Image

Build the image from the local Dockerfile:

docker build -t ffmpeg-ampere-n1 .

Note on High Core Counts: This Dockerfile uses $(nproc) to maximize parallelism. On machines with a very high number of cores (e.g., 128+), ensure you have sufficient RAM (roughly 2GB per core) to avoid Out-Of-Memory (OOM) errors during compilation. If the build fails due to memory, you may need to limit the parallelism by replacing $(nproc) with a fixed number in the Dockerfile.

⁠Usage

The image is configured with ffmpeg as the entrypoint.

⁠Check Version

Verify the installation and optimization:

docker run --rm ffmpeg-ampere-n1 -version
⁠Basic Transcoding

To transcode a video file, mount your local media directory to the container. We recommend using --shm-size=2g and --privileged for maximum performance on Ampere N1:

docker run --rm --shm-size=2g --privileged -v $(pwd):/media ffmpeg-ampere-n1 \
  -i /media/input.mp4 \
  -c:v libx264 \
  -preset medium \
  -crf 23 \
  -c:a aac \
  /media/output.mp4
⁠H.265 (HEVC) Encoding

Utilizing the optimized libx265:

docker run --rm --shm-size=2g --privileged -v $(pwd):/media ffmpeg-ampere-n1 \
  -i /media/input.mp4 \
  -c:v libx265 \
  -crf 28 \
  /media/output_hevc.mp4
⁠AV1 Encoding

Utilizing libsvtav1 for high-performance, scalable encoding.

Note: This image uses SVT-AV1 instead of the reference libaom-av1 implementation. SVT-AV1 is specifically designed for massive multi-core parallelism, making it significantly faster and more efficient on high-core-count systems (e.g., 128+ cores).

docker run --rm --shm-size=2g --privileged -v $(pwd):/media ffmpeg-ampere-n1 \
  -i /media/input.mp4 \
  -c:v libsvtav1 \
  -crf 30 \
  -preset 6 \
  /media/output_av1.mp4

⁠Large-Scale Encoding (Chunked Parallelism)

For maximum throughput on high-core-count systems, use chunked encoding. This process splits the input into segments, encodes them in parallel across multiple FFmpeg instances, and concatenates the results.

Use the provided helper script:

chmod +x chunked_encode.sh
./chunked_encode.sh <input_file> <output_file> <codec> <crf> <preset> <chunks> <image_tag>

# Example: Split into 10 chunks using libx265
./chunked_encode.sh input.mp4 output.mp4 libx265 28 slow 10 ffmpeg-ampere-n1

⁠Benchmarks

Performance comparison between a generic FFmpeg image and the optimized ffmpeg-ampere-n1 image on Ampere Neoverse-N1 architecture.

Quality Metric (PSNR): PSNR (Peak Signal-to-Noise Ratio) measures reconstruction quality.

  • > 40 dB: Excellent (imperceptible difference from source)
  • 30-40 dB: Good to Very Good
  • < 30 dB: Noticeable quality loss

Test Configuration:

  • Source: BigBuckBunny_512kb.mp4
  • Settings: CRF 23, Preset Slow (or Preset 8 for SVT-AV1)
  • Runtime: --ipc=host, --privileged
⁠Standard Mode
ImageCodecTime (s)Size (KB)PSNR (dB)FPS
Genericlibx26451.8426,09241.36286.00
Optimizedlibx26450.9326,09241.36286.00
Genericlibx265312.6217,26839.4446.00
Optimizedlibx265253.0117,42039.3657.00
Genericlibsvtav136.6229,18841.94412.00
Optimizedlibsvtav128.8229,23241.95515.00

Standard Mode Performance Improvement:

  • libx264: 1.00% faster (Speedup: 0.91s)
  • libx265: 19.00% faster (Speedup: 59.61s)
  • libsvtav1: 21.00% faster (Speedup: 7.80s)
⁠Chunked Mode (Parallelism)
ImageCodecTime (s)Size (KB)PSNR (dB)FPS
Genericlibx26420.5826,11641.36280.00
Optimizedlibx26418.3426,11641.36279.25
Genericlibx26599.4417,28839.4444.50
Optimizedlibx26577.5717,44039.3656.00
Genericlibsvtav115.8229,22441.94391.75
Optimizedlibsvtav113.5729,24441.95475.50

Chunked Mode Performance Improvement:

  • libx264: 10.00% faster (Speedup: 2.24s)
  • libx265: 21.00% faster (Speedup: 21.87s)
  • libsvtav1: 14.00% faster (Speedup: 2.25s)

⁠Optimizations applied

  • Target CPU: -mcpu=neoverse-n1 (Optimized for the Ampere Neoverse-N1 architecture)
  • Link-Time Optimization: -flto=auto enabled across all libraries and FFmpeg for improved inter-procedural optimization.
  • Compiler Flags:
    • Standard optimization level (Default -O2) used for stability and peak performance on N1.
  • Libraries: libx264, libx265, libvpx, libsvtav1
  • Architecture: Built specifically for ARM64 / Ampere N1.

⁠Credits

This project was created by Some Watson⁠ with the assistance of opencode, an AI software engineering agent.

Tag summary

Content type

Image

Digest

sha256:8eca7aef3…

Size

248.4 MB

Last updated

3 months ago

docker pull somewatson/ffmpeg-ampere-n1