DeepSpeed

A training optimisation library for huge models covering memory and parallelism

What it is

Microsoft’s open training optimisation library makes very large model training possible through memory sharding, pipeline parallelism and related techniques. It is an important component of training infrastructure.

Highlights

  • Memory sharding and parallelism lower the barrier
  • Supports training and fine-tuning very large models
  • Thorough optimiser and communication optimisation
  • Integrated by many training frameworks

How to use

Install it, configure strategies to match your hardware, then launch distributed training.

License

Released under Apache-2.0. Read the terms before commercial use or redistribution, especially if you plan to offer it as a service.