exllamav2

An efficient inference back-end for consumer GPUs with strong quantisation support

What it is

An inference back-end optimised for consumer GPUs with solid quantisation support and memory efficiency, able to fit larger models on a single ordinary card. It is a common choice among local enthusiasts.

Highlights

  • Rich quantisation support with efficient memory use
  • Larger models fit on one consumer card
  • Fast inference with low latency
  • Integrated as a back-end by several local UIs

How to use

Install it, download matching quantised weights, then load via script to run the model locally.

License

Released under MIT. Read the terms before commercial use or redistribution, especially if you plan to offer it as a service.