llama.cpp

A pure C++ inference engine that puts models on ordinary computers

What it is

An inference engine written in C and C++ with no heavy framework dependencies, running quantised models on CPUs and consumer GPUs. It is one of the most important foundations for running open models locally.

Highlights

  • Pure C++ with minimal dependencies and strong portability
  • Supports CPU and multiple acceleration backends
  • Its quantisation format became a de facto standard
  • Many local apps are built on top of it

How to use

Compile from source or grab a release, download quantised weights, then run from the command line or a local server.

License

Released under MIT. Read the terms before commercial use or redistribution, especially if you plan to offer it as a service.