What it is
An inference engine written in C and C++ with no heavy framework dependencies, running quantised models on CPUs and consumer GPUs. It is one of the most important foundations for running open models locally.
Highlights
- Pure C++ with minimal dependencies and strong portability
- Supports CPU and multiple acceleration backends
- Its quantisation format became a de facto standard
- Many local apps are built on top of it
How to use
Compile from source or grab a release, download quantised weights, then run from the command line or a local server.
License
Released under MIT. Read the terms before commercial use or redistribution, especially if you plan to offer it as a service.