llama.cpp
High-performance local inference engine for LLMs and multimodal models, supporting CPU/GPU execution, quantization, and broad model compatibility.
Freemium Free tier available
Learn More - Category
- Agents & Automation
- Pricing
- Freemium
- Website
- github.com
About llama.cpp
High-performance local inference engine for LLMs and multimodal models, supporting CPU/GPU execution, quantization, and broad model compatibility.