Trending repositories: llm-inference
3 tracked repositories tagged with llm-inference, ordered by stars. Use the topic filters below to narrow further.
Filter by topic
3 of 3 repositories
lyogavin/airllm
AirLLM 70B inference with single 4GB GPU
AI summary: Run massive language models like 70B on a single 4GB GPU using layer-by-layer inference.
29,821ai-mlJupyter NotebookApache-2.0TheTom/turboquant_plus
AI summary: Advanced KV cache compression techniques merged into major LLM inference engines.
7,002ai-mlPythonApache-2.00xSojalSec/airllm
Runs 405B LLMs on 8GB VRAM
AI summary: Memory optimization framework that runs massive LLMs like Llama3.1 405B on low-VRAM GPUs.
3,054ai-mlJupyter NotebookApache-2.0