PythonApache-2.0Created Feb 20, 2025Last push 4d agoLatest release v0.6.0+-19 stars this week+-10 this month
Quick answers
What is chitu?
A high-performance LLM inference framework optimized for efficiency and diverse hardware compatibility, including Chinese accelerators.
What does chitu do?
Chitu is a production-grade inference engine engineered to serve massive Large Language Models (LLMs) efficiently across a wide spectrum of computational hardware. It provides robust support for mainstream NVIDIA GPUs alongside deep, native integration with domestic Chinese AI accelerators like Ascend, Moore Threads, and Hygon. The framework is highly scalable, handling everything from heterogeneous CPU+GPU execution on a single machine to complex multi-node enterprise cluster deployments. It employs advanced optimizations like highly efficient FP4 and FP8 online quantization operators to minimize memory footprint and maximize token throughput for models as massive as 671B parameters.
Who is chitu for?
AI engineers, infrastructure architects, and researchers deploying massive LLMs in production, particularly on Chinese domestic hardware.
How do I get started with chitu?
Download the chitu.run executable from the Releases page.
How popular is chitu on GitHub?
thu-pacman/chitu has 2,987 stars and 246 forks on GitHub, and gained -19 stars in the last 7 days.
What license does chitu use?
thu-pacman/chitu is released under the Apache-2.0 license.
Star history
since Mar 9, 2025
3K stars as of Oct 3, 2026. Before Jul 29, 2026, reconstructed from public GitHub event archives (checked against the repository's real star total); since then measured daily.
Contribution activity
commits per day, last 52 weeks
753 commits in the last yearLessMore
Signals and awards
derived from tracked data
Very active
753 commits in 52 weeks
Permissive license
Apache-2.0
Continuous integration
Automated checks passing
What chitu does
Chitu is a production-grade inference engine engineered to serve massive Large Language Models (LLMs) efficiently across a wide spectrum of computational hardware. It provides robust support for mainstream NVIDIA GPUs alongside deep, native integration with domestic Chinese AI accelerators like Ascend, Moore Threads, and Hygon. The framework is highly scalable, handling everything from heterogeneous CPU+GPU execution on a single machine to complex multi-node enterprise cluster deployments. It employs advanced optimizations like highly efficient FP4 and FP8 online quantization operators to minimize memory footprint and maximize token throughput for models as massive as 671B parameters.
AI engineers, infrastructure architects, and researchers deploying massive LLMs in production, particularly on Chinese domestic hardware.
Broad hardware compatibility: Natively supports NVIDIA GPUs alongside Chinese AI accelerators like Huawei Ascend and Moore Threads.
Heterogeneous execution: Enables running massive LLMs on single machines by mixing CPU and GPU compute resources dynamically.
Advanced quantization operations: Features highly optimized operators for online conversion between FP4, FP8, and BF16 formats to save VRAM.
Cluster deployment scalability: Built for production environments, scaling seamlessly from local test instances to large enterprise deployments.
Massive model support: Optimized specifically to handle the inference of extremely large parameter models, including the 671B DeepSeek-R1.
Where teams use it
Enterprise Model Serving
AI engineering teams deploy large foundational models stably in high-traffic, production-grade API environments.
Domestic Hardware Migration
Organizations seamlessly migrate and execute complex LLMs on domestic Chinese hardware architectures like the Ascend 910B.
Low-Resource Massive Inference
Researchers run extremely large models on limited hardware through efficient CPU+GPU heterogeneous processing.
High-Throughput Processing
Infrastructure teams leverage advanced FP8 quantization to maximize the token throughput of their LLM service clusters.
Getting started: Download the chitu.run executable from the Releases page.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.