PythonApache-2.0Created Feb 20, 2025Last push 2d agoLatest release v0.6.0+16 stars this week+16 this month
Star history
since Mar 9, 2025
3.2K stars as of Aug 6, 2026, tracked back to Mar 9, 2025. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.
Contribution activity
commits per day, last 52 weeks
859 commits in the last yearLessMore
Signals and awards
derived from tracked data
Very active
859 commits in 52 weeks
Permissive license
Apache-2.0
Continuous integration
Automated checks passing
What chitu does
Chitu is a production-grade Large Language Model inference framework designed for efficiency and flexibility. It provides comprehensive support for diverse hardware, including NVIDIA GPUs and various domestic Chinese AI chips (like Ascend and Moore Threads). The engine scales from single-GPU setups to massive clusters and features advanced optimizations such as FP4 online quantization and heterogeneous CPU+GPU inference. It is built to maintain long-term stability in high-concurrency enterprise environments.
AI infrastructure engineers and enterprise teams looking to deploy LLMs efficiently at scale, particularly those utilizing diverse hardware ecosystems. Requires Linux and supported accelerators.
Broad Hardware Support: Adapts to NVIDIA, Huawei Ascend, Moore Threads, and other heterogeneous chips.
Advanced Quantization: Supports efficient online conversion of FP4 to FP8/BF16.
Heterogeneous Inference: Enables running massive models like DeepSeek-R1 on single nodes using CPU+GPU memory.
Cluster Scalability: Provides a unified `chitu.run` executable for deploying across complex multi-node clusters.
Production Stability: Designed to handle continuous, high-concurrency enterprise workloads without crashing.
Where teams use it
Enterprise LLM Deployment
Serve models like Qwen or DeepSeek stably to thousands of concurrent users.
Domestic Hardware Integration
Run large models efficiently on Chinese AI accelerators like Huawei Ascend 910B.
Massive Model Inference
Use CPU+GPU hybrid inference to run a 671B model on limited hardware.
Cluster Management
Easily spin up multi-instance deployments across multiple physical servers.
Getting started: Download the `chitu.run` executable from the Releases page and follow the deployment manual.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.