Zaneham/BoothPublic

Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.

AI summary: An open-source C99 compiler translating CUDA, HIP, and Triton code into CPU and multiple GPU binaries.

Stars
1.8K
Forks
92
Watchers
19
Open issues
33
Open PRs
2
Contributors
~7
Commits
289
Branches
2

CApache-2.0Created Feb 16, 2026Last push 20d agoLatest release v0.6.0+1 stars this week+15 this month

Quick answers

What is Booth?
An open-source C99 compiler translating CUDA, HIP, and Triton code into CPU and multiple GPU binaries.
What does Booth do?
Booth (formerly BarraCUDA) is a lightweight, open-source compiler that translates GPU-centric languages like CUDA C, HIP, and Triton into machine code. It acts as a versatile alternative to proprietary toolchains like nvcc, natively targeting diverse architectures including AMD RDNA, NVIDIA PTX, Tenstorrent Metalium, and standard x86-64 CPUs. Notably, it can compile complex GPU kernels, such as Triton matmuls, to execute directly on a CPU without requiring any underlying graphics hardware. Built entirely in C99, it avoids heavy LLVM dependencies and incorporates mainframe-style diagnostic features for robust debugging.
Who is Booth for?
Booth is tailored for systems programmers, high-performance computing engineers, and compiler researchers. It is an excellent fit for users seeking a lightweight, open-source toolchain to compile and test GPU kernels across a variety of hardware environments.
How do I get started with Booth?
make
How popular is Booth on GitHub?
Zaneham/Booth has 1,750 stars and 92 forks on GitHub, and gained 1 stars in the last 7 days.
What license does Booth use?
Zaneham/Booth is released under the Apache-2.0 license.

Star history

since Jul 28, 2026
05001K1.5KJul 2026Aug 2026Sep 2026Oct 2026
1.8K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Update history

1 recorded
  • Oct 4, 2026Previously tracked as Zaneham/BarraCUDA; its 1 daily snapshot and 1 trending appearance were merged into this profile. Stars: 1,721 on 2026-07-28 under the old name, 1,750 on 2026-10-04 (+29).

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 4 commits2026-02-17: 7 commits2026-02-18: 5 commits2026-02-19: 10 commits2026-02-20: 7 commits2026-02-21: 5 commits2026-02-22: 1 commit2026-02-23: 0 commits2026-02-24: 4 commits2026-02-25: 12 commits2026-02-26: 4 commits2026-02-27: 3 commits2026-02-28: 11 commits2026-03-01: 0 commits2026-03-02: 1 commit2026-03-03: 5 commits2026-03-04: 0 commits2026-03-05: 4 commits2026-03-06: 1 commit2026-03-07: 1 commit2026-03-08: 6 commits2026-03-09: 4 commits2026-03-10: 5 commits2026-03-11: 2 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 6 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 3 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 1 commit2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 1 commit2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 3 commits2026-05-20: 5 commits2026-05-21: 3 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 3 commits2026-05-25: 6 commits2026-05-26: 2 commits2026-05-27: 2 commits2026-05-28: 0 commits2026-05-29: 2 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 4 commits2026-06-04: 2 commits2026-06-05: 1 commit2026-06-06: 3 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 1 commit2026-06-15: 1 commit2026-06-16: 0 commits2026-06-17: 1 commit2026-06-18: 4 commits2026-06-19: 1 commit2026-06-20: 6 commits2026-06-21: 1 commit2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 1 commit2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 1 commit2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 1 commit2026-07-05: 1 commit2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 2 commits2026-07-09: 1 commit2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 6 commits2026-07-15: 0 commits2026-07-16: 3 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 7 commits2026-07-23: 5 commits2026-07-24: 0 commits2026-07-25: 1 commit2026-07-26: 4 commits2026-07-27: 4 commits2026-07-28: 5 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 1 commit2026-08-02: 2 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 3 commits2026-08-07: 1 commit2026-08-08: 2 commits2026-08-09: 1 commit2026-08-10: 0 commits2026-08-11: 1 commit2026-08-12: 2 commits2026-08-13: 1 commit2026-08-14: 0 commits2026-08-15: 0 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 2 commits2026-08-19: 1 commit2026-08-20: 0 commits2026-08-21: 4 commits2026-08-22: 0 commits2026-08-23: 2 commits2026-08-24: 1 commit2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 4 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 0 commits2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 0 commits2026-09-15: 4 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits2026-09-27: 0 commits2026-09-28: 0 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 0 commits2026-10-02: 0 commits2026-10-03: 0 commits
238 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

What Booth does

Booth (formerly BarraCUDA) is a lightweight, open-source compiler that translates GPU-centric languages like CUDA C, HIP, and Triton into machine code. It acts as a versatile alternative to proprietary toolchains like nvcc, natively targeting diverse architectures including AMD RDNA, NVIDIA PTX, Tenstorrent Metalium, and standard x86-64 CPUs. Notably, it can compile complex GPU kernels, such as Triton matmuls, to execute directly on a CPU without requiring any underlying graphics hardware. Built entirely in C99, it avoids heavy LLVM dependencies and incorporates mainframe-style diagnostic features for robust debugging.

Booth is tailored for systems programmers, high-performance computing engineers, and compiler researchers. It is an excellent fit for users seeking a lightweight, open-source toolchain to compile and test GPU kernels across a variety of hardware environments.

  • Multi-target compilation: Translates CUDA and Triton source code into binaries for AMD, NVIDIA, Tenstorrent, and x86-64 platforms.
  • CPU fallback execution: Enables complex GPU kernels to run natively on standard CPUs, bypassing the need for dedicated graphics hardware.
  • Minimal dependencies: Implemented purely in C99, requiring only a basic C compiler to build and eliminating the need for bulky LLVM installations.
  • Fortran support: Compiles Fortran 'do concurrent' kernels to various targets by integrating with the LFortran compiler.
  • Mainframe-style diagnostics: Generates structured ABEND dumps, SNAP outputs, and SYSPRINT logs to aid in debugging complex kernel faults.
  • Open-source transparency: Offers a fully auditable and modifiable alternative to closed-source, proprietary GPU compilation stacks.

Where teams use it

Cross-platform kernel deployment

Write a single CUDA kernel and compile it for execution on both NVIDIA and AMD graphics hardware architectures.

Hardware-independent testing

Develop and verify complex Triton workloads locally on a laptop CPU without relying on expensive GPU infrastructure.

Compiler architecture research

Examine a clean, dependency-free C99 codebase to study GPU compiler implementation details like register allocation.

Robust kernel debugging

Utilize the built-in mainframe-style crash dumps to accurately diagnose and resolve execution faults in production code.

Getting started: make

README

master branch
Booth logo

Booth

Booth is an open-source CUDA, HIP and Triton compiler targeting multiple GPU architectures, either natively by emitting machine code or as close as we can possibly get. Now with distinctly less fish.

It is named to honour Kathleen Booth: creator of the first assembly language, co-builder of the computer it first ran on, an early researcher into neural nets, a champion of women in computing, the daughter of a tax clerk, English by birth and Canadian by choice, and a mother. I believe it is fitting to name this after an incredible woman whose work this is built on. She built the machines her assembler ran on, which is exactly the bootstrapped "yeah, why not ay?" attitude this whole compiler is going for.

A running log of what's changed is in CHANGELOG.md.

Update: ggml-cuda, the CUDA half of llama.cpp, now lowers. All 67 of its files reach Booth's IR, and 47 of its kernels go through the new native SASS backend, which writes a cubin the card loads directly. Until now NVIDIA meant PTX and the driver's JIT. Try it with kath --nvidia-cubin. And I never want to have to read a machine dump ever again, lmao omg it's been rough.

What It Does

Takes CUDA C, HIP, or Triton source (the same files you'd hand to nvcc, ROCm, or Triton's JIT) and turns them into AMD RDNA 2/3/4 binaries, NVIDIA PTX, Tenstorrent Metalium C++ or native RV32IM, or just plain x86-64 you can run on a laptop with no GPU in it.

That last one still surprises me a bit. You can write a Triton kernel, matmul and all, and run it on a machine that's never seen a GPU, from scratch, no LLVM, straight to native. I haven't come across anyone else doing Triton like this, but I'd happily be proven wrong, so give me a yell if you've seen it somewhere.

Fortran do concurrent kernels go down the same path through LFortran, checked against SLATEC values in CI. See Using Fortran.

OCaml kernels go down it too, written as plain functions and type-checked by ocamlc before Booth reads the .cmt. See Using OCaml.

It also borrows a pile of operational discipline from the mainframe world: real crash dumps when a kernel faults, structured output routed by class, parameter snapshots on entry. See docs/mainframe.md if that sounds like your kind of thing.

Getting it

If you're lazy like me and just want to run and go, then you can install it here. It comes with no dependencies and you don't have to run make to use it. There's a build for Linux, macOS and Windows.

tar xzf booth-*-linux-x86_64.tar.gz
cd booth-*-linux-x86_64
./kath --version

Build

If you'd rather build it, or you're on something I don't ship a binary for:

make

That's the whole thing. You need a C99 compiler (gcc, clang, whatever you've got) and nothing else.

Two frontends read what another compiler produced, so you only need that compiler if you want that frontend. Neither is needed to build Booth:

  • Fortran kernels want LFortran, which emits the CUDA source Booth compiles the rest of the way.
  • OCaml kernels want OCaml 5.x and dune, which type-check the kernel and leave the .cmt Booth reads.
# compile a CUDA kernel to an AMD GPU binary
./kath --amdgpu-bin kernel.cu -o kernel.hsaco

# or compile and run it on whatever device you have, in one command
./kath run kernel.cu

# see what your machine can do, then self-test it
./kath doctor

The binary is kath (after Kathleen but if she picked Australia or New Zealand instead of Canada), not booth: there's already a booth in the Linux HA stack, so you'll likely end up with both on your PATH. The full command reference, every backend and flag, lives in docs/usage.md.

Documentation

  • Usage — every backend, every flag, and the runtime launcher
  • CMake — installing Booth and compiling kernels from a CMake project
  • Feature status — what compiles today, and what doesn't yet
  • Mainframe curios — ABEND dumps, SNAP, SYSPRINT, TDF
  • Validated hardware — the silicon it's been tested on, and the test suite
  • Roadmap — where it's headed
  • Contributing — style, naming, and where to help (PRs in any language welcome)

Supporting this compiler

If you're considering supporting this compiler please feel free to get in touch.

However if this compiler has been particularly helpful for you then please consider the Kate Edger Foundation which funds women in education across Auckland and Northland, named for Kate Edger, the first woman in New Zealand to earn a university degree (Latin and mathematics, 1877). If Booth's namesake means anything to you, they're well worth a look.

License

Apache 2.0. Do whatever you want. If this compiler somehow ends up in production, I'd love to hear about it, mostly so I can update my LinkedIn with something more interesting than wrote a CUDA compiler for fun.

Contact

Found a bug? Want to discuss the finer points of AMDGPU instruction encoding? Need someone to commiserate with about the state of GPU computing?

[email protected]

Open an issue if there's anything you want to discuss. Or don't. I'm not your mum.

Based in New Zealand, where it's already tomorrow and the GPUs are just as confused as everywhere else.

Acknowledgements

  • Fernando Magno Quintão Pereira and the Compilers Lab at UFMG (Universidade Federal de Minas Gerais). Fernando reached out after seeing the project, pointed me to the divergence analysis papers, and offered guidance. The SSA register allocator exists because of that conversation.
  • Jorge Galvez for sending me his do concurrent Fortran benchmarks and letting me run tests on them. Three real frontend bugs turned up in an afternoon, which is exactly what you want somebody else's code to do.
  • Jon Stevens from Hot Aisle who has very generously supported this compiler by providing access to AMD CDNA GPUs. You can find more about Hot Aisle here.
  • The academic community: Cooper, Harvey & Kennedy for dominators; Braun & Hack for SSA spilling; Sampaio, Souza, Collange & Pereira for divergence analysis. I'm just a hobbyist who reads papers and writes C. The actual hard work was done by the researchers.
  • Steven Muchnick for Advanced Compiler Design and Implementation. If this compiler does anything right, that book is why.
  • Low Level for the Zero to Hero C course and the YouTube channel. That's where I learnt C.
  • Abe Kornelis for being an amazing teacher. His work on the z390 Portable Mainframe Assembler project is well worth your time.
  • To the people who've sent messages of kindness and critique, thank you from a forever student and a happy hobbyist.
  • Lola, my sister, for the logo :-)
  • My Granny, Grandad, Nana and Baka. Love you x

He aha te mea nui o te ao. He tāngata, he tāngata, he tāngata.

What is the most important thing in the world? It is people, it is people, it is people.

View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Discussions

all 6

Releases and announcements

5 total
  1. Booth 0.6.0v0.6.0Sep 14, 202630 downloads

    Kia ora, G'day and hello! Here is Booth 0.6.0! This one has been a few months coming. You can actually run it now! `kath run kernel.cu` compiles a kernel and runs it on whatever device you have, `kath build` compiles it, and `kath doctor` tells you what your machine can do and then tests it. SASS makes an appearance with some of my own deciphering, and massive thanks to the wonderful folks at NAK, who have also reverse engineered some NVIDIA GPUs. `kath --nvidia-cubin` writes a cubin the card loads directly. Until now NVIDIA meant PTX and the driver's JIT. CUDA has yet again received way too much damn love and I don't want to look at C++ for a while. ggml-cuda, the CUDA half of llama.cpp, now compiles through Booth, woohoo! All 67 of its files reach Booth's IR, and 47 of its kernels go through the new SASS backend. Elsewhere: - `kath run`, `kath build` and `kath doctor`, with the docs updated to match. - The SASS backend lowers 32- and 64-bit division and remainder, float division, the int/float conversions, sqrt, sin, cos, exp2 and log2, min and max, and the 64-bit compares. - Class templates, explicit and partial specialisations, default template arguments, te

  2. Booth 0.5.3v0.5.3Sep 3, 2026238 downloads

    Kia ora, G'day and Hello, Here is Booth 0.5.3! There's been some big changes recently which I am happy to show ya'll! OCaml makes an appearance and you can begin using OCaml and have been using it to make a few kernels in my own spare time. Quite a bit of the code has been ported over from my other OCaml projects and luckily I've been working within the OCaml compiler. I even did a write up on my website at https://zanehambly.com/ocaml. If you'd like to see an exemplar Kernel I put some of my finance education to the test which you can see at `src/ocaml/Asian.ml`. An MLIR frontend has also been vendored in with huge thanks to @certik. The code is largely his with some edits here and there to fit in with the makefile and is released under his license. CUDA has also seen a tonne of love with the addition of multiple translation units allowing the compilation of multiple `.cu` files with one invocation. Elsewhere: - `(a) + (b)` adds again. Any parenthesised identifier was being read as a cast without checking whether it named a type, so the left operand vanished with no diagnostic. `((a) + (b))` is what every defensive macro expands to, and one of Booth's own test fixtures had b

  3. Booth 0.5.2v0.5.2Aug 8, 2026238 downloads

    Version 0.5.2. First, a correction. The last release went out tagged v5.01, which was meant to be 0.5.1 and wasn't, and it left Booth looking four major versions further along than it actually is. It isn't. This release puts the numbering back where it belongs, and sorry to anyone who pinned the old one or took the version at face value. The tag stays where it is so nothing breaks underneath you, but the compiler now reports what it is. The theme this cycle, without meaning to be, was the compiler telling the truth. Semantic errors used to be printed and then ignored by every mode except `--sema`, so the backend ran on source that had already been rejected, wrote an output file and exited zero. Asking for several backends at once wrote all of them over the same `-o` path and left you whichever finished last, under the name you chose, again exiting zero. Metal quietly narrowed a double-precision kernel to float and said nothing, which is a real problem if you were counting on the precision. `--amdgpu` ignored `-o` entirely and mixed a diagnostic into the assembly on stdout. All four are fixed, and all four had been sitting there being cheerfully wrong for a while. The structural

  4. Booth 5.01v5.01Jul 14, 2026

    So long, and thanks for all the fish. BarraCUDA was a good pun and a bad description, so it swam off: the compiler is Booth now, the binary is `kath`, and the version jumps to mark the line. This is the first release under the new name. Since 0.5 the compiler grew a spine on the CPU side. x86-64 and RV64 both gained most of their scalar arithmetic, so a CUDA, HIP or Triton kernel now compiles and runs on a machine with no GPU in it at all. Tensix learned to emit native machine code for the baby RISC-V cores instead of leaning on a C++ handoff. `__device__` calls are inlined away before isel, so device helpers with control flow finally work on the GPU and vector backends. Diagnostics were rebuilt to read like Clang's, carets and colour and all, and the frontend picked up a pile of coverage. Thanks to the people who sent patches this cycle: @GauthamMK-0 for `__hip_bfloat16` and the `warpSize` builtin, @nataliakokoromyti for lowering `tl.where` to select and keeping the PTX header ASCII, and @kstppd for fixing the Makefile on non-x86 machines. Much appreciated.

  5. BarraCUDA 0.5v0.5.0May 29, 2026

    # BarraCUDA 0.5 The first tagged release. The headline is that you can write a Triton kernel, matmul and all, and run it on a CPU with no GPU. The `--cpu` backend lowers BIR straight to x86-64 with the SIMT model collapsed into a thread loop, and the rank-2 tile path materialises and unrolls so `tl.dot` plus a K-loop sweeps an arbitrary contraction. ## New in this cycle - **CPU backend (`--cpu`)**. CUDA and Triton kernels compile to a host object and run natively. Headline demo: `examples/cpu_launch_matmul.c`. - **RISC-V backend (`--rv64`)**. Same idea, RV64IMFD objects that run under qemu. - **Cross-backend differential testing** (`tests/diff/`). Same BIR through two backends, diff the output buffers, CPU is the oracle. Every case runs `--inject` so a green result actually means something. - **Triton scalar math intrinsics**. `exp`, `log`, `sin`, `cos`, `tan`, `tanh`, `sqrt`, `rsqrt`, `abs`, `floor`, `ceil`, `maximum`, `minimum`, `fdiv`. Thanks to @shivam2931120 for the PR, radians-to-turns convention done right. - **Triton constexpr ABI compaction**. `tl.constexpr` params with defaults fold to literals and drop out of the runtime signature. - **CUDA fi

Code frequency

additions and deletions
+27.8K-27.8KWeek of 2026-02-15: +20,729 linesWeek of 2026-02-15: -386 linesWeek of 2026-02-22: +22,589 linesWeek of 2026-02-22: -15,924 linesWeek of 2026-03-01: +7,361 linesWeek of 2026-03-01: -2,996 linesWeek of 2026-03-08: +6,293 linesWeek of 2026-03-08: -752 linesWeek of 2026-03-15: +3,608 linesWeek of 2026-03-15: -54 linesWeek of 2026-03-22: +119 linesWeek of 2026-03-22: -20 linesWeek of 2026-03-29: +0 linesWeek of 2026-03-29: -0 linesWeek of 2026-04-05: +0 linesWeek of 2026-04-05: -0 linesWeek of 2026-04-12: +0 linesWeek of 2026-04-12: -0 linesWeek of 2026-04-19: +45 linesWeek of 2026-04-19: -1 linesWeek of 2026-04-26: +0 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +6,765 linesWeek of 2026-05-17: -322 linesWeek of 2026-05-24: +12,467 linesWeek of 2026-05-24: -434 linesWeek of 2026-05-31: +3,693 linesWeek of 2026-05-31: -294 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +2,004 linesWeek of 2026-06-14: -938 linesWeek of 2026-06-21: +75 linesWeek of 2026-06-21: -5 linesWeek of 2026-06-28: +191 linesWeek of 2026-06-28: -3 linesWeek of 2026-07-05: +670 linesWeek of 2026-07-05: -725 linesWeek of 2026-07-12: +1,412 linesWeek of 2026-07-12: -827 linesWeek of 2026-07-19: +1,981 linesWeek of 2026-07-19: -415 linesWeek of 2026-07-26: +2,930 linesWeek of 2026-07-26: -713 linesWeek of 2026-08-02: +3,015 linesWeek of 2026-08-02: -473 linesWeek of 2026-08-09: +27,837 linesWeek of 2026-08-09: -854 linesWeek of 2026-08-16: +5,504 linesWeek of 2026-08-16: -46 linesWeek of 2026-08-23: +3,932 linesWeek of 2026-08-23: -4,209 linesWeek of 2026-08-30: +0 linesWeek of 2026-08-30: -0 linesFeb 15, 2026Aug 30, 2026
+133.2K lines added, -30.4K removed over the last year.

Commits per week

last 52 weeks
380Week of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 38 commitsWeek of 2026-02-22: 35 commitsWeek of 2026-03-01: 12 commitsWeek of 2026-03-08: 23 commitsWeek of 2026-03-15: 3 commitsWeek of 2026-03-22: 1 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 1 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 11 commitsWeek of 2026-05-24: 15 commitsWeek of 2026-05-31: 10 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 14 commitsWeek of 2026-06-21: 2 commitsWeek of 2026-06-28: 2 commitsWeek of 2026-07-05: 4 commitsWeek of 2026-07-12: 9 commitsWeek of 2026-07-19: 13 commitsWeek of 2026-07-26: 14 commitsWeek of 2026-08-02: 8 commitsWeek of 2026-08-09: 5 commitsWeek of 2026-08-16: 7 commitsWeek of 2026-08-23: 3 commitsWeek of 2026-08-30: 4 commitsWeek of 2026-09-06: 0 commitsWeek of 2026-09-13: 4 commitsWeek of 2026-09-20: 0 commitsWeek of 2026-09-27: 0 commitsOct 4, 2025Sep 27, 2026
238 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 1 commitsSun 8:00 — 0 commitsSun 9:00 — 2 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 1 commitsSun 13:00 — 2 commitsSun 14:00 — 2 commitsSun 15:00 — 0 commitsSun 16:00 — 2 commitsSun 17:00 — 1 commitsSun 18:00 — 3 commitsSun 19:00 — 6 commitsSun 20:00 — 2 commitsSun 21:00 — 0 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 2 commitsMon 2:00 — 4 commitsMon 3:00 — 1 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 1 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 0 commitsMon 11:00 — 2 commitsMon 12:00 — 1 commitsMon 13:00 — 2 commitsMon 14:00 — 1 commitsMon 15:00 — 0 commitsMon 16:00 — 2 commitsMon 17:00 — 0 commitsMon 18:00 — 0 commitsMon 19:00 — 0 commitsMon 20:00 — 1 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 4 commitsTue 0:00 — 4 commitsTue 1:00 — 1 commitsTue 2:00 — 2 commitsTue 3:00 — 4 commitsTue 4:00 — 0 commitsTue 5:00 — 1 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 1 commitsTue 9:00 — 0 commitsTue 10:00 — 1 commitsTue 11:00 — 5 commitsTue 12:00 — 3 commitsTue 13:00 — 2 commitsTue 14:00 — 3 commitsTue 15:00 — 5 commitsTue 16:00 — 2 commitsTue 17:00 — 0 commitsTue 18:00 — 1 commitsTue 19:00 — 4 commitsTue 20:00 — 1 commitsTue 21:00 — 0 commitsTue 22:00 — 1 commitsTue 23:00 — 3 commitsWed 0:00 — 1 commitsWed 1:00 — 1 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 1 commitsWed 10:00 — 0 commitsWed 11:00 — 0 commitsWed 12:00 — 5 commitsWed 13:00 — 1 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 4 commitsWed 17:00 — 0 commitsWed 18:00 — 6 commitsWed 19:00 — 5 commitsWed 20:00 — 4 commitsWed 21:00 — 7 commitsWed 22:00 — 2 commitsWed 23:00 — 11 commitsThu 0:00 — 2 commitsThu 1:00 — 2 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 1 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 0 commitsThu 11:00 — 0 commitsThu 12:00 — 0 commitsThu 13:00 — 1 commitsThu 14:00 — 0 commitsThu 15:00 — 2 commitsThu 16:00 — 3 commitsThu 17:00 — 0 commitsThu 18:00 — 1 commitsThu 19:00 — 2 commitsThu 20:00 — 4 commitsThu 21:00 — 9 commitsThu 22:00 — 9 commitsThu 23:00 — 7 commitsFri 0:00 — 0 commitsFri 1:00 — 1 commitsFri 2:00 — 1 commitsFri 3:00 — 3 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 1 commitsFri 8:00 — 0 commitsFri 9:00 — 2 commitsFri 10:00 — 0 commitsFri 11:00 — 1 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 1 commitsFri 15:00 — 1 commitsFri 16:00 — 1 commitsFri 17:00 — 2 commitsFri 18:00 — 1 commitsFri 19:00 — 0 commitsFri 20:00 — 1 commitsFri 21:00 — 2 commitsFri 22:00 — 1 commitsFri 23:00 — 3 commitsSat 0:00 — 3 commitsSat 1:00 — 3 commitsSat 2:00 — 2 commitsSat 3:00 — 2 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 1 commitsSat 9:00 — 1 commitsSat 10:00 — 6 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 2 commitsSat 15:00 — 1 commitsSat 16:00 — 0 commitsSat 17:00 — 10 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 1 commitsSat 21:00 — 5 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits262 (91%)
Community commits27 (9%)

289 commits in total over the last year.

DateListRankStars gained
Feb 18, 2026daily#15+186
  • tensorflow/tensorflow

    An Open Source Machine Learning Framework for Everyone

    200.7K stars · C++

  • Genymobile/scrcpy

    Display and control your Android device

    151K stars · C

  • vercel/next.js

    The React Framework

    143.2K stars · JavaScript

  • microsoft/PowerToys

    Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows

    139.2K stars · C

  • pytorch/pytorch

    Tensors and Dynamic neural networks in Python with strong GPU acceleration

    103.8K stars · Python

  • colbymchenry/codegraph

    Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, CoPilot, and Hermes Agent — fewer tokens, fewer tool calls, 100% local

    73.2K stars · C