higgsfield-ai/higgsfieldPublic

Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters

AI summary: A fault-tolerant GPU orchestration and training framework for scaling trillion-parameter models across multiple nodes.

Stars
5.8K
+38 today
Forks
1K
Watchers
90
Open issues
7
Open PRs
9
Contributors
~4
Commits
44
Branches
7

Jupyter NotebookApache-2.0Created May 26, 2018Last push 20d agoLatest release v0.0.4-rc+183 stars this week+928 this month

Quick answers

What is higgsfield?
A fault-tolerant GPU orchestration and training framework for scaling trillion-parameter models across multiple nodes.
What does higgsfield do?
Higgsfield addresses the immense complexity of distributed machine learning by combining GPU resource management with a highly optimized training framework. It orchestrates exclusive node access and manages the inevitable hardware failures that occur during massive training runs. Under the hood, it deeply integrates with PyTorch's Fully Sharded Data Parallel and DeepSpeed ZeRO-3 to partition model states. This ensures efficient execution of foundational models while abstracting away the boilerplate of cluster networking and state recovery.
Who is higgsfield for?
ML infrastructure engineers and AI researchers training massive models. Requires access to multi-node GPU clusters and advanced PyTorch knowledge.
How do I get started with higgsfield?
pip install higgsfield
How popular is higgsfield on GitHub?
higgsfield-ai/higgsfield has 5,844 stars and 1,048 forks on GitHub, and gained 183 stars in the last 7 days.
What license does higgsfield use?
higgsfield-ai/higgsfield is released under the Apache-2.0 license.

Star history

since Sep 19, 2026
02K4KSep 2026Sep 2026Sep 2026Oct 2026
5.8K stars as of Oct 2, 2026. Measured daily since Sep 19, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks

Signals and awards

derived from tracked data
  • Battle-tested

    8 years of history

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    3 trending appearances

What higgsfield does

Higgsfield addresses the immense complexity of distributed machine learning by combining GPU resource management with a highly optimized training framework. It orchestrates exclusive node access and manages the inevitable hardware failures that occur during massive training runs. Under the hood, it deeply integrates with PyTorch's Fully Sharded Data Parallel and DeepSpeed ZeRO-3 to partition model states. This ensures efficient execution of foundational models while abstracting away the boilerplate of cluster networking and state recovery.

ML infrastructure engineers and AI researchers training massive models. Requires access to multi-node GPU clusters and advanced PyTorch knowledge.

  • Cluster orchestration: Manages exclusive and non-exclusive GPU node allocations for multi-tenant training environments.
  • Advanced sharding: Wraps DeepSpeed ZeRO-3 and FSDP to efficiently distribute model parameters across thousands of GPUs.
  • Fault tolerance: Automatically detects node failures and resumes training jobs to prevent catastrophic loss of progress.
  • Resource management: Handles hardware contention and ensures optimal utilization of expensive compute clusters.
  • Lifecycle monitoring: Provides a unified framework for initiating, executing, and tracking the telemetry of massive training runs.

Where teams use it

Foundation Model Training

Scale the pre-training of billion-parameter language models across dozens of nodes without writing custom networking code.

Cluster Management

Allocate GPU resources efficiently among a team of AI researchers working on concurrent distributed jobs.

Resilient Workloads

Ensure long-running machine learning tasks automatically recover from transient network or hardware failures.

Model Checkpointing

Manage the saving and loading of massive distributed model states seamlessly.

Getting started: pip install higgsfield

README

main branch

higgsfield - multi node training without crying

Higgsfield is an open-source, fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters, such as Large Language Models (LLMs).

PyPI version

architecture

Higgsfield serves as a GPU workload manager and machine learning framework with five primary functions:

  1. Allocating exclusive and non-exclusive access to compute resources (nodes) to users for their training tasks.
  2. Supporting ZeRO-3 deepspeed API and fully sharded data parallel API of PyTorch, enabling efficient sharding for trillion-parameter models.
  3. Offering a framework for initiating, executing, and monitoring the training of large neural networks on allocated nodes.
  4. Managing resource contention by maintaining a queue for running experiments.
  5. Facilitating continuous integration of machine learning development through seamless integration with GitHub and GitHub Actions. Higgsfield streamlines the process of training massive models and empowers developers with a versatile and robust toolset.

Install

$ pip install higgsfield==0.0.3

Train example

That's all you have to do in order to train LLaMa in a distributed setting:

from higgsfield.llama import Llama70b
from higgsfield.loaders import LlamaLoader
from higgsfield.experiment import experiment

import torch.optim as optim
from alpaca import get_alpaca_data

@experiment("alpaca")
def train(params):
    model = Llama70b(zero_stage=3, fast_attn=False, precision="bf16")

    optimizer = optim.AdamW(model.parameters(), lr=1e-5, weight_decay=0.0)

    dataset = get_alpaca_data(split="train")
    train_loader = LlamaLoader(dataset, max_words=2048)

    for batch in train_loader:
        optimizer.zero_grad()
        loss = model(batch)
        loss.backward()
        optimizer.step()

    model.push_to_hub('alpaca-70b')

How it's all done?

  1. We install all the required tools in your server (Docker, your project's deploy keys, higgsfield binary).
  2. Then we generate deploy & run workflows for your experiments.
  3. As soon as it gets into Github, it will automatically deploy your code on your nodes.
  4. Then you access your experiments' run UI through Github, which will launch experiments and save the checkpoints.

Design

We follow the standard pytorch workflow. Thus you can incorporate anything besides what we provide, deepspeed, accelerate, or just implement your custom pytorch sharding from scratch.

Enviroment hell

No more different versions of pytorch, nvidia drivers, data processing libraries. You can easily orchestrate experiments and their environments, document and track the specific versions and configurations of all dependencies to ensure reproducibility.

Config hell

No need to define 600 arguments for your experiment. No more yaml witchcraft. You can use whatever you want, whenever you want. We just introduce a simple interface to define your experiments. We have even taken it further, now you only need to design the way to interact.

Compatibility

We need you to have nodes with:

  • Ubuntu
  • SSH access
  • Non-root user with sudo privileges (no-password is required)

Clouds we have tested on:

  • Azure
  • LambdaLabs
  • FluidStack

Feel free to open an issue if you have any problems with other clouds.

Getting started

Here you can find the quick start guide on how to setup your nodes and start training.

API for common tasks in Large Language Models training.

Platform Purpose Estimated Response Time Support Level
Github Issues Bug reports, feature requests, install issues, usage issues, etc. < 1 day Higgsfield Team
Twitter For staying up-to-date on new features. Daily Higgsfield Team
Website Discussion, news. < 2 days Higgsfield Team
View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Releases and announcements

1 total
  1. v0.0.4-rcv0.0.4-rcMar 23, 2024

    ## What's Changed * fix(action-builder): default field gen by @higgsfield in https://github.com/higgsfield-ai/higgsfield/pull/33 * add: specify invoker version if needed by @higgsfield in https://github.com/higgsfield-ai/higgsfield/pull/34 * build(deps): bump asyncssh from 2.14.0 to 2.14.1 by @dependabot in https://github.com/higgsfield-ai/higgsfield/pull/35 * merge: dev by @arpanetus in https://github.com/higgsfield-ai/higgsfield/pull/36 * Feat/invoker exec of master host of no python by @arpanetus in https://github.com/higgsfield-ai/higgsfield/pull/42 ## New Contributors * @higgsfield made their first contribution in https://github.com/higgsfield-ai/higgsfield/pull/33 * @arpanetus made their first contribution in https://github.com/higgsfield-ai/higgsfield/pull/36 **Full Changelog**: https://github.com/higgsfield-ai/higgsfield/commits/v0.0.4-rc

Commits per week

last 52 weeks

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 1 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 0 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 0 commitsSun 22:00 — 0 commitsSun 23:00 — 4 commitsMon 0:00 — 2 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 0 commitsMon 11:00 — 0 commitsMon 12:00 — 1 commitsMon 13:00 — 3 commitsMon 14:00 — 0 commitsMon 15:00 — 1 commitsMon 16:00 — 0 commitsMon 17:00 — 0 commitsMon 18:00 — 0 commitsMon 19:00 — 0 commitsMon 20:00 — 0 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 0 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 0 commitsTue 12:00 — 0 commitsTue 13:00 — 0 commitsTue 14:00 — 0 commitsTue 15:00 — 0 commitsTue 16:00 — 0 commitsTue 17:00 — 0 commitsTue 18:00 — 0 commitsTue 19:00 — 0 commitsTue 20:00 — 1 commitsTue 21:00 — 1 commitsTue 22:00 — 0 commitsTue 23:00 — 0 commitsWed 0:00 — 1 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 1 commitsWed 6:00 — 1 commitsWed 7:00 — 1 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 0 commitsWed 11:00 — 0 commitsWed 12:00 — 0 commitsWed 13:00 — 0 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 0 commitsWed 17:00 — 0 commitsWed 18:00 — 0 commitsWed 19:00 — 0 commitsWed 20:00 — 0 commitsWed 21:00 — 0 commitsWed 22:00 — 0 commitsWed 23:00 — 0 commitsThu 0:00 — 0 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 1 commitsThu 4:00 — 1 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 0 commitsThu 11:00 — 0 commitsThu 12:00 — 0 commitsThu 13:00 — 0 commitsThu 14:00 — 0 commitsThu 15:00 — 0 commitsThu 16:00 — 0 commitsThu 17:00 — 0 commitsThu 18:00 — 1 commitsThu 19:00 — 0 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 0 commitsThu 23:00 — 0 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 1 commitsFri 8:00 — 6 commitsFri 9:00 — 6 commitsFri 10:00 — 0 commitsFri 11:00 — 0 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 0 commitsFri 15:00 — 0 commitsFri 16:00 — 0 commitsFri 17:00 — 0 commitsFri 18:00 — 3 commitsFri 19:00 — 1 commitsFri 20:00 — 0 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Sep 21, 2026daily#7+196
Sep 20, 2026daily#7+196
Sep 19, 2026daily#6+325
  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    272.8K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • tensorflow/tensorflow

    An Open Source Machine Learning Framework for Everyone

    200.7K stars · C++

  • firecrawl/firecrawl

    Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥

    188.6K stars · TypeScript

  • Significant-Gravitas/AutoGPT

    AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

    187.7K stars · Python

  • f/prompts.chat

    f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.

    172K stars · HTML