higgsfield
View on GitHubFault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
Higgsfield is an open-source GPU orchestration and ML training framework for multi-node training of billion-to-trillion parameter LLMs. It wraps PyTorch FSDP and DeepSpeed ZeRO-3 with node allocation, experiment queuing, monitoring, and GitHub Actions-driven deployment across cloud nodes.
Use Cases
Distributed multi-node training of billion-to-trillion parameter LLMsManaging GPU node allocation and exclusive/non-exclusive accessRunning fault-tolerant LLM training jobs across cloud nodesQueuing and scheduling resource-contended training experimentsSaving and pushing trained model checkpoints to Hugging Face HubCI/CD of ML training pipelines via GitHub and GitHub ActionsReproducible experiment environment and dependency trackingSharding large models with ZeRO-3 or FSDP
Built With
- Language
- Jupyter Notebook
- Frameworks
- PyTorch · DeepSpeed ZeRO-3 · PyTorch FSDP · Poetry · Click · asyncssh
Tags
distributed-training · gpu-orchestration · llm-training · pytorch · deepspeed · zero-3 · fsdp · fault-tolerance · cluster-management · mlops · multi-node · sharding · llama · training-framework · github-actions · checkpointing