Cerebras-GPT
Open-source, Apache-2.0 family of seven compute-optimal GPT models (111M-13B) trained on the Cerebras cluster.
What it does
Cerebras-GPT is a family of seven open GPT language models ranging from 111M to 13B parameters, released under Apache 2.0. Trained with the Chinchilla compute-optimal recipe on the public Pile dataset, they offer high accuracy per unit of compute. Weights and checkpoints are on Hugging Face and GitHub.
Core features
Seven model sizes from 111M to 13B parameters
Apache 2.0 license for research and commercial use
Chinchilla compute-optimal training
Trained on the public Pile dataset
Weights and checkpoints on Hugging Face/GitHub
New open scaling law
Best for
→Research on LLM scaling laws
→Building on open, royalty-free base models
→Reproducible compute-efficient training
Tutorials
Step-by-step: exactly how to get things done with it.