Run a 70B AI Model on a 4GB Graphics Card: Set Up the Free AirLLM Library
Big open AI models need an expensive graphics card, so most people never try one. AirLLM is a free, open-source Python library that runs a 70 billion parameter model on a single 4GB GPU by keeping only one layer of it in memory at a time. The install command, the code to paste in, ready-made prompts for Claude Code, and the model sizes from its own README
A 70 billion parameter model is about 140GB at 16-bit precision (70 billion numbers, two bytes each). That is why running one has meant buying a very expensive graphics card, and why most people never tried.
AirLLM is a free, open-source library (Apache-2.0 licence, over 35,000 stars on GitHub) that gets around the problem. Its README says it runs 70B models on a single 4GB GPU, without quantization, distillation, or pruning.
How it works
- A model is a stack of layers. A normal setup loads all of them into the graphics card at once, which is where the 140GB goes.
- AirLLM splits the model into its layers and saves them one by one on your disk the first time you run it.
- When you ask a question it loads one layer onto the GPU, runs it, frees it, and loads the next. Per its README it only ever keeps one layer on the GPU at a time.
- So the GPU memory you need depends on the size of one layer, not the whole model. Think of a giant book that you read one page at a time instead of cramming the whole book into your graphics card.
The trade-off is the disk. The README says the bottleneck is mainly disk loading, and that the original model is decomposed and saved layer-wise first, so you need free disk space for it.
Keep reading
Free, for an email.
Unlocks 6 more sections, 4 copy-paste prompts, the document version, and every other guide on the site, for free. Enter your email once to keep reading.
19,000+ follow where these guides come from.No spam. Unsubscribe in one click.