Tools & repos

Run a 70B AI Model on a 4GB Graphics Card: Set Up the Free AirLLM Library

Big open AI models need an expensive graphics card, so most people never try one. AirLLM is a free, open-source Python library that runs a 70 billion parameter model on a single 4GB GPU by keeping only one layer of it in memory at a time. The install command, the code to paste in, ready-made prompts for Claude Code, and the model sizes from its own README

4min read
4prompts to copy
4numbered steps
7sections

A 70 billion parameter model is about 140GB at 16-bit precision (70 billion numbers, two bytes each). That is why running one has meant buying a very expensive graphics card, and why most people never tried.

AirLLM is a free, open-source library (Apache-2.0 licence, over 35,000 stars on GitHub) that gets around the problem. Its README says it runs 70B models on a single 4GB GPU, without quantization, distillation, or pruning.

How it works

  1. A model is a stack of layers. A normal setup loads all of them into the graphics card at once, which is where the 140GB goes.
  2. AirLLM splits the model into its layers and saves them one by one on your disk the first time you run it.
  3. When you ask a question it loads one layer onto the GPU, runs it, frees it, and loads the next. Per its README it only ever keeps one layer on the GPU at a time.
  4. So the GPU memory you need depends on the size of one layer, not the whole model. Think of a giant book that you read one page at a time instead of cramming the whole book into your graphics card.

The trade-off is the disk. The README says the bottleneck is mainly disk loading, and that the original model is decomposed and saved layer-wise first, so you need free disk space for it.

Keep reading

Free, for an email.

Unlocks 6 more sections, 4 copy-paste prompts, the document version, and every other guide on the site, for free. Enter your email once to keep reading.

19,000+ follow where these guides come from.No spam. Unsubscribe in one click.

Read next