Tools & repos

DeepGEMM: The Free Engine DeepSeek Runs Its AI On, and Why It Makes AI Cheaper

DeepSeek open sourced the low-level GPU math library it trains and runs its own models on. It's why the same computers can answer far more questions per second, and a big part of why DeepSeek's prices are so low. What it is, who can actually use it, and what it means for your AI bill

2min read
2prompts to copy
8numbered steps
4sections

Every answer an AI model gives is millions of multiplications on a graphics card. Do that math faster on the same hardware, and every answer costs less to serve.

That's what DeepGEMM does. It's DeepSeek's own kernel library, the exact code it trains and runs its models on, and it's free and open source on GitHub under the MIT license.

What DeepGEMM is

  1. A library of the core math every large language model runs on, written for NVIDIA's newest data-center GPUs.
  2. Fast: DeepSeek reports up to 1,550 TFLOPS on a single H800, matching or beating libraries that experts tune by hand.
  3. Kept up to date: DeepSeek keeps adding the kernels behind its newest models, and on Sept 30 it released a version for Huawei's Ascend chips too.
  4. Free: MIT license, so anyone can use it, change it and ship it. Repo: github.com/deepseek-ai/DeepGEMM

Keep reading

Free, for an email.

Unlocks 3 more sections, 2 copy-paste prompts, the document version, and every other guide on the site, for free. Enter your email once to keep reading.

19,000+ follow where these guides come from.No spam. Unsubscribe in one click.

Read next