Faster, cheaper AI
How can large models run with less memory and compute, with a guarantee on what we give up?
Running large models is expensive, and most ways of cutting the cost make no promise about what you lose. We decide, prompt by prompt, how much of a model’s memory can be dropped within a set error budget, and compress weights in ways that lose less.
What we’ve found
Compressing a model’s memory with a guarantee
Deciding, prompt by prompt, how much of a model’s memory (its key-value cache) can be dropped while staying within a set error budget.
Better 4-bit models
Rearranging a model’s weights in mathematically exact ways so it loses less quality when compressed to 4 bits.
Triton on NVIDIA’s GB10
An open-source add-on that makes Triton, a widely used tool for writing fast GPU programs, work on NVIDIA’s GB10 (Blackwell) desktop hardware until official support arrives.
Where it stops working
- Recent work from other groups uses the same family of weight rearrangements, so our next step is a head-to-head comparison before we claim an advantage.
Projects
Compressing a model’s memory with a guarantee
CompletedDeciding, prompt by prompt, how much of a model’s memory can be dropped while staying within a set error budget.
Better 4-bit models with exact rearrangements
CompletedRearranging a model’s weights in mathematically exact ways so it loses less quality when compressed.
Running Triton on NVIDIA’s GB10
CompletedAn open-source add-on that makes Triton work on NVIDIA’s GB10 (Blackwell) desktop hardware.