Faster LLM decoding for free: how EAGLE-3 speeds up generation without changing a single output
Speculative decoding is a rare thing in machine learning, a speed-up with no cost to output quality. Here is how it works, and how EAGLE-3 pushes it furthest.
Jun 29, 202611 min read1



