Accelerating Gemma 4: faster inference with multi-token prediction drafters

(blog.google)

569 points | by amrrs 18 hours ago ago

272 comments