CodeRabbit CLI can fix your agent’s code before it ever opens a PR - https://coderabbit.link/fireship Free forever for any open source project. Last week, Google surprised us all by shipping their latest micro model Gemma 4 under a truly open source license. But what's the catch? Let's run it... #coding #programming #programming 🔖 Topics Covered - How Gemma 4 works - Gemma 4 benchmarks - TurboQuant 📌 Resources - https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-gemma-4 Want more Fireship? 🗞️ Newsletter: https://bytes.dev 🧠 Courses: https://fireship.dev
ADVERTISEMENT
i saw a comment under google's official video saying he will wait for fireship video, well this is it😂
Google: "Open-source" Lawyers: "Define open"
Actual computer science.
❌ FAANG ✅ GAYMMAN
"as a former GAYMMAN engineer" hits different from "as a former FAANG engineer".
it's good to see big companies actually release open-source software
Google saved ... RAM prices???
I may be wrong but something tells me that gemma 4 may be a distilled version of the next version of gemini which they would be working on. If that's the case, gemini's gonna hit new records on the trust-me-bro benchmarks.
can't wait for basement fine-tuners to distill Mythos into Gemma 4
Wild to see models getting compressed to the point where a single high-end GPU can handle what used to require clusters. Feels less like scaling up and more like finally learning how to scale smart.
The PiedPiper compression algo is here!
GAYMMAN killed me. I needed an intermission in this one. Holy shit dude lol
TurboQuant isn't used for compressing model weights, it’s designed to compress the KV Cache.
0:57 Fireship just casually calling big companies gay by nothing more than arranging their logos.
Legendary refresh pull
0:59 I see what you did there...
I hate the fact that nobody gets what turbo quant really does. It does not shrink the model size you need to keep in memory, it just allows you to keep more context while using less VRAM.Turbo quant is applied to the Key-Value cache and not to any model weights, this cache is only relevant for the context window. if you have a 20GB model and use turbo quant you still have to fit the complete 20GB into VRAM, you just can fit a bigger context window into your VRAM then you could before.
0:02 Insane opening
I’ll wait for the mythos video…
0:48 I can't tell if he recorded this part separately from the rest or if the voice is AI generated