Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
gen
ginigini
23
210
Follow
aiqtech's profile picture
fantos's profile picture
Anserwise's profile picture
25 followers
ยท
38 following
AI & ML interests
None yet
Recent Activity
upvoted
an
article
1 day ago
The Fast Gemma Challenge: our verified-SOTA recipe, in full
reacted
to
SeaWolf-AI
's
post
with ๐ง
1 day ago
We wrote up our run in The Fast Gemma Challenge โ as vidraft-darwin โ and wanted to share the recipe. ๐ https://huggingface.co/spaces/gemma-challenge/gemma-dashboard Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify โ what we're proud of is the fastest result that keeps quality. The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the publicโprivate gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups. Huge thanks to @firfir-cast, @gemma-slayer, @chiku-inu, @kenyan-duma, @dixie-flatline and everyone who shared their experiments. Full write-up ๐ https://huggingface.co/blog/FINAL-Bench/fast-gemma
reacted
to
SeaWolf-AI
's
post
with ๐
1 day ago
We wrote up our run in The Fast Gemma Challenge โ as vidraft-darwin โ and wanted to share the recipe. ๐ https://huggingface.co/spaces/gemma-challenge/gemma-dashboard Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify โ what we're proud of is the fastest result that keeps quality. The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the publicโprivate gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups. Huge thanks to @firfir-cast, @gemma-slayer, @chiku-inu, @kenyan-duma, @dixie-flatline and everyone who shared their experiments. Full write-up ๐ https://huggingface.co/blog/FINAL-Bench/fast-gemma
View all activity
Organizations
None yet
spaces
2
Sort:ย Recently updated
pinned
Build error
Agents
HeartMuLa
๐
A Family of Open Sourced Music Foundation Models
Build error
FinePDFs: Liberating 3T of the finest tokens from PDFs
๐
models
0
None public yet
datasets
0
None public yet