Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Luong Huu Thanh
fisherman611
4
1
19
Follow
Muigai1's profile picture
DuyguJones's profile picture
webxos's profile picture
5 followers
·
15 following
fisherman611
AI & ML interests
None yet
Recent Activity
reacted
to
salma-remyx
's
post
with 🔥
about 8 hours ago
Inspired by the methods described in "Batch-wise Adaptive Pruning" (arxiv 2608.14003, COLM '26), we implemented a training-free FFN-neuron-pruning knob for SGLang. Authors were motivated by the reality that decode is HBM-bandwidth-bound; the gated MLP is the bulk of weights read per step. Threshold methods (TEAL/CATS) collapse under batching; BWAP's periodic top-k over a max-aggregated score keeps the shared batch mask stable. The method is complementary to KV-sparsity (FFN-weight bandwidth vs KV read, context-length-independent). Our implementation uses an adaptive mask under a captured graph (topology static = k-wide GEMM; mask change = between-replay buffer update via version-gated post_fill; prune steps replay, explore steps eager). Preliminary Results GSM8K n=50, ±6pp: 7B dense 92% → ρ=0.5 84% (−8pp) at up to 1.40× (probe ceiling; ~10% realistic under the adaptive schedule; smaller models need lower ρ as accuracy scales with size). Read more in the upstream issue: https://github.com/sgl-project/sglang/issues/35987
updated
a model
30 days ago
distillation-sql/nothing
published
a model
about 1 month ago
Dream-AI-HUST/synid_ckpt
View all activity
Organizations
fisherman611
's datasets
2
Sort: Recently updated
fisherman611/synthesize_reasoning_data_prompt
Viewer
•
Updated
May 10
•
17.1k
•
47
fisherman611/benchmarks
Preview
•
Updated
Apr 28
•
1