Learning Functional Subspaces for Neural Network Compression Paper • 2609.40127 • Published 12 days ago • 25
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing Paper • 2610.05842 • Published 7 days ago • 13
position-specialist-speculative-decoding/Speed-E3-Llama3.1-8B-Instruct-vllm 0.9B • Updated Jan 28 • 76 • 6
ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context Paper • 2609.36684 • Published 13 days ago • 20
open-llm-leaderboard-old/details_saarvajanik__facebook-opt-6.7b-gqa-ub-16-best-for-KV-cache Updated Jan 28, 2024 • 557 • 9
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning Paper • 2610.02181 • Published 11 days ago • 28
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs Paper • 2609.32259 • Published 13 days ago • 101