-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 110 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
Cinny
cinnybun02
AI & ML interests
None yet
Recent Activity
liked a model 2 days ago
InternScience/Agents-A1 updated a collection 6 days ago
Cool-Research-Mad updated a collection 6 days ago
Cool-Research-MadOrganizations
None yet
Cool-Research-Mad
-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 110 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
models 7
cinnybun02/heretic-e2b-fp16
Any-to-Any • 5B • Updated • 6
cinnybun02/Qwen3-0.9B-A0.6B-Coder-Q4_K_M-GGUF
0.9B • Updated • 25
cinnybun02/Anonymizer-4B-Q4_K_M-GGUF
4B • Updated • 15
cinnybun02/gpt-oss-9.0b-specialized-math-pruned-moe-only-12-experts-Q5_K_M-GGUF
Text Generation • 9B • Updated • 7 • 1
cinnybun02/gpt-oss-9.0b-specialized-math-pruned-moe-only-12-experts-Q8_0-GGUF
Text Generation • 9B • Updated • 6
cinnybun02/gpt-oss-12.0b-specialized-science-pruned-moe-only-17-experts-Q8_0-GGUF
Text Generation • 12B • Updated • 34 • 1
cinnybun02/gpt-oss-6.6b-specialized-all-pruned-moe-only-8-experts-Q8_0-GGUF
Text Generation • 7B • Updated • 22