-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 110 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
Cinny
cinnybun02
AI & ML interests
None yet
Recent Activity
upvoted a paper 2 days ago
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight
Compression upvoted a paper 2 days ago
LLM in a flash: Efficient Large Language Model Inference with Limited
Memory new activity 5 days ago
ultimatechris/Ornith-1.5-9B-DFlash-SGLang:Do one for 35bOrganizations
None yet