-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 110 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
Cinny
cinnybun02
AI & ML interests
None yet
Recent Activity
upvoted a paper about 17 hours ago
Accurate Expert Predictions in MoE Inference via Cross-Layer Gate upvoted a paper about 17 hours ago
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE
Inference upvoted a paper about 17 hours ago
MoE-Infinity: Activation-Aware Expert Offloading for Efficient MoE
ServingOrganizations
None yet