Paired-4:8 + NVFP4 W4A4 expert compressed MoE models for NVIDIA Blackwell Sparse Tensor Cores.
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Disaggregated Quantization: Specializing LLM Prefill and Decode
models 165
ISTA-DASLab/Qwen3-8B-MatGPTQ
8B • Updated • 16
ISTA-DASLab/Llama-3.1-8B-Instruct-MatGPTQ
8B • Updated • 17
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Image-Text-to-Text • 177B • Updated • 755k • 405
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
Image-Text-to-Text • 117B • Updated • 33.3k • 152
ISTA-DASLab/Kimi-K2.5-P48NVFP4-MoESQ
Text Generation • 646B • Updated • 109
ISTA-DASLab/Qwen3.5-397B-A17B-P48NVFP4-MoESQ
Image-Text-to-Text • 258B • Updated • 23
ISTA-DASLab/Qwen3-30B-A3B-P48NVFP4-MoESQ
Text Generation • 20B • Updated • 100 • 1
ISTA-DASLab/Qwen3.8-27B-NVFP4-prefiller
Text Generation • 24B • Updated • 427 • 21
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
Image-Text-to-Text • 27B • Updated • 1.68M • 1.86k
ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
Image-Text-to-Text • 27B • Updated • 8.08k • 20