Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Loser Cheems
JingzeShi
26
23
31
Follow
kroeke's profile picture
YIA990ss's profile picture
derrickzhu's profile picture
49 followers
·
21 following
https://github.com/LoserCheems
LoserCheems
AI & ML interests
I like training small languge models.
Recent Activity
authored
a paper
about 21 hours ago
MassAlloc Attention: Let Attention Allocate Its Own Compute
authored
a paper
about 21 hours ago
CoWindow Attention: Full Causal Coverage Is a Collective Property
posted
an
update
2 days ago
Sharing two recent explorations in attention design from our team. We started with two straightforward questions: Does every attention head need to repeatedly attend to the entire causal history? Once attention scores have been computed, do regions with very little contribution still need the full subsequent computation? We explored two approaches: CoWindow Attention (CoWA): Let heads share the work of accessing history. Heads share local context and divide distant context into complementary windows. Each head attends sparsely, while the heads collectively cover the full causal history. https://huggingface.co/papers/2609.32704 MassAlloc Attention (MALA): Let attention allocate its own compute. MALA preserves full causal QK scoring, then uses attention’s own softmax statistics to reduce subsequent computation in low-contribution regions. https://huggingface.co/papers/2609.32712 Both approaches support training forward and backward passes, as well as inference prefill and decoding. In attention-operator benchmarks at 128K tokens on 8×H100 with TP=8, compared with FullAttn: CoWA: 7.4× forward, 8.6× backward, and 3.0× decoding speedups. MALA: 2.2× forward, 3.0× backward, and 1.6× decoding speedups. We also conducted scaling experiments from 0.6B to 14B, alongside separate continued-training experiments at 32B. During 14B training with 32K context, CoWA and MALA reduced total training FLOPs by 28.5% and 23.1%, respectively, while maintaining performance comparable to FullAttn on the evaluated model capabilities. From method design to kernel implementation to model training, our goal was to explore which attention computations can be eliminated, and how to turn those savings into practical gains in ML infrastructure.
View all activity
Organizations
JingzeShi
's models
8
Sort: Recently updated
JingzeShi/flash-sparse-attention
Updated
Jun 7
JingzeShi/OpenSeek-1.4B-A0.4B-KTO
Text Generation
•
1B
•
Updated
Sep 9, 2025
•
10
JingzeShi/OpenSeek-1.4B-A0.4B
Text Generation
•
1B
•
Updated
Aug 24, 2025
•
23
JingzeShi/Doge-20M
Text Generation
•
37.6M
•
Updated
Jul 5, 2025
•
25
JingzeShi/Doge-320M-Reason-checkpoint
0.4B
•
Updated
May 15, 2025
•
8
JingzeShi/Doge-320M-Reason-Distill
Text Generation
•
0.3B
•
Updated
Mar 29, 2025
•
20
JingzeShi/Doge-120M-MoE
0.1B
•
Updated
Mar 20, 2025
•
9
JingzeShi/Mixtral-7B-v0.1
Text Generation
•
7B
•
Updated
Mar 4, 2025
•
14