Qwen2.5-3B checkpoints for Defects4J unit-test generation with Online Policy Distillation and GRPO.
tomsawyer
tomhu
·
AI & ML interests
None yet
Recent Activity
updated a collection about 12 hours ago
RL4TG updated a model about 13 hours ago
tomhu/RL4TG-DeepSeek-Coder-1.3B-Coder6.7B-Offline-SFT-GRPO published a model about 13 hours ago
tomhu/RL4TG-DeepSeek-Coder-1.3B-Coder6.7B-Offline-SFT-GRPOOrganizations
None yet
models 42
tomhu/RL4TG-DeepSeek-Coder-1.3B-Coder6.7B-Offline-SFT-GRPO
Text Generation • 1B • Updated
tomhu/RL4TG-DeepSeek-Coder-1.3B-GRPO
Text Generation • 1B • Updated
tomhu/RL4TG-DeepSeek-Coder-1.3B-OPD6.7B-GRPO
Text Generation • 1B • Updated
tomhu/RL4TG-DeepSeek-Coder-1.3B-OPD-33B-Teacher
Text Generation • 1B • Updated
tomhu/RL4TG-DeepSeek-Coder-1.3B-OPD-6.7B-Teacher
Text Generation • 1B • Updated
tomhu/RL4TG-DeepSeek-Coder-1.3B-Official-D4J-SFT
Text Generation • 1B • Updated
tomhu/RL4TG-DeepSeek-Coder-1.3B-Coder6.7B-Offline-SFT
Text Generation • 0.7B • Updated
tomhu/RL4TG-Qwen2.5-3B-OPD-7B-Teacher
Text Generation • 3B • Updated • 418
tomhu/RL4TG-Qwen2.5-3B-Coder7B-Offline-SFT-GRPO
Text Generation • 3B • Updated • 26
tomhu/RL4TG-Qwen2.5-3B-Official-D4J-SFT
Text Generation • 3B • Updated • 28