view article Article **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** nvidia • 8 days ago • 65
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 20 days ago • 264
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 9 days ago • 101
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 14 days ago • 57
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 14 days ago • 91
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 14 days ago • 138
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 14 days ago • 75
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 14 days ago • 110
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 14 days ago • 191
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 22 days ago • 45
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 21 days ago • 48
FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation Paper • 2608.24168 • Published Aug 25 • 4
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 23 days ago • 165
DriveZero: End-to-End Driving Beyond Human Demonstrations Paper • 2609.06055 • Published 26 days ago • 57