AI & ML interests
LLM
Recent Activity
View all activity
Papers
View all PapersA unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription.
-
MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization
Paper β’ 2601.01554 β’ Published β’ 66 -
MOSS Transcribe Diarize
π’105Transcribe audio/video with speaker diarization
-
OpenMOSS-Team/MOSS-Transcribe-preview-2B
Automatic Speech Recognition β’ 2B β’ Updated β’ 522 β’ 52 -
OpenMOSS-Team/MOSS-Transcribe-Diarize
Audio-Text-to-Text β’ 0.9B β’ Updated β’ 158k β’ 453
-
OpenMOSS-Team/moss-video-preview-base
Video-Text-to-Text β’ 11B β’ Updated β’ 42 β’ 15 -
OpenMOSS-Team/moss-video-preview-sft
Video-Text-to-Text β’ 11B β’ Updated β’ 81 β’ 17 -
OpenMOSS-Team/moss-video-preview-realtime-sft
Video-Text-to-Text β’ 11B β’ Updated β’ 55 β’ 27 -
OpenMOSS-Team/Realtime-QA-100K
Viewer β’ Updated β’ 100k β’ 211 β’ 8
-
OpenMOSS-Team/MOSS-TTS
Text-to-Speech β’ 8B β’ Updated β’ 3.7k β’ 430 -
OpenMOSS-Team/MOSS-TTS-Local-Transformer
Text-to-Speech β’ 3B β’ Updated β’ 13.7k β’ 31 -
OpenMOSS-Team/MOSS-TTS-Realtime
Text-to-Speech β’ 2B β’ Updated β’ 10.6k β’ 107 -
OpenMOSS-Team/MOSS-TTS-Nano-100M
Text-to-Speech β’ Updated β’ 101k β’ 239
Opensource Lorsas and Transcoders
-
OpenMOSS-Team/MOSS-TTSD-v1.0
Text-to-Speech β’ 8B β’ Updated β’ 7.67k β’ 59 -
OpenMOSS-Team/MOSS-TTSD-v0.7
Text-to-Speech β’ 2B β’ Updated β’ 343 β’ 18 -
OpenMOSS-Team/MOSS-TTSD-v0.5
Text-to-Speech β’ 2B β’ Updated β’ 293 β’ 54 -
OpenMOSS-Team/MOSS-TTSD-v0
Text-to-Speech β’ 2B β’ Updated β’ 33 β’ 28
Evaluating Agentic Backend Coding Capabilities in Real-World Development Scenarios
-
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
Paper β’ 2601.11077 β’ Published β’ 67 -
OpenMOSS-Team/ABC-Bench
Viewer β’ Updated β’ 224 β’ 206 β’ 4 -
OpenMOSS-Team/Qwen3-32B-ABC
Text Generation β’ 33B β’ Updated β’ 19 β’ 3 -
OpenMOSS-Team/Qwen3-8B-ABC
Text Generation β’ 8B β’ Updated β’ 30 β’ 3
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
An Efficient Training Framework for Diffusion Language Models
-
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
Paper β’ 2503.10480 β’ Published β’ 57 -
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
Paper β’ 2506.23127 β’ Published β’ 2 -
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
Paper β’ 2506.21230 β’ Published β’ 1 -
OpenMOSS-Team/Embodied_R1-ScienceWorld
8B β’ Updated β’ 11 β’ 1
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_8-refactor
Text Generation β’ 0.1B β’ Updated β’ 22 β’ 1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_32-refactor
Text Generation β’ 0.1B β’ Updated β’ 18 β’ 1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_16-refactor
Text Generation β’ 0.1B β’ Updated β’ 25 β’ 1 -
OpenMOSS-Team/SmolLM-360M-MLA-d_kv_8-refactor
Text Generation β’ 0.3B β’ Updated β’ 17 β’ 1
-
OpenMOSS-Team/moss-moon-003-sft-plugin
Text Generation β’ Updated β’ 99 β’ 72 -
OpenMOSS-Team/moss-moon-003-sft
Text Generation β’ Updated β’ 123 β’ 129 -
OpenMOSS-Team/moss-moon-003-base
Text Generation β’ Updated β’ 103 β’ 132 -
OpenMOSS-Team/moss-moon-003-sft-int4
Text Generation β’ Updated β’ 102 β’ 41
openeta: embodied task agent
An open-source audio understanding model supporting speech recognition, environmental sound analysis, music understanding, time-aware QA, and complex
-
MOSS Audio 8B Thinking
π’29Generate answers to audio or video prompts
-
OpenMOSS-Team/MOSS-Audio-4B-Instruct
Audio-Text-to-Text β’ 5B β’ Updated β’ 15.3k β’ 84 -
OpenMOSS-Team/MOSS-Audio-4B-Thinking
Audio-Text-to-Text β’ 5B β’ Updated β’ 747 β’ 38 -
OpenMOSS-Team/MOSS-Audio-8B-Instruct
Audio-Text-to-Text β’ 9B β’ Updated β’ 4.06k β’ 50
-
OpenMOSS-Team/MOSS-VL-Instruct-0408
Video-Text-to-Text β’ 11B β’ Updated β’ 775 β’ 107 -
OpenMOSS-Team/MOSS-VL-Base-0408
Video-Text-to-Text β’ 11B β’ Updated β’ 115 β’ 63 -
OpenMOSS-Team/MOSS-VL-Instruct-0708
Video-Text-to-Text β’ 11B β’ Updated β’ 525 β’ 31 -
OpenMOSS-Team/MOSS-VL-Base-0708
Video-Text-to-Text β’ 11B β’ Updated β’ 95 β’ 18
-
AI Can Learn Scientific Taste
Paper β’ 2603.14473 β’ Published β’ 316 -
OpenMOSS-Team/SciJudgeBench
Viewer β’ Updated β’ 732k β’ 243 β’ 11 -
OpenMOSS-Team/SciJudge-4B-2605
Text Generation β’ 4B β’ Updated β’ 65 β’ 6 -
OpenMOSS-Team/SciJudge-30B-2605
Text Generation β’ 31B β’ Updated β’ 534 β’ 3
True Speech-to-Speech Langugage Model
First Omni-modal Future Forecasting Benchmark
https://github.com/OpenMOSS/FRoM-W1
Proactive Robot Manipulation in Omni-modal Context
Open source weights of Lorsa modules introduced in "Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition".
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
Paper β’ 2502.14837 β’ Published β’ 4 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_16
Text Generation β’ 6B β’ Updated β’ 33 β’ 1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_32
Text Generation β’ 6B β’ Updated β’ 32 β’ 1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_64
Text Generation β’ 7B β’ Updated β’ 39 β’ 1
openeta: embodied task agent
A unified multimodal large language model for end-to-end speaker-attributed, time-stamped transcription.
-
MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization
Paper β’ 2601.01554 β’ Published β’ 66 -
MOSS Transcribe Diarize
π’105Transcribe audio/video with speaker diarization
-
OpenMOSS-Team/MOSS-Transcribe-preview-2B
Automatic Speech Recognition β’ 2B β’ Updated β’ 522 β’ 52 -
OpenMOSS-Team/MOSS-Transcribe-Diarize
Audio-Text-to-Text β’ 0.9B β’ Updated β’ 158k β’ 453
An open-source audio understanding model supporting speech recognition, environmental sound analysis, music understanding, time-aware QA, and complex
-
MOSS Audio 8B Thinking
π’29Generate answers to audio or video prompts
-
OpenMOSS-Team/MOSS-Audio-4B-Instruct
Audio-Text-to-Text β’ 5B β’ Updated β’ 15.3k β’ 84 -
OpenMOSS-Team/MOSS-Audio-4B-Thinking
Audio-Text-to-Text β’ 5B β’ Updated β’ 747 β’ 38 -
OpenMOSS-Team/MOSS-Audio-8B-Instruct
Audio-Text-to-Text β’ 9B β’ Updated β’ 4.06k β’ 50
-
OpenMOSS-Team/moss-video-preview-base
Video-Text-to-Text β’ 11B β’ Updated β’ 42 β’ 15 -
OpenMOSS-Team/moss-video-preview-sft
Video-Text-to-Text β’ 11B β’ Updated β’ 81 β’ 17 -
OpenMOSS-Team/moss-video-preview-realtime-sft
Video-Text-to-Text β’ 11B β’ Updated β’ 55 β’ 27 -
OpenMOSS-Team/Realtime-QA-100K
Viewer β’ Updated β’ 100k β’ 211 β’ 8
-
OpenMOSS-Team/MOSS-VL-Instruct-0408
Video-Text-to-Text β’ 11B β’ Updated β’ 775 β’ 107 -
OpenMOSS-Team/MOSS-VL-Base-0408
Video-Text-to-Text β’ 11B β’ Updated β’ 115 β’ 63 -
OpenMOSS-Team/MOSS-VL-Instruct-0708
Video-Text-to-Text β’ 11B β’ Updated β’ 525 β’ 31 -
OpenMOSS-Team/MOSS-VL-Base-0708
Video-Text-to-Text β’ 11B β’ Updated β’ 95 β’ 18
-
OpenMOSS-Team/MOSS-TTS
Text-to-Speech β’ 8B β’ Updated β’ 3.7k β’ 430 -
OpenMOSS-Team/MOSS-TTS-Local-Transformer
Text-to-Speech β’ 3B β’ Updated β’ 13.7k β’ 31 -
OpenMOSS-Team/MOSS-TTS-Realtime
Text-to-Speech β’ 2B β’ Updated β’ 10.6k β’ 107 -
OpenMOSS-Team/MOSS-TTS-Nano-100M
Text-to-Speech β’ Updated β’ 101k β’ 239
-
AI Can Learn Scientific Taste
Paper β’ 2603.14473 β’ Published β’ 316 -
OpenMOSS-Team/SciJudgeBench
Viewer β’ Updated β’ 732k β’ 243 β’ 11 -
OpenMOSS-Team/SciJudge-4B-2605
Text Generation β’ 4B β’ Updated β’ 65 β’ 6 -
OpenMOSS-Team/SciJudge-30B-2605
Text Generation β’ 31B β’ Updated β’ 534 β’ 3
Opensource Lorsas and Transcoders
-
OpenMOSS-Team/MOSS-TTSD-v1.0
Text-to-Speech β’ 8B β’ Updated β’ 7.67k β’ 59 -
OpenMOSS-Team/MOSS-TTSD-v0.7
Text-to-Speech β’ 2B β’ Updated β’ 343 β’ 18 -
OpenMOSS-Team/MOSS-TTSD-v0.5
Text-to-Speech β’ 2B β’ Updated β’ 293 β’ 54 -
OpenMOSS-Team/MOSS-TTSD-v0
Text-to-Speech β’ 2B β’ Updated β’ 33 β’ 28
True Speech-to-Speech Langugage Model
Evaluating Agentic Backend Coding Capabilities in Real-World Development Scenarios
-
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
Paper β’ 2601.11077 β’ Published β’ 67 -
OpenMOSS-Team/ABC-Bench
Viewer β’ Updated β’ 224 β’ 206 β’ 4 -
OpenMOSS-Team/Qwen3-32B-ABC
Text Generation β’ 33B β’ Updated β’ 19 β’ 3 -
OpenMOSS-Team/Qwen3-8B-ABC
Text Generation β’ 8B β’ Updated β’ 30 β’ 3
First Omni-modal Future Forecasting Benchmark
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
https://github.com/OpenMOSS/FRoM-W1
An Efficient Training Framework for Diffusion Language Models
Proactive Robot Manipulation in Omni-modal Context
-
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
Paper β’ 2503.10480 β’ Published β’ 57 -
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
Paper β’ 2506.23127 β’ Published β’ 2 -
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
Paper β’ 2506.21230 β’ Published β’ 1 -
OpenMOSS-Team/Embodied_R1-ScienceWorld
8B β’ Updated β’ 11 β’ 1
Open source weights of Lorsa modules introduced in "Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition".
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_8-refactor
Text Generation β’ 0.1B β’ Updated β’ 22 β’ 1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_32-refactor
Text Generation β’ 0.1B β’ Updated β’ 18 β’ 1 -
OpenMOSS-Team/SmolLM-135M-MLA-d_kv_16-refactor
Text Generation β’ 0.1B β’ Updated β’ 25 β’ 1 -
OpenMOSS-Team/SmolLM-360M-MLA-d_kv_8-refactor
Text Generation β’ 0.3B β’ Updated β’ 17 β’ 1
The MHA2MLA model published in the paper "Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-Based LLMs"
-
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
Paper β’ 2502.14837 β’ Published β’ 4 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_16
Text Generation β’ 6B β’ Updated β’ 33 β’ 1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_32
Text Generation β’ 6B β’ Updated β’ 32 β’ 1 -
OpenMOSS-Team/Llama-2-7B-MLA-d_kv_64
Text Generation β’ 7B β’ Updated β’ 39 β’ 1
-
OpenMOSS-Team/moss-moon-003-sft-plugin
Text Generation β’ Updated β’ 99 β’ 72 -
OpenMOSS-Team/moss-moon-003-sft
Text Generation β’ Updated β’ 123 β’ 129 -
OpenMOSS-Team/moss-moon-003-base
Text Generation β’ Updated β’ 103 β’ 132 -
OpenMOSS-Team/moss-moon-003-sft-int4
Text Generation β’ Updated β’ 102 β’ 41