MultiBBQ Fairness benchmark for multimodal LLMs: dataset, image perturbations, and results (paper: Fairness Failure Modes of Multimodal LLMs). MLL-Lab/MultiBBQ Viewer • Updated 29 days ago • 1.64k • 208 MLL-Lab/MultiBBQ-perturbations Viewer • Updated Jul 17 • 8.59k • 4.68k MLL-Lab/MultiBBQ-results Preview • Updated 21 days ago • 666 MLL-Lab/MultiBBQ-realworld Viewer • Updated Jul 17 • 79 • 389
viewsuite-models MLL-Lab/viewagent-all-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 7 MLL-Lab/viewagent-ivp-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 11
viewagent-datasets MLL-Lab/viewsuite Updated May 28 • 267 • 1 MLL-Lab/viewagent-ai2thor Preview • Updated 22 days ago • 36
mindcube-models MLL-Lab/MindCube-RL-PlainCogMap-FromSFT 4B • Updated Mar 23 • 116 MLL-Lab/MindCube-Qwen2.5VL-RawQA-SFT 4B • Updated Jun 20, 2025 • 20
viewagent-models MLL-Lab/viewagent-all-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 7 MLL-Lab/viewagent-ivp-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 11
MultiBBQ Fairness benchmark for multimodal LLMs: dataset, image perturbations, and results (paper: Fairness Failure Modes of Multimodal LLMs). MLL-Lab/MultiBBQ Viewer • Updated 29 days ago • 1.64k • 208 MLL-Lab/MultiBBQ-perturbations Viewer • Updated Jul 17 • 8.59k • 4.68k MLL-Lab/MultiBBQ-results Preview • Updated 21 days ago • 666 MLL-Lab/MultiBBQ-realworld Viewer • Updated Jul 17 • 79 • 389
viewagent-datasets MLL-Lab/viewsuite Updated May 28 • 267 • 1 MLL-Lab/viewagent-ai2thor Preview • Updated 22 days ago • 36
viewsuite-models MLL-Lab/viewagent-all-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 7 MLL-Lab/viewagent-ivp-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 11
mindcube-models MLL-Lab/MindCube-RL-PlainCogMap-FromSFT 4B • Updated Mar 23 • 116 MLL-Lab/MindCube-Qwen2.5VL-RawQA-SFT 4B • Updated Jun 20, 2025 • 20
viewagent-models MLL-Lab/viewagent-all-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 7 MLL-Lab/viewagent-ivp-qwen25vl7b Image-Text-to-Text • 8B • Updated Jun 29 • 11