Mobile Vision Perception Lab

Team

community

https://github.com/mvp-ai-lab

mvp-ai-lab

Activity Feed

AI & ML interests

multi-modal foundation models

Recent Activity

xiangan updated a collection about 14 hours ago

LLaVA-OneVision-1.5

geoffreychen777 published a dataset 13 days ago

mvp-lab/LLaVA-OneVision-1.5-RL-Data

geoffreychen777 new activity 13 days ago

mvp-lab/LLaVA-OneVision-1.5-RL-Data:[bot] Conversion to Parquet

View all activity

xiangan

updated a collection about 14 hours ago

LLaVA-OneVision-1.5

Collection

5 items • Updated about 14 hours ago

geoffreychen777

published a dataset 13 days ago

mvp-lab/LLaVA-OneVision-1.5-RL-Data

Viewer • Updated 14 days ago • 69.2k • 246 • 2

geoffreychen777

in mvp-lab/LLaVA-OneVision-1.5-RL-Data 13 days ago

[bot] Conversion to Parquet

#1 opened 14 days ago by

parquet-converter

geoffreychen777

updated a dataset 14 days ago

mvp-lab/LLaVA-OneVision-1.5-RL-Data

Viewer • Updated 14 days ago • 69.2k • 246 • 2

xiangan

published a model 21 days ago

mvp-lab/LLaVA-OneVision-1.5-8B-RL

9B • Updated 21 days ago • 13 • 1

DidiZhu

updated a model 21 days ago

mvp-lab/LLaVA-OneVision-1.5-8B-RL

9B • Updated 21 days ago • 13 • 1

winking636

updated a dataset about 1 month ago

mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M

Viewer • Updated about 1 month ago • 91.5M • 464k • 48

xiangan

authored 4 papers 2 months ago

ForCenNet: Foreground-Centric Network for Document Image Rectification

Paper • 2507.19804 • Published Jul 26 • 11

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

Paper • 2509.09118 • Published Sep 11 • 8

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

Paper • 2510.13515 • Published Oct 15 • 11

ORID: Organ-Regional Information Driven Framework for Radiology Report Generation

Paper • 2411.13025 • Published Nov 20, 2024 • 2

luodian

authored 2 papers 3 months ago

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Paper • 2411.15296 • Published Nov 22, 2024 • 21

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Paper • 2501.13826 • Published Jan 23 • 24

xiangan

authored a paper 3 months ago

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Paper • 2509.23661 • Published Sep 28 • 47

luodian

authored 2 papers 3 months ago

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Paper • 2509.23661 • Published Sep 28 • 47

Visual Jigsaw Post-Training Improves MLLMs

Paper • 2509.25190 • Published Sep 29 • 36

oliveryanzuolu

authored 2 papers 3 months ago

Seedream 4.0: Toward Next-generation Multimodal Image Generation

Paper • 2509.20427 • Published Sep 24 • 81

Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation

Paper • 2509.18824 • Published Sep 23 • 22

iyop45

authored a paper 3 months ago

Region-based Cluster Discrimination for Visual Representation Learning

Paper • 2507.20025 • Published Jul 26 • 19

iyop45

authored a paper 4 months ago

MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge Distillation

Paper • 2303.07815 • Published Mar 14, 2023 • 1

AI & ML interests

Recent Activity

Team members 16

mvp-lab's activity

[bot] Conversion to Parquet