A comprehensive framework designed to cultivate VLMs with human-like visuospatial abilities.
Ray Yang
rayruiyang
AI & ML interests
None yet
Recent Activity
upvoted a paper about 12 hours ago
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training upvoted a paper about 1 month ago
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone upvoted a paper about 2 months ago
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image GenerationOrganizations
None yet