flâneur — a map of the web's best reading

MinD-Vis

mind-vis.github.io · 430 words · saved by 1 readers

Decoding visual stimuli from brain recordings aims to deepen our understanding of the human visual system and build a solid foundation for bridging human vision and computer vision through the Brain-Computer Interface. However, due to the scarcity of data annotations and the complexity of underlying brain information, it is challenging to decode images with faithful details and meaningful semantics. In this work, we present MinD-Vis: Sparse Masked Brain Modeling with Double-Conditioned Diffusion Model for Vision Decoding. Specifically, by boosting the information capacity of representations learned in a large-scale resting-state fMRI dataset, we show that our MinD-Vis framework reconstructed highly plausible images with semantically matching details from brain recordings with very few training pairs. We benchmarked our model and our method outperformed state-of-the-arts in both semantic mapping (100-way semantic classification) and generation quality (FID) by 66% and 41%, respectively.

MinD-Vis Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision Decoding CVPR2023 Zijiao Chen 1* --> Jiaxin Qing 2* Tiange Xiang 3 Wan Lin Yue 1 Juan Helen Zhou 1 1 National University of Singapore, Center for Sleep and Cognition, Centre for Translational Magnetic Resonance Research 2 The Chinese University of Hong Kong, Department of Information Engineering 3 Standford University, Vision and Learning Lab * Equal Contribution Main Paper Main Paper (Low Resolution) Supp. Materials Video --> Data --> Code BibTeX Overview Motivation Decoding visual stimuli fro

Explore this link on the map →

related reading