Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning
1 Institute of Automation, Chinese Academy of Sciences, Beijing, China
2 School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
3 Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences, Shanghai, China
4 State Key Laboratory of Brain Cognition and Brain-Inspired Intelligence Technology, Shanghai, China
Abstract
Method
Experiments
REAL-ROBOT MANIPULATION
All videos in this main experiment section are shown at 3× speed.
Manipulate SimilarFO with Shared Policy
The frozen policy remains effective under background changes and cluttered scenes, and can be reused across different registered SimilarFOs without retraining by switching only the corresponding FO Memory Tokens.
Visual Analysis of FSAE Representations
Citation
BibTeXarXiv 2026
@article{meng2026finegrained,
title={Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning},
author={Meng, Haolong and Qin, Fangbo and Bai, Mengchen and Wang, Houwu and Liu, Cirong and Yu, Shan},
journal={arXiv preprint arXiv:2609.21621},
year={2026}
}


