Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning

Haolong Meng1,2Fangbo Qin1,2,4,*Mengchen Bai1Houwu Wang1Cirong Liu3,4Shan Yu1,2,4

1 Institute of Automation, Chinese Academy of Sciences, Beijing, China

2 School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China

3 Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences, Shanghai, China

4 State Key Laboratory of Brain Cognition and Brain-Inspired Intelligence Technology, Shanghai, China

* Corresponding author: Fangbo Qin

Comparison of scene-level visual conditioning, category-level conditioning, and fine-grained object-focused conditioning with FOM-SAM3 and FSAE.

Abstract

Method

Persistent FO Memory Learning: concept-level prompt tokens are transformed into persistent FO memories using a one-vs-rest training loss with SAM3 frozen.
Persistent FO Memory Learning
Focused Spatial-Appearance Encoding: retrieved FO memories prompt frozen SAM3 to extract boxes, masks, and FPN features, which form spatial and appearance conditions for DP and ACT.
Focused Spatial-Appearance Encoding (FSAE)

Experiments

REAL-ROBOT MANIPULATION

Complete real-robot manipulation results, including the overall and per-task success rates for Collect Can, Push Box, and Lid Cup.
Real-robot manipulation results.

All videos in this main experiment section are shown at 3× speed.

Manipulate SimilarFO with Shared Policy

The frozen policy remains effective under background changes and cluttered scenes, and can be reused across different registered SimilarFOs without retraining by switching only the corresponding FO Memory Tokens.

Visual Analysis of FSAE Representations

FSAE representation visualization along six manipulation waypoints and t-SNE trajectories for FO1, FO1 with distractors, and a Similar FO.
FSAE representations across six manipulation waypoints (Fig. 7). Boxes preserve spatial grounding; focused appearance features track task progress, remain stable with distractors, and distinguish SimilarFOs.

Citation

BibTeXarXiv 2026
@article{meng2026finegrained,
  title={Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning},
  author={Meng, Haolong and Qin, Fangbo and Bai, Mengchen and Wang, Houwu and Liu, Cirong and Yu, Shan},
  journal={arXiv preprint arXiv:2609.21621},
  year={2026}
}