Where Do Humans Look When Demonstrating to Robots?
Human Gaze Behavior in Pick-And-Place Tasks across Demonstration Devices

35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026)
1Toyota Motor Corporation
2Nara Institute of Science and Technology

Illustration of the research question. Left: Insights from cognitive science—gaze behavior analyzed under natural conditions. Right: Device-imposed demonstration conditions—how demonstration devices influence gaze behavior when demonstrations are collected for imitation learning in pick-and-place tasks.

Overview Video


Abstract


Imitation learning for generalizable performance often requires a large volume of demonstration data, making the process significantly costly. One promising strategy to address this challenge is to leverage the cognitive skills of human demonstrators with strong generalization capability, particularly by revealing the underlying task demands reflected in their gaze behavior. However, imitation learning typically involves humans collecting data using demonstration devices that emulate a robot's embodiment and visual condition. This raises the question of how such devices influence gaze behavior. We propose an experimental framework that systematically analyzes human demonstrators' gaze behavior across a spectrum of robot-emulating demonstration devices. Our experimental results show that certain device properties shift gaze from task-goal cues (e.g., objects) toward control-monitoring cues (e.g., the end-effector). Furthermore, these shifts directly affect the performance of typical gaze-based imitation learning models, sometimes degrading it below non-gaze baselines.

Method



(a)
(b)
(c)
(d)
(e)
(f)
(g)


To investigate how demonstration devices influence demonstrators' gaze behavior, we conduct two comparisons: (1) an embodiment comparison across four conditions—(a) Natural, (b) UMI, (c) Leader, and (d) Leader–Follower, and (2) a visual condition comparison across three conditions—(a) Natural, (e) HMD-Ego, and (f) HMD-Top-down. To provide foundational insights and empirical validation, all comparisons are conducted using a canonical pick-and-place benchmark with (g) commonly available household objects.

Results


Our experimental results indicate that stronger embodiment or visual condition constraints shift demonstrators' gaze from task-goal cues (such as objects) toward control-monitoring cues (such as the end-effector). Furthermore, we found that the influence of these constraints on gaze behavior depends on the verbal instructions given to the demonstrator. Consistent with these findings, policies trained with gaze behavior collected under Natural condition consistently achieved the best performance, whereas gaze collected using constrained demonstration conditions reduced the effectiveness of gaze-based imitation learning.

Citation


@inproceedings{Ishida2026WhenDoHumansLook,
    title = {Where Do Humans Look When Demonstrating to Robots? Human Gaze Behavior in Pick-And-Place Tasks across Demonstration Devices},
    author = {Ishida, Yutaro and Matsubara, Takamitsu and Kanai, Takayuki and Shintani, Kazuhiro and Bito, Hiroshi},
    booktitle = {IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)},
    year = {2026}
}

Notification


The project page was solely developed for and published as part of the publication, titled ``Where Do Humans Look When Demonstrating to Robots? Human Gaze Behavior in Pick-And-Place Tasks across Demonstration Devices'' for its visualization. We do not ensure the future maintenance and monitoring of this page.

Contents might be updated or deleted without notice regarding the original manuscript update and policy change.

This webpage template was adapted from DiffusionNOCS -- we thank Takuya Ikeda for additional support and making their source available.