Research
My long-term goal is to enable robots to acquire human-level manipulation skills. My doctoral dissertation addresses two connected questions: How do humans perform manipulation? and How can robots learn those skills? The resulting framework moves from multimodal human modeling to online intent decoding and sample-efficient robot learning.
Cross-modal brain–muscle modulation
A Siamese and contrastive-learning framework aligns EEG and EMG in a shared space, while conditional generation models the mapping from brain activity to muscular response. Generated and real signals achieve a maximum mean discrepancy below 0.04, enabling temporal, spatial, and frequency-domain modulation analysis.
Multi-scale motor modeling
Complementary modalities describe manipulation from coarse to fine: monocular video for sequence-level kinematics (Pearson correlation > 0.75), EEG for primitive-level brain activity (accuracy > 80%), and EMG for action segmentation and rhythm modeling (accuracy > 90%; Pearson correlation > 0.8).
Online skill decoding
An “external-data pretraining → local denoising → individual adaptation” pipeline addresses limited samples, low signal-to-noise ratio, and inter-subject variability. Cross-dataset alignment improves accuracy by more than 10%, task-oriented denoising raises SNR by 0.82 dB, and online adaptation improves motor-imagery decoding by 3.6%.
Multi-system human–robot skill transfer
Inspired by memory-, reward-, and error-driven learning in the cortex, basal ganglia, and cerebellum, I combine self-supervised pretraining, reinforcement-learning post-training, and supervised human–robot collaboration. Relative to pure imitation learning, task success improves by 23% and operation steps fall by 36%.
Part I · Representing how humans manipulate
Human manipulation is a hierarchical control process in which high-level planning in the brain is translated into low-level muscular execution. My work makes this implicit structure measurable through cross-modal EEG–EMG representation and generation, then uses video, EEG, and EMG to model behavior at the sequence, primitive, and action-command levels. This work has appeared in IEEE Transactions on Cybernetics, IEEE Transactions on Industrial Informatics, IEEE Transactions on Emerging Topics in Computational Intelligence, IEEE Transactions on Medical Robotics and Bionics, ACM Multimedia, and IJCNN.
Part II · Transferring how robots learn
Human–robot skill transfer requires both robust intent decoding and efficient learning from scarce supervision. I first stabilize online EEG decoding through cross-dataset source alignment, reference-free denoising, and unsupervised adaptation. I then combine pretrained vision–language–action models with action-chunked policy optimization and sparse expert collaboration. Under sparse demonstrations, collaboration increases task success by 13.6%, while the collected expert interventions improve the robot policy by a further 3.8%.
Future directions
- Modeling human skill acquisition: characterize how behavioral and neural representations evolve from novice to expert.
- Multimodal online decoding: combine EEG with EMG, eye tracking, and other wearable signals for more reliable intent transfer.
- Process-level alignment: align human and robot learning trajectories, error-correction patterns, and convergence behavior—not only final task performance.