Research Overview
My research focuses on developing innovative solutions to enhance multimodal perception and learning for intelligent systems. I am particularly interested in addressing the challenges of missing modality in multimodal learning, sensor fusion, robust perception across multiple sensor modalities, and human motion analysis. My contributions span a wide range of domains, including human-robot interaction, intelligent mobility, and autonomous driving, where I have applied deep learning and probabilistic methods to integrate and interpret data from visible, thermal, audio, radar, and LIDAR sensors. Through my work, I aim to enable more robust and adaptive intelligent systems, particularly in the fields of robotics, autonomous vehicles, and human-computer interaction.
At Lawrence Tech, I am involved in the research and development of robust multimodal learning for robotics, intelligent mobility, and health care applications.
At RIKEN, I was involved in the Guardian Robot Research project, a collaborative effort involving six teams. The objective is to develop a social robot with human-like characteristics, with my contributions focusing on multimodal learning — addressing the missing modality problem, weak supervised multimodal learning, and multimodal perception.
The research addresses the missing modality problem in multimodal person classification and recognition tasks, where deep learning-based frameworks often require the presence of all modalities to achieve state-of-the-art accuracy. However, in real-world applications, there is the possibility of missing modalities due to conditions such as sensor malfunction or failure, resulting in missing data. An overview of the missing modality problem is presented below.
In this research we formulate novel deep learning frameworks using metric learning, cascade learning, transfer learning, and progressive learning to address the missing modality problem. Using metric learning, we proposed novel latent loss functions to obtain latent embeddings for multimodal data even when one or more modalities are missing.
We also developed a cascaded framework comprising three deep learning models that utilize metric learning to generate complete multimodal data from incomplete inputs in the latent feature space. Additionally, KModNet was proposed to tackle both the missing modality and multimodal portability problems, enabling classification using any subset of available modalities. An audio-visual person recognition framework was also introduced, leveraging audio-based person attributes and a multi-head attention transformer to enhance recognition accuracy in the absence of visual data.
Training for multimodal perception tasks typically requires annotating frame-level strong labels, which is a laborious and challenging task. Conversely, annotating sequence-level weak labels is comparatively simpler. However, employing weak labels for training a multimodal perception framework poses challenges. We proposed a novel framework for training video-based frame-level action recognition models utilizing sequence-level weak labels. Sequence-level annotations are used to generate view-specific latent embeddings, which significantly enhance the performance of downstream models for action recognition and detection tasks. The framework incorporates a unique latent loss function that demonstrated improved accuracy despite the limited supervision provided by weak labels.
Our research in human-robot interaction focuses on integrating visible cameras, thermal cameras, and microphones to improve person classification and emotion recognition. We developed a visible-thermal person classifier that employs transfer learning, knowledge distillation, and vision transformers. The classifier utilizes teacher models to guide the multimodal classifier, incorporating a novel loss function for knowledge distillation. Furthermore, a transformer-based model for audio-visual emotion recognition was created, featuring audio self-attention, video self-attention, and audio-video cross-attention mechanisms to effectively combine modalities.
As part of a collaborative effort involving Toyota Technological Institute, Denso, Nippon Soken, Aisin, Toyota R&D Labs, and Tier IV, our research in intelligent mobility focused on vision-based perception and the integration of diverse sensor modalities for autonomous driving applications. I was the principal contributor in all aforementioned projects.
For the different road environment perception problems, we utilised deep learning and observed state-of-the-art detection results. For pedestrian, traffic light, and lane detection, we achieve nearly 100% detection accuracy with very few false positives. In the case of lane detection, using our proposed method we estimate the ego-lane image location even under lane occlusion and absence of lane markers. We have also developed a road surface and small object detection algorithm using convolutional neural networks, where geometric and spatial priors are incorporated to enhance classification accuracy.
We developed deep sensor fusion frameworks, namely RVNet and SO-Net, which combined radar and appearance information to enhance vehicle detection and semantic segmentation. Additionally, we advanced methods for integrating visible and thermal cameras for scene forecasting and pedestrian behavior estimation, while formulating sensor fusion techniques to integrate LIDAR and stereo cameras, achieving improved depth precision and predicting lane markers in complex driving scenarios. Sample results of the fusion frameworks are presented below.
Results of the ChiNet and RVNet sensor fusion frameworks.
In a project aimed at enhancing touchless automotive user interfaces, we developed a deep learning and sparse modeling framework for gesture recognition from video data, facilitating intuitive user interaction in vehicles. Our efforts also focused on localizing autonomous vehicles using stereo vision-based depth information, introducing a novel particle swarm optimization framework in conjunction with Kalman filtering. Moreover, methods for generating virtual depth images from 3D point clouds were proposed, significantly improving localization accuracy in complex environments.
In collaboration with the University of Amsterdam, Philips Research, and Eagle Vision, we addressed the extrinsic calibration of camera networks with non-overlapping fields of view. Our work involved developing a probabilistic approach for appearance-based person re-identification across camera networks, accounting for varying illumination and camera gain. An automatic camera calibration algorithm was introduced, leveraging multiple trajectories and particle swarm optimization, coupled with a Bayesian framework to correct calibration errors.
For my PhD thesis at the University of Dundee, we focused on markerless human motion tracking and classification using multiple-view video sequences. We developed a method employing particle swarm optimization and charting to analyze human motion without physical markers. This research contributed to practical applications of human motion analysis across various domains and included the maintenance and updating of the motion capture studio, encompassing software development, synchronized camera systems, and professional lighting control. This research was featured in Science Scotland, Royal Society of Edinburgh.
My future research endeavours to extend the literature in multimodal learning in robotics, intelligent mobility, and human motion analysis. As intelligent systems evolve to operate autonomously in complex, dynamic environments, it becomes imperative to develop algorithms capable of utilising multimodal data.
I intend to further explore methods for overcoming the missing modality problem, allowing systems to function optimally even when sensor data is noisy or incomplete. We will explore graph neural networks, representation learning, self-supervised learning, and metric techniques to integrate information from various sensor modalities, including visible, thermal, and audio data.
Using data obtained from multiple multimodal sensors for perception is a challenging task. We will solve perception problems using multiple sensors while addressing the missing modality problem, and continue research on weak supervised learning. For multi-view multimodal datasets, it is difficult to annotate modality-specific frame-level labels for various perception tasks; however, it is easier to obtain sequence-level weak labels. I will focus on developing algorithms to utilise the easily available weak labels for frame-level perception tasks.
Sign language recognition using multimodal data is a challenging task within human motion analysis. The research will focus on interpreting complex gestures using multiple sensor inputs — including visual and audio data — while accounting for temporal dependencies. We will also focus on generating sign language gestures using audio and text inputs.
Perception is a critical task for autonomous driving. Recent deep learning frameworks using vision sensors report state-of-the-art performance. However, their performance is limited by appearance variations, illumination variations, and occlusion. To address these issues, a sensor fusion-based approach will be adopted where the complementary advantages of different sensors are incorporated in the perception tasks.
Perception of human motion from video sequences is an important research problem with numerous applications in security, surveillance, biomedical, animation, and autonomous driving. I will investigate the problems of abnormal behaviour analysis and gait analysis. Automatically identifying suspicious human behaviour requires implementing: human detection; human tracking; human re-identification across networks; and identifying abnormal behaviour. Key challenges include variation in appearance owing to body shape and clothing, illumination variations, motion variations, and occlusions.