Integrating camera, LIDAR, radar, thermal, and audio data for robust scene understanding in dynamic environments.
Developing frameworks that maintain high accuracy when one or more sensor modalities are unavailable or corrupted.
Vision-based and multimodal perception for lane detection, obstacle detection, and semantic segmentation in autonomous vehicles.
Multimodal learning for social robots, including person classification, emotion recognition, and gesture understanding.
CNNs, RNNs, Transformers, Vision Transformers, Graph Neural Networks, and metric learning for perception and recognition tasks.
Markerless motion capture, gait analysis, abnormal behavior detection, and sign language recognition using multi-view video.
Transformer-based frameworks for audio-visual emotion recognition, sound event detection, and speech-gesture generation.
Automatic extrinsic calibration of LIDAR-stereo and non-overlapping camera networks using probabilistic and optimization methods.
Leveraging sequence-level weak labels and self-supervised techniques to reduce annotation burden for frame-level perception tasks.
I actively seek partnerships with researchers and industry partners in autonomous systems, robotics, and multimodal AI.
vjohn@ltu.edu