Publications
Full list on Google Scholar ↗
Patents
- Vehicle Control Device. Inventors: M. Harada, Y. Tanaka, Teraji, S. Mita, V. John. Applicant: SOKEN Co., Ltd., DENSO Corporation, Toyota School Foundation. Patent No. JP7303521B2, Japan. Filed: June 28, 2019. Published: January 28, 2021.
- Steering Angle Determination Device and Autonomous Vehicle. Inventors: K. Ishimaru, H. Tehrani, M. Konishi, V. John, S. Mita. Applicant: SOKEN Co., Ltd., DENSO Corporation, Toyota School Foundation. Patent No. JP7078458B2, Japan. Filed: May 30, 2018. Published: December 5, 2019.
- Recognition Device. Inventors: B. Chen, H. Tehrani, V. John, S. Mita, S. Nishino, K. Ishimaru. Applicant: DENSO Corporation, Toyota School Foundation, SOKEN Co., Ltd. Patent No. JP7071215B2, Japan. Filed: May 24, 2018. Published: November 28, 2019.
- Steering Angle Determination Device, Autonomous Vehicle. Inventors: K. Ishimaru, H. Tehrani, V. John, S. Mita. Applicant: SOKEN Co., Ltd., DENSO Corporation, Toyota School Foundation. Patent No. JP6923362B2, Japan. Filed: May 30, 2017. Published: December 20, 2018.
- Position Estimation Device, Mobile Device. Inventors: S. Nishino, H. Tehrani, K. Ishimaru, S. Yu-Chuan, V. John, S. Mita. Applicant: SOKEN Co., Ltd., DENSO Corporation, Toyota School Foundation. Patent No. JP6959032B2, Japan. Filed: May 17, 2017. Published: December 6, 2018.
- Boundary Line Estimation Device. Inventors: K. Ishimaru, H. Tehrani, T. Shirai, V. John, S. Mita. Applicant: SOKEN Co., Ltd., DENSO Corporation, Toyota School Foundation. Patent No. JP6798860B2, Japan. Filed: November 29, 2016. Published: June 7, 2018.
- Recognition Device, Program. Inventors: K. Ishimaru, H. Tehrani, T. Shirai, V. John, S. Mita. Applicant: SOKEN Co., Ltd., DENSO Corporation, Toyota School Foundation. Patent No. JP6701057B2, Japan. Filed: November 4, 2016. Published: May 10, 2018.
Book Chapters
- N. Ramesh, P. Joe, V. John. "Surgical Robots in Urology and Gynecology." Surgical Robots in Smart Hospitals, Wiley, 2025. [DOI]
- Y. Jawale, P. Joe, V. John. "Next-Generation Industrial Robots." Surgical Robots in Smart Hospitals, Wiley, 2025. [DOI]
- L. Shuo, V. John, Z. Liu. "On the Prospects of Using Deep Learning for Surveillance and Security Applications." Deep Learning for Image Processing Applications, Advances in Parallel Computing, 2017.
- V. John. "Human Gait Signature for Biometric Authentication." Advances in Biometrics for Secure Human Authentication, CRC Press, 2013.
- V. John, S. Ivekovic, E. Trucco. "Markerless Human Motion Capture using Hierarchical Particle Swarm Optimisation." Computer Vision, Imaging and Computer Graphics: Theory and Applications, CCIS Vol. 68, 2010.
Journal Articles
- T. Nguyen, Y. Kawanishi, V. John, T. Komamizu, I. Ide. "MultiSensor-Home: Multimodal multi-view dataset and benchmarks for action recognition in home environments." Pattern Recognition, Vol. 179, Apr 2026. [DOI]
- V. John, Y. Kawanishi. "Hierarchical graph attention networks with spatio-temporal class tokens for distributed audio-visual event classification." Multimedia Tools and Applications, Vol. 85, Issue 4, Apr 2026. [DOI]
- T. Nguyen, Y. Kawanishi, V. John, T. Komamizu, I. Ide. "Action Selection Learning for Weakly Labeled Multi-modal Multi-view Action Recognition." ACM Trans. Multimedia Computing, Communications, and Applications, June 2025.
- L. Zhang, Q. Liang, V. John, H. Chen, S. Li, W. Li, Y. Chen. "Intelligent Psyllid Monitoring Based on DiTs-YOLOv10-SOD." IEEE Trans. AgriFood Electronics, Vol. 3, Issue 1, Apr 2025.
- V. John, Y. Kawanishi. "Multimodal Cascaded Framework with Multimodal Latent Loss Functions Robust to Missing Modalities." ACM Trans. Multimedia Computing, Communication, and Applications, Jan 2025.
- C. Yang, Z. Tian, X. You, K. Jia, T. Liu, Z. Pan, V. John. "Polylanenet++: enhancing the polynomial regression lane detection based on spatio-temporal fusion." Signal, Image and Video Processing, 18(4), 2024.
- V. John, Y. Kawanishi. "Progressive Learning of a Multimodal Classifier Accounting for Different Modality Combinations." Sensors, 23(10), 2023.
- Z. Bao, W. Li, J. Chen, H. Chen, V. John, C. Xiao, Y. Chen. "Predicting and Visualizing Citrus Color Transformation Using a Deep Mask-Guided Generative Network." Plant Phenomics, 5, 2023.
- X. Zhao, W. Li, H. Chen, Y. Wang, Y. Chen, V. John. "Distribution dependent feature selection for deep neural networks." Applied Intelligence, 52(4), 2022.
- V. John, A. Lakshmanan, S. Mita, A. Boyali, S. Thompson. "Deep Visible and Thermal Camera-based Optimal Semantic Segmentation using Semantic Forecasting." ASME J. Autonomous Vehicles and Systems, 1(2), 2021.
- V. John, S. Mita. "Deep Feature-Level Sensor Fusion Using Skip Connections for Real-Time Object Detection in Autonomous Driving." Electronics, 10(4), 2021.
- S. Liu, M. Gao, V. John, Z. Liu, E. Blasch. "Deep Learning Thermal Image Translation for Night Vision Perception." IEEE Trans. Intelligent Systems and Technology, 12(9), 2021.
- S. Liu, H. Liu, V. John, Z. Liu, E. Blasch, Y. Huang. "Enhanced Situation Awareness through CNN-based Deep Multi-Modal Image Fusion." Optical Engineering, 59(5), 2020.
- M. Gao, J. Jiang, G. Zou, V. John, Z. Liu. "RGB-D-based Object Recognition using Multimodal Convolutional Neural Networks: A Survey." IEEE Access, Vol. 7, 2019.
- V. John, Y. Xu, S. Mita. "Stereo Vision based Vehicle Localization in Point Cloud Maps Using Multi-Swarm Particle Swarm Optimization." Signal Image and Video Processing, Vol. 13, 2019.
- V. John, A. Boyali, H. Tehrani, K. Ishimaru, M. Konishi, Z. Liu, S. Mita. "Estimation of Steering Angle and Collision Avoidance for Automated Driving using Deep Mixture of Experts." IEEE Trans. Intelligent Vehicles, 3(4), 2018.
- V. John, Z. Liu, C. Guo, S. Mita, K. Kidono, H. Tehrani, K. Ishimaru. "Real-time Road Surface and Semantic Lane Estimation using Deep Features." Signal, Image and Video Processing, 12(4), 2018.
- Z. Liu, E. Blasch, G. Bhatnagar, V. John, W. Wu, R. Blum. "Fusing Synergistic Information from Multi-Sensor Images: from Implementation to Performance Assessment." Information Fusion, Vol. 42, 2017.
- V. John, A. Boyali, S. Mita. "Gabor Filter and Gershgorin Disk-based Convolutional Filter Constraining for Image Classification." Int. J. Machine Learning and Computing, 1(4), 2017.
- S. S. Rathour, A. Boyali, L. Zheming, S. Mita, V. John. "A Map-based Lateral and Longitudinal DGPS/DR Bias Estimation Method for Autonomous Driving." Int. J. Machine Learning and Computing, 7(4), 2017.
- Z. Liu, E. Blasch, V. John. "Statistical Comparison of Pixel-Level Fusion Algorithms." Information Fusion, Vol. 36, 2017.
- V. John, S. Tsuchizawa, Z. Liu, S. Mita. "Fusion of Thermal and Visible Cameras for the Application of Pedestrian Detection." Signal Image and Video Processing, Vol. 11, 2017.
- V. John, Q. Long, Y. Xu, Z. Liu, S. Mita. "Sensor Fusion of Lidar and Stereo Camera without Calibration Objects." IEICE Trans. Fundamentals, Vol. E100-A, No. 2, 2017. [Invited Paper]
- B. Qi, V. John, Z. Liu, S. Mita. "Pedestrian Detection from Thermal Images: A Sparse Representation Based Approach." Infrared Physics and Technology, Vol. 76, 2016.
- V. John, K. Yoneda, Z. Liu, S. Mita. "Saliency Map Generation by Convolutional Neural Network for Real-time Traffic Light Detection." IEEE Trans. Computational Imaging, 1(3), 2015.
- Q. Xie, C. Ma, C. Guo, V. John, S. Mita, Q. Long. "Image Fusion Based on the △−1TV0 Energy Function." Entropy, Vol. 16, 2014.
- V. John, E. Trucco. "Charting-based subspace learning for video-based human action classification." Machine Vision and Applications, 25(1), 2014.
- V. John, E. Trucco, S. Ivekovic. "Markerless human articulated tracking using hierarchical particle swarm optimisation." Image and Vision Computing, 28(11), 2010.
Conference Papers
- T. Nguyen, Y. Kawanishi, V. John, T. Komamizu, I. Ide. "View-aware Cross-modal Distillation for Multi-view Action Recognition." IEEE/CVF WACV, 2026.
- E. Malathy, V. John, S. Selcia, S. Vikraman. "Bidirectional Translation System for Tamil Sign Language Recognition." Intelligent Vision and Computing, 2025.
- Y. Fang, V. John, Y. Kawanishi. "SemGest: A Multimodal Feature Space Alignment and Fusion Framework for Semantic-aware Co-speech Gesture Generation." GENEA Workshop, 2025.
- Z. Zhu et al., V. John. "Reproducibility Companion Paper: Enhancing Model Interpretability with Local Attribution over Global Exploration." ACM Multimedia, 2025.
- T. Nguyen, Y. Kawanishi, V. John, T. Komamizu, I. Ide. "MultiSensor-Home: A Wide-area Multimodal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion." IEEE FG, 2025. 🏆 Best Student Paper
- J. Y. Chen, V. John, Y. Kawanishi. "Cross-modal Emotion-specific Attention Model for Multimodal Emotion Recognition." IEEE FG, 2025.
- V. John, Y. Kawanishi. "Modelling Spatio-Temporal Dynamics by Graph Attention Network for Distributed Multi-Microphone Sound Event Classification." IEEE IAVSS, 2025.
- V. John, Y. Kawanishi. "Generating Pseudo-Strong Labels from Weak Labels for Multi-Source Sound Event Detection." ICPR, 2024.
- V. John, Y. Kawanishi. "Frame-Level Latent Embedding using Weak Labels for Multi-view Action Recognition." IEEE MIPR, 2024.
- V. John, Y. Kawanishi. "Combining Knowledge Distillation and Transfer Learning for Sensor Fusion in Visible and Thermal Camera-based Person Classification." MVA, 2023.
- V. John, Y. Kawanishi. "Multimodal Cascaded Framework with Metric Learning Robust to Missing Modalities for Person Classification." ACM MMSys, 2023.
- V. John, Y. Kawanishi. "Audio-Visual Sensor Fusion Framework using Person Attributes Robust to Missing Visual Modality for Person Recognition." MMM, 2023.
- V. John, Y. Kawanishi. "A Multimodal Sensor Fusion Framework Robust to Missing Modalities for Person Recognition." ACM Multimedia Asia, 2022.
- V. John, Y. Kawanishi. "Audio and Video-based Emotion Recognition using Multimodal Transformers." ICPR, 2022.
- V. John, A. Lakshmanan, S. Mita, A. Boyali, S. Thompson. "Deep Fusion-based Visible and Thermal Camera Forecasting using Seq2Seq GAN." IEEE IV Symposium, 2021.
- W. Li, V. John, S. Mita. "Enhancing Depth Quality of Stereo Vision using Deep Learning-based Prior Information of the Driving Environment." ICPR, 2021.
- V. John, S. Mita, A. Boyali, S. Thompson. "BVTNet: Multi-label Multi-class Fusion of Visible and Thermal Camera for Free Space and Pedestrian Segmentation." ICPR Workshop, 2021.
- V. John, S. Mita, A. Boyali, S. Thompson. "Visible and Thermal Camera-based Jaywalking Estimation using a Hierarchical Deep Learning Framework." ACCV Workshop, 2020.
- V. John, S. Mita. "RVNet: Deep Sensor Fusion of Monocular Camera and Radar for Real-Time Obstacle Detection in Challenging Environments." PSIVT, 2019.
- V. John, M. K. Nithilan, S. Mita, H. Tehrani, R. S. Sudheesh, P. P. Lalu. "SO-Net: Joint Semantic Segmentation and Obstacle Detection using Deep Fusion of Monocular Camera and Radar." PSIVT Workshops, 2019.
- A. Boyali, T. Akarman, N. Hashimoto, V. John. "Multi-Agent Reinforcement Learning for Autonomous On Demand Vehicles." IEEE IV Symposium, 2019.
- V. John, N. Karunakaran, C. Guo, K. Kidono, S. Mita. "Free Space, Visible and Missing Lane Marker Estimation using the PsiNet and Extra Trees Regression." ICPR, 2018.
- V. John, N. Karunakaran, S. Mita, H. Tehrani, M. Konishi, K. Ishimaru, Y. Xu. "Sensor Fusion of Intensity and Depth Cues using the ChiNet for Semantic Segmentation of Road Scenes." IEEE IV Symposium, 2018.
- S. Rathour, V. John, N. Karunakaran, S. Mita. "Vision and Dead Reckoning-based End-to-End Parking for Autonomous Vehicles." IEEE IV Symposium, 2018.
- V. John et al. "Vision-based Steering Angle Prediction by the Fusion of Depth and Intensity Deep Features." ICCVGIP, 2018.
- A. Boyali, L. Zheming, V. John, S. Mita, T. Yoshikawa. "Self-Scheduling Robust Preview Controllers for Path Tracking and Autonomous Vehicles." ITSC, 2017.
- V. John, Q. Long, Y. Xu, S. Mita. "Registration of GPS and Stereo Vision for Point Cloud Localization in Intelligent Vehicles Using Particle Swarm Optimization." ICSI, 2017.
- V. John, M. Umetsu, A. Boyali, S. Mita, N. Imanashi, S. Sanma. "Real-time Hand Posture and Gesture-based Touchless Automotive User Interface using Deep Learning." IEEE IV Symposium, 2017.
- V. John, S. Mita, H. Tehrani, K. Kidono. "Automated Driving by Monocular Camera Using Deep Mixture of Experts." IEEE IV Symposium, 2017. [Invited for IEEE Transactions]
- V. John, A. Boyali, S. Mita, M. Imanishi, N. Sanma. "Deep Learning-based Fast Hand Gesture Recognition using Representative Frames." DICTA, 2016.
- V. John, C. Guo, S. Mita, K. Kidono, H. Tehrani, K. Ishimaru. "Fast Road Scene Segmentation using Deep Learning and Scene-based Models." ICPR, 2016.
- V. John, L. Zheming, S. Mita. "Robust Traffic Light and Arrow Detection using Optimal Camera Parameters and GPS-based Prior." IEEE APIRS, 2016. 🏆 Best Image Processing Oral Presentation
- V. John, Z. Liu, C. Guo, S. Mita, K. Kidono. "Real-Time Lane Estimation using Deep Features and Extra Trees Regression." PSIVT, 2015. 🏆 Best Applications Paper
- V. John, Q. Long, Z. Liu, S. Mita. "Automatic Calibration and Registration of Lidar and Stereo Camera without Calibration Objects." IEEE IVES, 2015.
- V. John, Z. Liu, S. Mita, B. Qi. "Pedestrian Detection in Thermal Images Using Adaptive Fuzzy C-Means Clustering and Convolutional Neural Networks." MVA, 2015. 🏆 Most Influential Paper of the Decade, 2025
- V. John, K. Yoneda, Z. Liu, S. Mita, B. Qi. "Traffic Light Recognition in Varying Illumination using Deep Learning and Saliency Map." IEEE ITSC, 2014.
- B. Qi, V. John, Z. Liu, S. Mita. "Pedestrian Detection from Thermal Images with A Scattered Difference of Directional Gradients Feature Descriptor." IEEE ITSC, 2014.
- B. Qi, V. John, Z. Liu, S. Mita. "Use of Sparse Representation for Pedestrian Detection in Thermal Images." CVPR Workshops, 2014.
- V. John, G. Englebienne, B. Krose. "Solving Person Re-identification using Efficient Gibbs Sampling." BMVC, 2013. [Oral, 7% acceptance rate]
- V. John, G. Englebienne, B. Krose. "Person Re-identification using Height-based Gait in Colour Depth Cameras." ICIP, 2013.
- V. John, G. Englebienne, B. Krose. "Relative Camera Localisation in Non-Overlapping Camera Networks using Multiple Trajectories." ECCV Workshops, 2012.
- V. John, E. Trucco, S. J. McKenna. "Markerless human motion capture using charting and manifold constrained particle swarm optimisation." BMVC Workshop, 2010. 🏆 Best Paper
- S. Ivekovic, V. John, E. Trucco. "Markerless Multi-view Articulated Pose Estimation Using Adaptive Hierarchical Particle Swarm Optimisation." EvoApplications, 2010.
- V. John, E. Trucco. "Multiple view human articulated tracking using charting and particle swarm optimisation." ACM MM 3DVP Workshop, 2010.
- V. John, S. Ivekovic, E. Trucco. "Articulated Human Motion Tracking with HPSO." VISSAPP, 2009.