Khoa Vo
This is a likely match — the affiliation was inferred from OpenAlex, ORCID, and web sources but has not been fully confirmed. Treat with appropriate caution.
Researcher
Also affiliated: Vietnam National University Ho Chi Minh City (2024–2025); Medical University of South Carolina (2006); University of Wollongong (2014–2015); CRC for Rail Innovation (2014–2015); Ho Chi Minh City University of Technology (2024–2025); Sungkyunkwan University (2025)
Faculty Researcher
Research Areas
Biography and Research Information
OverviewAI-generated summary
Khoa Vo's research centers on advancing machine learning techniques, particularly within the domains of computer vision and natural language processing. His work investigates the development of transformer-based models for complex tasks such as video paragraph captioning, amodal instance segmentation, and weakly-supervised video anomaly detection. Vo has co-authored publications exploring visual-linguistic transformers, multimodal fusion for 3D scene representation, and contextual explainable video representations that align with human perception.
His recent publications include "VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning" (2023, 2022), "AISFormer: Amodal Instance Segmentation with Transformer" (2022), and "Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation" (2024). Vo collaborates with researchers at the University of Arkansas at Fayetteville, including Taisei Hanyu, Chase Rainwater, Ngan Le, and Anthony L. Gunderman, with whom he has co-authored multiple publications. His research has garnered 315 citations and an h-index of 8 across 26 publications.
Metrics
- h-index: 8
- Publications: 26
- Citations: 319
Selected Publications
-
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective (2026)
-
CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling (2025)
-
VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning (2023)
Collaboration Network
Top Collaborators
- VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
- VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
- VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
Similar Researchers
Based on overlapping research topics