Ngan Le
Affiliation confirmed via AI analysis of OpenAlex, ORCID, and web sources.
Associate Professor
Research Areas
Biomedical Subjects
Links
Biography and Research Information
OverviewAI-generated summary
Ngan Le is an Associate Professor and Director of the Artificial Intelligence & Computer Vision (AICV) Lab in the Department of Electrical Engineering & Computer Science at the University of Arkansas. Her research interests lie in robotics, machine learning, computer vision, and medical analysis, with a focus on trusted decision-making, handling imperfect data, and real-time applications on edge devices. Dr. Le has expertise in processing various data modalities, including image, video, point cloud, volumetric data, time series, and remote sensing data, and her work spans image processing, scene understanding, multiple object tracking, behavior analysis, medical image analysis, and 3D reconstruction.
Her federal grant funding includes three awards totaling $6,242,025. These include a CAREER award from the NSF for a multimodal framework for video analytics, and two Co-PI awards from the NSF Convergence Accelerator Track J for projects focused on regional food systems and data-driven agriculture. Dr. Le has published 244 works with 3,950 citations and an h-index of 27. She leads a research group and collaborates with several faculty members at the University of Arkansas, including Trong Thang Pham, Taisei Hanyu, Khoa Vo, and Huyen Tran.
Metrics
- h-index: 27
- Publications: 244
- Citations: 3,950
Positions
-
Associate Professor 2025–presentUniversity of Arkansas at Fayetteville Electrical & Computer Engineering ORCID
-
Assistant Professor 2023–presentUniversity of Arkansas Department of Electrical Engineering and Computer Science ORCID
-
Postdoc 2018–2019Carnegie Mellon University Electrical & Computer Engineering ORCID
Selected Publications
-
FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding (2026)
-
Learning to Adapt and Calibrate: Score Distribution Alignment for Few-Shot Uncertainty Prediction in Medical VLMs (2026)
-
GloVLA: Let Geometry Move and Local VLA Interact for Robust Object-Centric Manipulation in Unstructured Environments (2026)
-
ARGUSTRACK: A Multi-View Annotation System for Multi-Object Tracking (2026)
-
OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs (2026)
-
Self-Improving VLA Policies: Selected Diffusion Noise for Spurious-Robust Action Smoothing (2026)arXiv (Cornell University) OpenAlex
-
Self-Improving VLA Policies: Selected Diffusion Noise for Spurious-Robust Action Smoothing (2026)
-
VentVision: A multimodal vision-based system for automated vent-based chick sexing (2026)
-
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models (2026)arXiv (Cornell University) OpenAlex
-
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models (2026)
-
SSL-MTab: Self-Supervised Distillation for Missing Data in Tabular Prediction Tasks (2026)
-
MiGa: Multi-chicken gait assessment (2026)
-
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding (2026)arXiv (Cornell University) OpenAlex
-
GazeQwen: Lightweight Gaze-Conditioned LLM Modulation for Streaming Video Understanding (2026)
-
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images (2026)arXiv (Cornell University) OpenAlex
Federal Grants 3 $6,242,025 total
NSF Convergence Accelerator Track J Phase 2: Cultivate IQ - Empowering Regional Food Systems
CAREER: Trustworthy, Robust, and Efficient Multimodal Framework for Video Analytics.
Collaboration Network
Top Collaborators
- Deep reinforcement learning in computer vision: a comprehensive survey
- Non-volume preserving-based fusion to group-level emotion recognition on crowd videos
- Multi-camera multi-object tracking on the move via single-stage global association approach
- VLCAP: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning
- Narrow Band Active Contour Attention Model for Medical Segmentation
Showing 5 of 19 shared publications
- Spiking Neural Networks and Their Applications: A Review
- Deep reinforcement learning in computer vision: a comprehensive survey
- CLIP-TSA: Clip-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection
- AerialFormer: Multi-Resolution Transformer for Aerial Image Segmentation
- VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
Showing 5 of 15 shared publications
- Spiking Neural Networks and Their Applications: A Review
- CLIP-TSA: Clip-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection
- AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
- VLCAP: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning
- ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection
Showing 5 of 14 shared publications
- MEGANet: Multi-Scale Edge-Guided Attention Network for Weak Boundary Polyp Segmentation
- AOE-Net: Entities Interactions Modeling with Adaptive Attention Mechanism for Temporal Action Proposals Generation
- EmbryosFormer: Deformable Transformer and Collaborative Encoding-Decoding for Embryos Stage Development Classification
- ABN: Agent-Aware Boundary Networks for Temporal Action Proposal Generation
- Agent-Environment Network for Temporal Action Proposal Generation
Showing 5 of 11 shared publications
- Bridging human and machine intelligence: Reverse-engineering radiologist intentions for clinical trust and adoption
- Deep reinforcement learning in medical imaging: A literature review
- Enhance Portable Radiograph for Fast and High Accurate COVID-19 Monitoring
- ItpCtrl-AI: End-to-end interpretable and controllable artificial intelligence by modeling radiologists’ intentions
- Correction to: Interpretable and Annotation-Efficient Learning for Medical Image Computing
Showing 5 of 9 shared publications
- Language-Conditioned Affordance-Pose Detection in 3D Point Clouds
- Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
- Lightweight Language-driven Grasp Detection using Conditional Consistency Model
- Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation
- Language-driven Grasp Detection with Mask-guided Attention
Showing 5 of 9 shared publications
- Open-Vocabulary Affordance Detection in 3D Point Clouds
- Language-Conditioned Affordance-Pose Detection in 3D Point Clouds
- Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
- Lightweight Language-driven Grasp Detection using Conditional Consistency Model
- Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation
Showing 5 of 7 shared publications
- Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance
- Language-driven Grasp Detection with Mask-guided Attention
- WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
- Style Transfer for 2D Talking Head Generation
- Guide3D: A Bi-planar X-ray Dataset for 3D Shape Reconstruction
Showing 5 of 7 shared publications
- I-AI: A Controllable & Interpretable AI System for Decoding Radiologists’ Intense Focus for Accurate CXR Diagnoses
- FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
- Style Transfer for 2D Talking Head Generation
- GazeSearch: Radiology Findings Search Benchmark
- CattleFever: An automated cattle fever estimation system
Showing 5 of 7 shared publications
- Bridging human and machine intelligence: Reverse-engineering radiologist intentions for clinical trust and adoption
- FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
- ItpCtrl-AI: End-to-end interpretable and controllable artificial intelligence by modeling radiologists’ intentions
- GazeSearch: Radiology Findings Search Benchmark
- Modeling radiologists’ cognitive processes using a digital gaze twin to enhance radiology training
Showing 5 of 7 shared publications
- Non-volume preserving-based fusion to group-level emotion recognition on crowd videos
- Multi-camera multi-object tracking on the move via single-stage global association approach
- Active Contour Model in Deep Learning Era: A Revise and Review
- LIAAD: Lightweight attentive angular distillation for large-scale age-invariant face recognition
- Domain Generalization via Universal Non-volume Preserving Approach
Showing 5 of 6 shared publications
- AerialFormer: Multi-Resolution Transformer for Aerial Image Segmentation
- VLCAP: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning
- Land8Fire: A Complete Study on Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, and Extensive Benchmarking
- RSSep: Sequence-to-Sequence Model for Simultaneous Referring Remote Sensing Segmentation and Detection
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
Showing 5 of 6 shared publications
- Open-Vocabulary Affordance Detection in 3D Point Clouds
- Language-Conditioned Affordance-Pose Detection in 3D Point Clouds
- Lightweight Language-driven Grasp Detection using Conditional Consistency Model
- Open-Vocabulary Affordance Detection using Knowledge Distillation and Text-Point Correlation
- Language-driven Grasp Detection with Mask-guided Attention
Showing 5 of 6 shared publications
- Non-volume preserving-based fusion to group-level emotion recognition on crowd videos
- Multi-camera multi-object tracking on the move via single-stage global association approach
- Active Contour Model in Deep Learning Era: A Revise and Review
- A Multi-task Contextual Atrous Residual Network for Brain Tumor Detection & Segmentation
- LIAAD: Lightweight attentive angular distillation for large-scale age-invariant face recognition
- Teaching Yourself: A Self-Knowledge Distillation Approach to Action Recognition
- (2+1)D Distilled ShuffleNet: A Lightweight Unsupervised Distillation Network for Human Action Recognition
- Deep Learning for Human Action Recognition: A Comprehensive Review
- Self-Supervised Learning via multi-Transformation Classification for Action Recognition
- Self-Supervised Learning via Multi-Transformation Classification for Action Recognition
Similar Researchers
Based on overlapping research topics