Latest AI and machine learning research in ophthalmology for healthcare professionals.
In this paper, we propose BeamLLM, a vision-aided millimeter-wave (mmWave) beam prediction framework leveraging large language models (LLMs) to address the challenges of high training overhead and latency in mmWave communication systems. By combining computer vision (CV) with LLMs' cross-modal reasoning capabilities, the framework extracts user equipment (UE) positional features from RGB images ...
Analyzing animal behavior is crucial in advancing neuroscience, yet quantifying and deciphering its intricate dynamics remains a significant challenge. Traditional machine vision approaches, despite their ability to detect spontaneous behaviors, fall short due to limited interpretability and reliance on manual labeling, which restricts the exploration of the full behavioral spectrum. Here, we in...
Accurate staging of Diabetic Retinopathy (DR) is essential for guiding timely interventions and preventing vision loss. However, current staging mod...
In Daugman-style iris recognition, the textures of the left and right irises of the same person are traditionally considered as being as different a...
Vision-Language Models (VLMs) based on Mixture-of-Experts (MoE) architectures have emerged as a pivotal paradigm in multimodal understanding, offeri...
Climate change is one of the most pressing challenges of the 21st century, sparking widespread discourse across social media platforms. Activists, p...
With the development of robotics technology, some tactile sensors, such as vision-based sensors, have been applied to contact-rich robotics tasks. H...
Foundation models have emerged as a powerful paradigm in computational pathology (CPath), enabling scalable and generalizable analysis of histopatho...
Vision Transformer models exhibit immense power yet remain opaque to human understanding, posing challenges and risks for practical applications. Wh...
Eye tracking has been found to be useful in various tasks including diagnostic and screening tools. However, traditional eye trackers had a complica...
Visual attention plays a critical role when our visual system executes active visual tasks by interacting with the physical scene. However, how to e...
Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textual prom...
Deep learning models based on graph neural networks have emerged as a popular approach for solving computer vision problems. They encode the image i...
Despite their success, Large Vision-Language Models (LVLMs) remain vulnerable to hallucinations. While existing studies attribute the cause of hallu...
Human-centric visual perception (HVP) has recently achieved remarkable progress due to advancements in large-scale self-supervised pretraining (SSP)...
As the computational needs of Large Vision-Language Models (LVLMs) increase, visual token pruning has proven effective in improving inference speed ...
This paper explores vision-based localization through a biologically-inspired approach that mirrors how humans and animals link views or perspective...
Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase ...
Visual instruction tuning (VIT) for large vision-language models (LVLMs) requires training on expansive datasets of image-instruction pairs, which c...
Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current Large Vision-Language Models (L...