Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5701-5720 of 9,853 articles

BeamLLM: Vision-Empowered mmWave Beam Prediction with Large Language Models

In this paper, we propose BeamLLM, a vision-aided millimeter-wave (mmWave) beam prediction framework leveraging large language models (LLMs) to address the challenges of high training overhead and latency in mmWave communication systems. By combining computer vision (CV) with LLMs' cross-modal reasoning capabilities, the framework extracts user equipment (UE) positional features from RGB images ...

MouseGPT: A Large-scale Vision-Language Model for Mouse Behavior Analysis

Analyzing animal behavior is crucial in advancing neuroscience, yet quantifying and deciphering its intricate dynamics remains a significant challenge. Traditional machine vision approaches, despite their ability to detect spontaneous behaviors, fall short due to limited interpretability and reliance on manual labeling, which restricts the exploration of the full behavioral spectrum. Here, we in...

Fine-tuning Vision Language Models with Graph-based Knowledge for Explainable Medical Image Analysis

Accurate staging of Diabetic Retinopathy (DR) is essential for guiding timely interventions and preventing vision loss. However, current staging mod...

A Siamese Network to Detect If Two Iris Images Are Monozygotic

In Daugman-style iris recognition, the textures of the left and right irises of the same person are traditionally considered as being as different a...

Astrea: A MOE-based Visual Understanding Model with Progressive Alignment

Vision-Language Models (VLMs) based on Mixture-of-Experts (MoE) architectures have emerged as a pivotal paradigm in multimodal understanding, offeri...

Deep Learning for Climate Action: Computer Vision Analysis of Visual Narratives on X

Climate change is one of the most pressing challenges of the 21st century, sparking widespread discourse across social media platforms. Activists, p...

Tacchi 2.0: A Low Computational Cost and Comprehensive Dynamic Contact Simulator for Vision-based Tactile Sensors

With the development of robotics technology, some tactile sensors, such as vision-based sensors, have been applied to contact-rich robotics tasks. H...

Multi-Modal Foundation Models for Computational Pathology: A Survey

Foundation models have emerged as a powerful paradigm in computational pathology (CPath), enabling scalable and generalizable analysis of histopatho...

Discovering Influential Neuron Path in Vision Transformers

Vision Transformer models exhibit immense power yet remain opaque to human understanding, posing challenges and risks for practical applications. Wh...

The Detection of Saccadic Eye Movements and Per-Eye Comparisons using Virtual Reality Eye Tracking Devices

Eye tracking has been found to be useful in various tasks including diagnostic and screening tools. However, traditional eye trackers had a complica...

Visual Attention Graph

Visual attention plays a critical role when our visual system executes active visual tasks by interacting with the physical scene. However, how to e...

Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models

Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textual prom...

i-WiViG: Interpretable Window Vision GNN

Deep learning models based on graph neural networks have emerged as a popular approach for solving computer vision problems. They encode the image i...

Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation

Despite their success, Large Vision-Language Models (LVLMs) remain vulnerable to hallucinations. While existing studies attribute the cause of hallu...

Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models

Human-centric visual perception (HVP) has recently achieved remarkable progress due to advancements in large-scale self-supervised pretraining (SSP)...

Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models

As the computational needs of Large Vision-Language Models (LVLMs) increase, visual token pruning has proven effective in improving inference speed ...

STRMs: Spatial Temporal Reasoning Models for Vision-Based Localization Rivaling GPS Precision

This paper explores vision-based localization through a biologically-inspired approach that mirrors how humans and animals link views or perspective...

Should VLMs be Pre-trained with Image Data?

Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase ...

Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning

Visual instruction tuning (VIT) for large vision-language models (LVLMs) requires training on expansive datasets of image-instruction pairs, which c...

When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning

Efficient vision-language understanding of large Remote Sensing Images (RSIs) is meaningful but challenging. Current Large Vision-Language Models (L...

Browse Categories