Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6001-6020 of 9,853 articles

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

The incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a cropping-based approach to process images, which leads to fragmented visual encoding and a sharp increase in redundant tokens. To tackle these issues, we propose the FALCON model...

PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer

Source camera identification has emerged as a vital solution to unlock incidents involving critical cases like terrorism, violence, and other criminal activities. The ability to trace the origin of an image/video can aid law enforcement agencies in gathering evidence and constructing the timeline of events. Moreover, identifying the owner of a certain device narrows down the area of search in a ...

CILP-FGDI: Exploiting Vision-Language Model for Generalizable Person Re-Identification

The Visual Language Model, known for its robust cross-modal capabilities, has been extensively applied in various computer vision tasks. In this pap...

MM-Retinal V2: Transfer an Elite Knowledge Spark into Fundus Vision-Language Pretraining

Vision-language pretraining (VLP) has been investigated to generalize across diverse downstream tasks for fundus image analysis. Although recent met...

Leveraging Video Vision Transformer for Alzheimer's Disease Diagnosis from 3D Brain MRI

Alzheimer's disease (AD) is a neurodegenerative disorder affecting millions worldwide, necessitating early and accurate diagnosis for optimal patien...

Vision-Aided Channel Prediction Based on Image Segmentation at Street Intersection Scenarios

Intelligent vehicular communication with vehicle road collaboration capability is a key technology enabled by 6G, and the integration of various vis...

DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions

By overlaying time-synced user comments on videos, Danmu creates a co-watching experience for online viewers. However, its visual-centric design pos...

FlatTrack: Eye-tracking with ultra-thin lensless cameras

Existing eye trackers use cameras based on thick compound optical elements, necessitating the cameras to be placed at focusing distance from the eye...

Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis

Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researc...

Vision Aided Channel Prediction for Vehicular Communications: A Case Study of Received Power Prediction Using RGB Images

The communication scenarios and channel characteristics of 6G will be more complex and difficult to characterize. Conventional methods for channel p...

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models

This work investigates the capabilities of current vision-language models (VLMs) in visual understanding and attribute measurement of primitive shap...

Complementary Subspace Low-Rank Adaptation of Vision-Language Models for Few-Shot Classification

Vision language model (VLM) has been designed for large scale image-text alignment as a pretrained foundation model. For downstream few shot classif...

What if Eye...? Computationally Recreating Vision Evolution

Vision systems in nature show remarkable diversity, from simple light-sensitive patches to complex camera eyes with lenses. While natural selection ...

Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing

Vision language models have achieved impressive results across various fields. However, adoption in remote sensing remains limited, largely due to t...

Approach to Designing CV Systems for Medical Applications: Data, Architecture and AI

This paper introduces an innovative software system for fundus image analysis that deliberately diverges from the conventional screening approach, o...

Improved Vessel Segmentation with Symmetric Rotation-Equivariant U-Net

Automated segmentation plays a pivotal role in medical image analysis and computer-assisted interventions. Despite the promising performance of exis...

Deep-BrownConrady: Prediction of Camera Calibration and Distortion Parameters Using Deep Learning and Synthetic Data

This research addresses the challenge of camera calibration and distortion parameter prediction from a single image using deep learning models. The ...

Characterizing Visual Intents for People with Low Vision through Eye Tracking

Accessing visual information is crucial yet challenging for people with low vision due to their visual conditions (e.g., low visual acuity, limited ...

Automatic detection and prediction of nAMD activity change in retinal OCT using Siamese networks and Wasserstein Distance for ordinality

Neovascular age-related macular degeneration (nAMD) is a leading cause of vision loss among older adults, where disease activity detection and progr...

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models

As the demand for high-resolution image processing in Large Vision-Language Models (LVLMs) grows, sub-image partitioning has become a popular approa...

Browse Categories