Latest AI and machine learning research in ophthalmology for healthcare professionals.
Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to a quadratic increase in the number of vision tokens and significantly raises computational costs. We attribute this inefficiency to substantial visual regions r...
Tabular data are central to biomedical research, from liquid biopsy and bulk and single-cell transcriptomics to electronic health records and phenotypic profiling. Unlike images or sequences, however, tabular datasets lack intrinsic spatial organization: features are treated as unordered dimensions, and their relationships must be inferred implicitly by the model. This limits the ability of vision...
Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large L...
Automated radiology report generation from 3D computed tomography (CT) volumes is challenging due to extreme sequence lengths, severe class imbalance,...
Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming vid...
Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. T...
Sparse Autoencoders uncover thousands of features in vision models, yet explaining these features without requiring human intervention remains an open...
Vision-Language Model (VLM)-based image quality assessment (IQA) has been significantly advanced by incorporating Chain-of-Thought (CoT) reasoning. Re...
Continual unlearning poses the challenge of enabling large vision-language models to selectively refuse specific image-instruction pairs in response t...
We present CataractSAM-2, a domain-adapted extension of Meta's Segment Anything Model 2, designed for real-time semantic segmentation of cataract opht...
Large Vision-Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Althou...
Steel surface defect detection is essential for ensuring product quality and reliability in modern manufacturing. Current methods often rely on basic ...
Optical coherence tomography (OCT) is a non-invasive volumetric imaging modality with high spatial and temporal resolution. For imaging larger tissue ...
Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are pr...
Recent advancements in video generation models have significantly improved their ability to follow text prompts. However, the customization of dynamic...
Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to associate objects with their p...
Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresented ...
In this paper, we present CornOrb, a publicly accessible multimodal dataset of Orbscan corneal topography images and clinical annotations collected fr...
This paper presents the development of a documented program capable of solving idealized beam models, such as those commonly used in textbooks and aca...
Monocular depth estimation remains challenging for transparent objects, where refraction and transmission are difficult to model and break the appeara...