Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6301-6320 of 9,853 articles

Paleoinspired Vision: From Exploring Colour Vision Evolution to Inspiring Camera Design

The evolution of colour vision is captivating, as it reveals the adaptive strategies of extinct species while simultaneously inspiring innovations in modern imaging technology. In this study, we present a simplified model of visual transduction in the retina, introducing a novel opsin layer. We quantify evolutionary pressures by measuring machine vision recognition accuracy on colour images shap...

CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models

Recent advances in Large Vision-Language Models (LVLMs) have enabled general-purpose vision tasks through visual instruction tuning. While existing LVLMs can generate segmentation masks from text prompts for single images, they struggle with segmentation-grounded reasoning across images, especially at finer granularities such as object parts. In this paper, we introduce the new task of part-focu...

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment

Current multimodal large language models (MLLMs) struggle with fine-grained or precise understanding of visuals though they give comprehensive perce...

Multi-Head Attention Driven Dynamic Visual-Semantic Embedding for Enhanced Image-Text Matching

With the rapid development of multimodal learning, the image-text matching task, as a bridge connecting vision and language, has become increasingly...

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

Open-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representatio...

WeatherGS: 3D Scene Reconstruction in Adverse Weather Conditions via Gaussian Splatting

3D Gaussian Splatting (3DGS) has gained significant attention for 3D scene reconstruction, but still suffers from complex outdoor environments, espe...

Unified Local and Global Attention Interaction Modeling for Vision Transformers

We present a novel method that extends the self-attention mechanism of a vision transformer (ViT) for more accurate object detection across diverse ...

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

Multi-head self-attention (MHSA) is a key component of Transformers, a widely popular architecture in both language and vision. Multiple heads intui...

ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science

Street view imagery (SVI) has been instrumental in many studies in the past decade to understand and characterize street features and the built envi...

Expand VSR Benchmark for VLLM to Expertize in Spatial Rules

Distinguishing spatial relations is a basic part of human cognition which requires fine-grained perception on cross-instance. Although benchmarks li...

Semantics Disentanglement and Composition for Versatile Codec toward both Human-eye Perception and Machine Vision Task

While learned image compression methods have achieved impressive results in either human visual perception or machine vision tasks, they are often s...

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective

Recent Large Vision-Language Models (LVLMs) have shown promising reasoning capabilities on text-rich images from charts, tables, and documents. Howe...

Kernel-Aware Graph Prompt Learning for Few-Shot Anomaly Detection

Few-shot anomaly detection (FSAD) aims to detect unseen anomaly regions with the guidance of very few normal support images from the same class. Exi...

V$^2$-SfMLearner: Learning Monocular Depth and Ego-motion for Multimodal Wireless Capsule Endoscopy

Deep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization...

Retention Score: Quantifying Jailbreak Risks for Vision Language Models

The emergence of Vision-Language Models (VLMs) is a significant advancement in integrating computer vision with Large Language Models (LLMs) to enha...

Multimodal Preference Data Synthetic Alignment with Reward Model

Multimodal large language models (MLLMs) have significantly advanced tasks like caption generation and visual question answering by integrating visu...

Comparison of Spatiotemporal Characteristics of Eye Movements in Non-experts and the Skill Transfer Effects of Gaze Guidance and Annotation Guidance

Methods for converting the tacit knowledge of experts into explicit knowledge have drawn increasing attention. Gaze data has emerged as a valuable a...

ChromaGazer: Unobtrusive Visual Modulation using Imperceptible Color Vibration for Visual Guidance

Visual guidance (VG) is critical for directing user attention in virtual and augmented reality applications. However, conventional methods using exp...

GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image Captioning

Retinal image analysis is crucial for diagnosing and treating eye diseases, yet generating accurate medical reports from images remains challenging ...

VilBias: A Study of Bias Detection through Linguistic and Visual Cues , presenting Annotation Strategies, Evaluation, and Key Challenges

The integration of Large Language Models (LLMs) and Vision-Language Models (VLMs) opens new avenues for addressing complex challenges in multimodal ...

Browse Categories