Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4561-4580 of 9,853 articles

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods require explicit safety labels or contrastive data; yet, threat-related concepts are concrete and visually depictable, while safety concepts, like helpfulness, are abstract and lack visual referents. Inspired by the Self-Fulfilling mechanism underlying em...

Mar 9 2026 2603.08486v1

Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models

Vision-Language Models achieve near-perfect accuracy at reading text in images, yet prove largely typography-blind: capable of recognizing what text says, but not how it looks. We systematically investigate this gap by evaluating font family, size, style, and color recognition across 26 fonts, four scripts, and three difficulty levels. Our evaluation of 15 state-of-the-art VLMs reveals a striking ...

Mar 9 2026 2603.08497v1
FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection

In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or t...

Mar 9 2026 2603.08611v1
UNBOX: Unveiling Black-box visual models with Natural-language

Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet modern ...

Mar 9 2026 2603.08639v1
Prompt-Based Caption Generation for Single-Tooth Dental Images Using Vision-Language Models

Digital dentistry has made significant advances with the advent of deep learning. However, the majority of these deep learning-based dental image anal...

Mar 8 2026 2603.07403v1
Looking Into the Water by Unsupervised Learning of the Surface Shape

We address the problem of looking into the water from the air, where we seek to remove image distortions caused by refractions at the water surface. O...

Mar 8 2026 2603.07614v1
GazeShift: Unsupervised Gaze Estimation and Dataset for VR

Gaze estimation is instrumental in modern virtual reality (VR) systems. Despite significant progress in remote-camera gaze estimation, VR gaze researc...

Mar 8 2026 2603.07832v1
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval

Dense image retrieval is accurate but offers limited interpretability and attribution, and it can be compute-intensive at scale. We present \textbf{BM...

Mar 6 2026 2603.05781v1
EventGeM: Global-to-Local Feature Matching for Event-Based Visual Place Recognition

Dynamic vision sensors, also known as event cameras, are rapidly rising in popularity for robotic and computer vision tasks due to their sparse activa...

Mar 6 2026 2603.05807v1
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues

Vision-Language Models (VLMs) have achieved remarkable progress on a wide range of challenging multimodal understanding and reasoning tasks. However, ...

Mar 6 2026 2603.05869v1
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models

Visual token reduction is critical for accelerating Vision-Language Models (VLMs), yet most existing approaches rely on a fixed budget shared across a...

Mar 6 2026 2603.05950v1
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models

Recent advances in large-scale pretrained vision models have demonstrated impressive capabilities across a wide range of downstream tasks, including c...

Mar 6 2026 2603.05963v1
MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs

Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time intervals across va...

Mar 6 2026 2603.05997v1
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision

Reasoning-based text-to-image (T2I) generation requires models to interpret complex prompts accurately. Existing reasoning frameworks can be broadly c...

Mar 6 2026 2603.06032v1
Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving

The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging their reasoning...

Mar 6 2026 2603.06054v1
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models

As Vision-Language Models (VLMs) become increasingly sophisticated and widely used, it becomes more and more crucial to understand their decision-maki...

Mar 6 2026 2603.06302v1
Non-invasive Growth Monitoring of Small Freshwater Fish in Home Aquariums via Stereo Vision

Monitoring fish growth behavior provides relevant information about fish health in aquaculture and home aquariums. Yet, monitoring fish sizes poses di...

Mar 6 2026 2603.06421v1
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization

Autoregressive (AR) language models rely on causal tokenization, but extending this paradigm to vision remains non-trivial. Current visual tokenizers ...

Mar 6 2026 2603.06449v1
Causal Interpretation of Neural Network Computations with Contribution Decomposition

Understanding how neural networks transform inputs into outputs is crucial for interpreting and manipulating their behavior. Most existing approaches ...

Mar 6 2026 2603.06557v1
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Vision Language Model (VLM) development has largely relied on scaling model size, which hinders deployment on compute-constrained mobile and edge devi...

Mar 6 2026 2603.06569v1
Browse Categories