Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3821-3840 of 9,853 articles

Visual Grounding in Zero-Shot Vision-Language Control

Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis ref...

Aug 6 2026 2608.06154v1

Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by the notorious Nazi propaganda picture book \emph{Der Giftpilz}. Recently, frontier text-to-image (T2I) systems such as Gemini and GPT-Image have enabled conversational generation with consistent characters and scenes across turns, making hateful visu...

Aug 5 2026 2608.05210v1
A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from...

Aug 5 2026 2608.05260v1
IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers

Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it rema...

Aug 5 2026 2608.05122v2
Adapting Vision Foundation Models with Cascaded Semantics

Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) ada...

Aug 5 2026 2608.05393v1
ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely inc...

Aug 5 2026 2608.04385v1
EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment

The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-sup...

Aug 5 2026 2608.04472v1
Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles

Vision-language models (VLMs) process an image as a sequence of visual tokens, which creates a substantial computational bottleneck during inference. ...

Aug 5 2026 2608.04483v1
DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models

Visual inputs in vision-language models (VLMs) are often encoded into substantially longer token sequences than text, making visual tokens a major bot...

Aug 5 2026 2608.04496v1
GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction

Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visu...

Aug 5 2026 2608.04504v1
Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are availabl...

Aug 5 2026 2608.04554v1
Simile Understanding in Text-to-Image Models: An Evaluation Framework

Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce ...

Aug 5 2026 2608.04750v1
IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers

Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it rema...

Aug 5 2026 2608.05122v1
RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the ...

Aug 4 2026 2608.04132v1
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks ...

Aug 4 2026 2608.04244v1
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to...

Aug 4 2026 2608.02980v1
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs

GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resol...

Aug 4 2026 2608.03270v1
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large paramet...

Aug 4 2026 2608.03580v1
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead...

Aug 4 2026 2608.03649v1
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, ac...

Aug 3 2026 2608.02039v2
Browse Categories