Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 3841-3860 of 9,853 articles

HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams

Much of the recent progress in image and video recognition has come at the cost of memory: larger models, increased resolution, and longer temporal contexts. An inevitable component is the quadratic (or larger) growth of memory and compute based on image resolution, which is a property of the grid sampling used in convolutional networks and vision transformers. In this work we study residual netwo...

Aug 3 2026 2608.02140v2

In-Context Collapse in Vision-Language Models and How to Mitigate it?

Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demonstrations are supplied. We show the opposite: as demonstrations accumulate, a subset of VLMs undergo an \emph{in-context collapse}, a sharp, sometimes catastrophic accuracy drop spanning synthetic classification, natural...

Aug 3 2026 2608.02830v1
FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering

We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answ...

Aug 3 2026 2608.01664v1
SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries requi...

Aug 3 2026 2608.01709v1
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability ...

Aug 3 2026 2608.01827v1
Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection

Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insuffici...

Aug 3 2026 2608.01930v1
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation

Multimodal large models are increasingly used to generate scalable vector graphics (SVG), but reliable evaluation remains underexplored. Existing prot...

Aug 3 2026 2608.01977v1
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, ac...

Aug 3 2026 2608.02039v1
Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression

Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens i...

Aug 3 2026 2608.02109v1
HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models

Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spec...

Aug 3 2026 2608.02124v1
Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression

Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and ...

Aug 3 2026 2608.02134v1
Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attac...

Aug 3 2026 2608.02137v1
HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams

Much of the recent progress in image and video recognition has come at the cost of memory: larger models, increased resolution, and longer temporal co...

Aug 3 2026 2608.02140v1
DerainSplat: Feed-Forward Clean 3D Gaussian Splatting from Sparse Rainy Views

Although image deraining has advanced substantially, existing methods mainly focus on 2D image restoration. As spatial intelligence applications such ...

Aug 3 2026 2608.02191v1
MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compa...

Aug 3 2026 2608.02449v1
Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models

Fine-grained visual representations are essential for medical image analysis, particularly when diagnostically relevant evidence is subtle and spatial...

Aug 2 2026 2608.00976v1
It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers po...

Aug 2 2026 2608.01207v1
Visual Distribution Anchoring for Efficient Prompt Tuning

Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textua...

Jul 31 2026 2607.28967v1
DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of annotated da...

Jul 31 2026 2607.29337v1
QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering...

Jul 31 2026 2607.29445v1
Browse Categories