Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4401-4420 of 9,853 articles

Universal computational thermal imaging overcoming the ghosting effect

Thermal imaging is crucial for night vision but fundamentally hampered by the ghosting effect, a loss of detailed texture in cluttered photon streams. While conventional ghosting mitigation has relied on data post-processing, the recent breakthrough in heat-assisted detection and ranging (HADAR) opens a promising frontier for hyperspectral computational thermal imaging that produces night vision w...

Apr 2 2026 2604.01542v1

Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning

Large Vision-Language Models (LVLMs) have achieved remarkable proficiency in explicit visual recognition, effectively describing what is directly visible in an image. However, a critical cognitive gap emerges when the visual input serves only as a clue rather than the answer. We identify that current models struggle with the complex, multi-step reasoning required to solve problems where informatio...

Apr 2 2026 2604.01764v1
Enhancing Medical Visual Grounding via Knowledge-guided Spatial Prompts

Medical Visual Grounding (MVG) aims to identify diagnostically relevant phrases from free-text radiology reports and localize their corresponding regi...

Apr 2 2026 2604.01915v1
Curia-2: Scaling Self-Supervised Learning for Radiology Foundation Models

The rapid growth of medical imaging has fueled the development of Foundation Models (FMs) to reduce the growing, unsustainable workload on radiologist...

Apr 2 2026 2604.01987v1
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline

Transformer-based architectures have become the shared backbone of natural language processing and computer vision. However, understanding how these m...

Apr 2 2026 2604.02182v1
Steerable Visual Representations

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such ...

Apr 2 2026 2604.02327v1
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a cha...

Mar 30 2026 2603.27942v1
SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering

A sketch is a distilled form of visual abstraction that conveys core concepts through simplified yet purposeful strokes while omitting extraneous deta...

Mar 30 2026 2603.28363v1
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains

The connoisseurship of antique Chinese porcelain demands extensive historical expertise, material understanding, and aesthetic sensitivity, making it ...

Mar 30 2026 2603.28474v1
Domain-Invariant Prompt Learning for Vision-Language Models

Large pre-trained vision-language models like CLIP have transformed computer vision by aligning images and text in a shared feature space, enabling ro...

Mar 30 2026 2603.28555v1
XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

Vision-language models (VLMs) rely on a shared visual-textual representation space to perform tasks such as zero-shot classification, image captioning...

Mar 30 2026 2603.28568v1
Streamlined Open-Vocabulary Human-Object Interaction Detection

Open-vocabulary human-object interaction (HOI) detection aims to localize and recognize all human-object interactions in an image, including those uns...

Mar 29 2026 2603.27500v1
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typical...

Mar 29 2026 2603.27577v1
OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery

Open-vocabulary change detection (OVCD) seeks to recognize arbitrary changes of interest by enabling generalization beyond a fixed set of predefined c...

Mar 29 2026 2603.27645v1
When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models

Visual-Language Models (VLMs) have demonstrated exceptional cross-modal understanding across various tasks, including zero-shot classification, image ...

Mar 29 2026 2603.27759v1
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants

While powerful in image-conditioned generation, multimodal large language models (MLLMs) can display uneven performance across demographic groups, hig...

Mar 27 2026 2603.26008v1
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation

Human age estimation from facial images represents a challenging computer vision task with significant applications in biometrics, healthcare, and hum...

Mar 27 2026 2603.26015v1
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays

Despite recent advances in medical vision-language pretraining, existing models still struggle to capture the diagnostic workflow: radiographs are typ...

Mar 27 2026 2603.26049v1
TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life

Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle v...

Mar 27 2026 2603.26128v1
Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones

Vision backbone networks play a central role in modern computer vision. Enhancing their efficiency directly benefits a wide range of downstream applic...

Mar 27 2026 2603.26551v1
Browse Categories