Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5861-5880 of 9,853 articles

A Training-Free Framework for Precise Mobile Manipulation of Small Everyday Objects

Many everyday mobile manipulation tasks require precise interaction with small objects, such as grasping a knob to open a cabinet or pressing a light switch. In this paper, we develop Servoing with Vision Models (SVM), a closed-loop training-free framework that enables a mobile manipulator to tackle such precise tasks involving the manipulation of small objects. SVM employs an RGB-D wrist camera...

A Chain-of-Thought Subspace Meta-Learning for Few-shot Image Captioning with Large Vision and Language Models

A large-scale vision and language model that has been pretrained on massive data encodes visual and linguistic prior, which makes it easier to generate images and language that are more natural and realistic. Despite this, there is still a significant domain gap between the modalities of vision and language, especially when training data is scarce in few-shot settings, where only very limited da...

Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images

Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting i...

Qwen2.5-VL Technical Report

We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational ...

Freezing of Gait as a Complication of Pallidal Deep Brain Stimulation in DYT- KMT2B Patients with Evidence of Striatonigral Degeneration

Background: Mutations in KMT2B are a recognized cause of early-onset complex dystonia, with deep brain stimulation (DBS) of the internal globus pall...

MobileViM: A Light-weight and Dimension-independent Vision Mamba for 3D Medical Image Analysis

Efficient evaluation of three-dimensional (3D) medical images is crucial for diagnostic and therapeutic practices in healthcare. Recent years have s...

Fundus2Globe: Generative AI-Driven 3D Digital Twins for Personalized Myopia Management

Myopia, projected to affect 50% population globally by 2050, is a leading cause of vision loss. Eyes with pathological myopia exhibit distinctive sh...

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

The emergence of large Vision Language Models (VLMs) has broadened the scope and capabilities of single-modal Large Language Models (LLMs) by integr...

MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Consequently, evaluations of large vision-language mo...

DAMamba: Vision State Space Model with Dynamic Adaptive Scan

State space models (SSMs) have recently garnered significant attention in computer vision. However, due to the unique characteristics of image data,...

Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining

Generalization remains a significant challenge for low-level vision models, which often struggle with unseen degradations in real-world scenarios de...

Computational Safety for Generative AI: A Signal Processing Perspective

AI safety is a rapidly growing area of research that seeks to prevent the harm and misuse of frontier AI technology, particularly with respect to ge...

OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging

Optical Coherence Tomography (OCT) provides high-resolution cross-sectional images useful for diagnosing various diseases, but their distinct charac...

LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models

Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders the...

VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues

Visually linking matching cues is a crucial ability in daily life, such as identifying the same person in multiple photos based on their cues, even ...

Range and Bird's Eye View Fused Cross-Modal Visual Place Recognition

Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are ...

Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models

Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language model...

A Physics-Informed Blur Learning Framework for Imaging Systems

Accurate blur estimation is essential for high-performance imaging across various applications. Blur is typically represented by the point spread fu...

VLDBench: Vision Language Models Disinformation Detection Benchmark

The rapid rise of AI-generated content has made detecting disinformation increasingly challenging. In particular, multimodal disinformation, i.e., o...

Multi-Faceted Multimodal Monosemanticity

Humans experience the world through multiple modalities, such as, vision, language, and speech, making it natural to explore the commonality and dis...

Browse Categories