Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5581-5600 of 9,853 articles

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLAs primarily focus on direct input--output mappings, lacking the intermediate reasoning steps cruci...

JEEM: Vision-Language Understanding in Four Arabic Dialects

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. JEEM includes the tasks of image captioning and visual question answering, and features culturally rich and regionally diverse content. This dataset aims to assess the ability of VLMs to generalize across dialec...

StarFlow: Generating Structured Workflow Outputs From Sketch Images

Workflows are a fundamental component of automation in enterprise platforms, enabling the orchestration of tasks, data processing, and system integr...

Test-Time Visual In-Context Tuning

Visual in-context learning (VICL), as a new paradigm in computer vision, allows the model to rapidly adapt to various tasks with only a handful of p...

Fwd2Bot: LVLM Visual Token Compression with Double Forward Bottleneck

In this work, we aim to compress the vision tokens of a Large Vision Language Model (LVLM) into a representation that is simultaneously suitable for...

CTRL-O: Language-Controllable Object-Centric Visual Representation Learning

Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot cap...

AlignDiff: Learning Physically-Grounded Camera Alignment via Diffusion

Accurate camera calibration is a fundamental task for 3D perception, especially when dealing with real-world, in-the-wild environments where complex...

Retinal Fundus Multi-Disease Image Classification using Hybrid CNN-Transformer-Ensemble Architectures

Our research is motivated by the urgent global issue of a large population affected by retinal diseases, which are evenly distributed but underserve...

InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language m...

Vision-to-Music Generation: A Survey

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonst...

Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping

Transformer-based models have driven significant advancements in Multimodal Large Language Models (MLLMs), yet their computational costs surge drast...

VinaBench: Benchmark for Faithful and Consistent Visual Narratives

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual ...

Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). R...

Vision as LoRA

We introduce Vision as LoRA (VoRA), a novel paradigm for transforming an LLM into an MLLM. Unlike prevalent MLLM architectures that rely on external...

Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model and Training Strategy

The rapid development of multimodal large language models has resulted in remarkable advancements in visual perception and understanding, consolidat...

AutoRad-Lung: A Radiomic-Guided Prompting Autoregressive Vision-Language Model for Lung Nodule Malignancy Prediction

Lung cancer remains one of the leading causes of cancer-related mortality worldwide. A crucial challenge for early diagnosis is differentiating unce...

Pupillary reactions depend on disgust sensitivity in conceptual pavlovian disgust conditioning

Exposure-based interventions rely on inhibitory learning, often studied through Pavlovian conditioning. While disgust conditioning is increasingly l...

Vision-Amplified Semantic Entropy for Hallucination Detection in Medical Visual Question Answering

Multimodal large language models (MLLMs) have demonstrated significant potential in medical Visual Question Answering (VQA). Yet, they remain prone ...

Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards

Can Visual Language Models (VLMs) effectively capture human visual preferences? This work addresses this question by training VLMs to think about pr...

Scaling Vision Pre-Training to 4K Resolution

High-resolution perception of visual details is crucial for daily tasks. Current vision pre-training, however, is still limited to low resolutions (...

Browse Categories