Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5801-5820 of 9,853 articles

VLEER: Vision and Language Embeddings for Explainable Whole Slide Image Representation

Recent advances in vision-language models (VLMs) have shown remarkable potential in bridging visual and textual modalities. In computational pathology, domain-specific VLMs, which are pre-trained on extensive histopathology image-text datasets, have succeeded in various downstream tasks. However, existing research has primarily focused on the pre-training process and direct applications of VLMs ...

Visual Attention Exploration in Vision-Based Mamba Models

State space models (SSMs) have emerged as an efficient alternative to transformer-based models, offering linear complexity that scales better than transformers. One of the latest advances in SSMs, Mamba, introduces a selective scan mechanism that assigns trainable weights to input tokens, effectively mimicking the attention mechanism. Mamba has also been successfully extended to the vision domai...

Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow

Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffe...

Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models

Omnimodal Large Language Models (OLLMs) have shown significant progress in integrating vision and text, but still struggle with integrating vision a...

Towards Statistical Factuality Guarantee for Large Vision-Language Models

Advancements in Large Vision-Language Models (LVLMs) have demonstrated promising performance in a variety of vision-language tasks involving image-c...

LIFT-GS: Cross-Scene Render-Supervised Distillation for 3D Language Grounding

Our approach to training 3D vision-language understanding models is to train a feedforward model that makes predictions in 3D, but never requires 3D...

Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation

Autoregressive (AR) modeling, known for its next-token prediction paradigm, underpins state-of-the-art language and visual generative models. Tradit...

Visual Adaptive Prompting for Compositional Zero-Shot Learning

Vision-Language Models (VLMs) have demonstrated impressive capabilities in learning joint representations of visual and textual data, making them po...

Explainable, Multi-modal Wound Infection Classification from Images Augmented with Generated Captions

Infections in Diabetic Foot Ulcers (DFUs) can cause severe complications, including tissue death and limb amputation, highlighting the need for accu...

Vector-Quantized Vision Foundation Models for Object-Centric Learning

Perceiving visual scenes as objects and background -- like humans do -- Object-Centric Learning (OCL) aggregates image or video feature maps into ob...

Do computer vision foundation models learn the low-level characteristics of the human visual system?

Computer vision foundation models, such as DINO or OpenCLIP, are trained in a self-supervised manner on large image datasets. Analogously, substanti...

RURANET++: An Unsupervised Learning Method for Diabetic Macular Edema Based on SCSE Attention Mechanisms and Dynamic Multi-Projection Head Clustering

Diabetic Macular Edema (DME), a prevalent complication among diabetic patients, constitutes a major cause of visual impairment and blindness. Althou...

SegLocNet: Multimodal Localization Network for Autonomous Driving via Bird's-Eye-View Segmentation

Robust and accurate localization is critical for autonomous driving. Traditional GNSS-based localization methods suffer from signal occlusion and mu...

A high-performance and portable implementation of the SISSO method for CPUs and GPUs

SISSO (sure-independence screening and sparsifying operator) is an artificial intelligence (AI) method based on symbolic regression and compressed s...

Night-Voyager: Consistent and Efficient Nocturnal Vision-Aided State Estimation in Object Maps

Accurate and robust state estimation at nighttime is essential for autonomous robotic navigation to achieve nocturnal or round-the-clock tasks. An i...

Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents

To improve Multimodal Large Language Models' (MLLMs) ability to process images and complex instructions, researchers predominantly curate large-scal...

InPK: Infusing Prior Knowledge into Prompt for Vision-Language Models

Prompt tuning has become a popular strategy for adapting Vision-Language Models (VLMs) to zero/few-shot visual recognition tasks. Some prompting tec...

Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack

Multimodal Large Language Models (MLLMs), built upon LLMs, have recently gained attention for their capabilities in image recognition and understand...

Image-Based Roadmaps for Vision-Only Planning and Control of Robotic Manipulators

This work presents a motion planning framework for robotic manipulators that computes collision-free paths directly in image space. The generated pa...

GONet: A Generalizable Deep Learning Model for Glaucoma Detection

Glaucomatous optic neuropathy (GON) is a prevalent ocular disease that can lead to irreversible vision loss if not detected early and treated. The t...

Browse Categories