Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5681-5700 of 9,853 articles

GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing

The application of Vision-Language Models (VLMs) in remote sensing (RS) has demonstrated significant potential in traditional tasks such as scene classification, object detection, and image captioning. However, current models, which excel in Referring Expression Comprehension (REC), struggle with tasks involving complex instructions (e.g., exists multiple conditions) or pixel-level operations li...

BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries

Encoder-free multimodal large language models(MLLMs) eliminate the need for a well-trained vision encoder by directly processing image tokens before the language model. While this approach reduces computational overhead and model complexity, it often requires large amounts of training data to effectively capture the visual knowledge typically encoded by vision models like CLIP. The absence of a ...

Empirical Privacy Variance

We propose the notion of empirical privacy variance and study it in the context of differentially private fine-tuning of language models. Specifical...

From Eye to Mind: brain2text Decoding Reveals the Neural Mechanisms of Visual Semantic Processing

Deciphering the neural mechanisms that transform sensory experiences into meaningful semantic representations is a fundamental challenge in cognitiv...

Minuscule Cell Detection in AS-OCT Images with Progressive Field-of-View Focusing

Anterior Segment Optical Coherence Tomography (AS-OCT) is an emerging imaging technique with great potential for diagnosing anterior uveitis, a visi...

Hyperbolic Safety-Aware Vision-Language Models

Addressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current eff...

Att-Adapter: A Robust and Precise Domain-Specific Multi-Attributes T2I Diffusion Adapter via Conditional Variational Autoencoder

Text-to-Image (T2I) Diffusion Models have achieved remarkable performance in generating high quality images. However, enabling precise control of co...

Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection

Vision-Language Models (VLMs) leverage aligned visual encoders to transform images into visual tokens, allowing them to be processed similarly to te...

Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense

Deploying large vision-language models (LVLMs) introduces a unique vulnerability: susceptibility to malicious attacks via visual inputs. However, ex...

Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models

Current Cross-Modality Generation Models (GMs) demonstrate remarkable capabilities in various generative tasks. Given the ubiquity and information r...

PARIC: Probabilistic Attention Regularization for Language Guided Image Classification from Pre-trained Vison Language Models

Language-guided attention frameworks have significantly enhanced both interpretability and performance in image classification; however, the relianc...

DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models

Visual Question Answering (VQA) models, which fall under the category of vision-language models, conventionally execute multiple downsampling proces...

Aerial Vision-and-Language Navigation with Grid-based View Selection and Map Construction

Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following huma...

Observation-Graph Interaction and Key-Detail Guidance for Vision and Language Navigation

Vision and Language Navigation (VLN) requires an agent to navigate through environments following natural language instructions. However, existing m...

Eye Movement Characteristics for Predicting a Transition to Psychosis: Longitudinal Changes and Implications.

BACKGROUND AND HYPOTHESIS: Substantive inquiry into the predictive power of eye movement (EM) features for clinical high-risk (CHR) conversion and the...

Mar 14 2025 38245498
The Power of One: A Single Example is All it Takes for Segmentation in VLMs

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities b...

HeightFormer: Learning Height Prediction in Voxel Features for Roadside Vision Centric 3D Object Detection via Transformer

Roadside vision centric 3D object detection has received increasing attention in recent years. It expands the perception range of autonomous vehicle...

Unifying 2D and 3D Vision-Language Understanding

Progress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture fo...

Learning Interpretable Logic Rules from Deep Vision Models

We propose a general framework called VisionLogic to extract interpretable logic rules from deep vision models, with a focus on image classification...

Learning Disease State from Noisy Ordinal Disease Progression Labels

Learning from noisy ordinal labels is a key challenge in medical imaging. In this work, we ask whether ordinal disease progression labels (better, w...

Browse Categories