Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6401-6420 of 9,853 articles

Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses

Vision-Language Models (VLMs) implicitly learn to associate image regions with words from large-scale training data, demonstrating an emergent capability for grounding concepts without dense annotations[14,18,51]. However, the coarse-grained supervision from image-caption pairs is often insufficient to resolve ambiguities in object-concept correspondence, even with enormous data volume. Rich sem...

Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation

Large Vision-Language Models (VLMs) have demonstrated remarkable performance across multimodal tasks by integrating vision encoders with large language models (LLMs). However, these models remain vulnerable to adversarial attacks. Among such attacks, Universal Adversarial Perturbations (UAPs) are especially powerful, as a single optimized perturbation can mislead the model across various input i...

Machine learning-driven conservative-to-primitive conversion in hybrid piecewise polytropic and tabulated equations of state

We present a novel machine learning (ML) method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabul...

Learning Visual Generative Priors without Text

Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling...

DRUM: Learning Demonstration Retriever for Large MUlti-modal Models

Recently, large language models (LLMs) have demonstrated impressive capabilities in dealing with new tasks with the help of in-context learning (ICL...

Virtual Reflections on a Dynamic 2D Eye Model Improve Spatial Reference Identification

The visible orientation of human eyes creates some transparency about people's spatial attention and other mental states. This leads to a dual role ...

Primary visual cortex contributes to color constancy by predicting rather than discounting the illuminant: evidence from a computational study

Color constancy (CC) is an important ability of the human visual system to stably perceive the colors of objects despite considerable changes in the...

Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models

While large vision-language models (LVLMs) have shown impressive capabilities in generating plausible responses correlated with input visual content...

Visual Lexicon: Rich Image Features in Language Space

We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intr...

Convolution goes higher-order: a biologically inspired mechanism empowers image classification

We propose a novel approach to image classification inspired by complex nonlinear biological visual processing, whereby classical convolutional neur...

You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

Recent 3D generation models typically rely on limited-scale 3D `gold-labels' or 2D diffusion priors for 3D content creation. However, their performa...

The Narrow Gate: Localized Image-Text Communication in Vision-Language Models

Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. Thi...

MAVias: Mitigate any Visual Bias

Mitigating biases in computer vision models is an essential step towards the trustworthiness of artificial intelligence models. Existing bias mitiga...

Fundus Image-based Visual Acuity Assessment with PAC-Guarantees

Timely detection and treatment are essential for maintaining eye health. Visual acuity (VA), which measures the clarity of vision at a distance, is ...

Inverting Transformer-based Vision Models

Understanding the mechanisms underlying deep neural networks in computer vision remains a fundamental challenge. While many previous approaches have...

MSCrackMamba: Leveraging Vision Mamba for Crack Detection in Fused Multispectral Imagery

Crack detection is a critical task in structural health monitoring, aimed at assessing the structural integrity of bridges, buildings, and roads to ...

Evaluating Model Perception of Color Illusions in Photorealistic Scenes

We study the perception of color illusions by vision-language models. Color illusion, where a person's visual system perceives color differently fro...

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) enc...

HSDA: High-frequency Shuffle Data Augmentation for Bird's-Eye-View Map Segmentation

Autonomous driving has garnered significant attention in recent research, and Bird's-Eye-View (BEV) map segmentation plays a vital role in the field...

Language Model as Visual Explainer

In this paper, we present Language Model as Visual Explainer LVX, a systematic approach for interpreting the internal workings of vision models usin...

Browse Categories