Latest AI and machine learning research in ophthalmology for healthcare professionals.
Vision-Language Models (VLMs) implicitly learn to associate image regions with words from large-scale training data, demonstrating an emergent capability for grounding concepts without dense annotations[14,18,51]. However, the coarse-grained supervision from image-caption pairs is often insufficient to resolve ambiguities in object-concept correspondence, even with enormous data volume. Rich sem...
Large Vision-Language Models (VLMs) have demonstrated remarkable performance across multimodal tasks by integrating vision encoders with large language models (LLMs). However, these models remain vulnerable to adversarial attacks. Among such attacks, Universal Adversarial Perturbations (UAPs) are especially powerful, as a single optimized perturbation can mislead the model across various input i...
We present a novel machine learning (ML) method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabul...
Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling...
Recently, large language models (LLMs) have demonstrated impressive capabilities in dealing with new tasks with the help of in-context learning (ICL...
The visible orientation of human eyes creates some transparency about people's spatial attention and other mental states. This leads to a dual role ...
Color constancy (CC) is an important ability of the human visual system to stably perceive the colors of objects despite considerable changes in the...
While large vision-language models (LVLMs) have shown impressive capabilities in generating plausible responses correlated with input visual content...
We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intr...
We propose a novel approach to image classification inspired by complex nonlinear biological visual processing, whereby classical convolutional neur...
Recent 3D generation models typically rely on limited-scale 3D `gold-labels' or 2D diffusion priors for 3D content creation. However, their performa...
Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. Thi...
Mitigating biases in computer vision models is an essential step towards the trustworthiness of artificial intelligence models. Existing bias mitiga...
Timely detection and treatment are essential for maintaining eye health. Visual acuity (VA), which measures the clarity of vision at a distance, is ...
Understanding the mechanisms underlying deep neural networks in computer vision remains a fundamental challenge. While many previous approaches have...
Crack detection is a critical task in structural health monitoring, aimed at assessing the structural integrity of bridges, buildings, and roads to ...
We study the perception of color illusions by vision-language models. Color illusion, where a person's visual system perceives color differently fro...
The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) enc...
Autonomous driving has garnered significant attention in recent research, and Bird's-Eye-View (BEV) map segmentation plays a vital role in the field...
In this paper, we present Language Model as Visual Explainer LVX, a systematic approach for interpreting the internal workings of vision models usin...