Latest AI and machine learning research in ophthalmology for healthcare professionals.
Recent advancements in multimodal reward models (RMs) have significantly propelled the development of visual generation. Existing frameworks typically adopt Bradley-Terry-style preference modeling or leverage generative VLMs as judges, and subsequently optimize visual generation models via reinforcement learning. However, current RMs suffer from inherent limitations: they often follow a one-size-f...
Vision-language models (VLMs) extend large language models (LLMs) with vision encoders, enabling text generation conditioned on both images and text. However, this multimodal integration expands the attack surface by exposing the model to image-based jailbreaks crafted to induce harmful responses. Existing gradient-based jailbreak methods transfer poorly, as adversarial patterns overfit to a singl...
Large Vision-Language Models (LVLMs) can reason effectively from image-text inputs and perform well in various multimodal tasks. Despite this success,...
Vision-Language-Action (VLA) models have emerged as a dominant paradigm for generalist robotic manipulation, unifying perception and control within a ...
This paper offers a mini review of Visual Word Sense Disambiguation (VWSD), which is a multimodal extension of traditional Word Sense Disambiguation (...
Developing 3D vision-language models with robust clinical reasoning remains a challenge due to the inherent complexity of volumetric medical imaging, ...
We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a set ...
The problem of corrupted data, missing features, or missing modalities continues to plague the modern machine learning landscape. To address this issu...
Endoscopic image analysis is vital for colorectal cancer screening, yet real-world conditions often suffer from lens fogging, motion blur, and specula...
Multimodal large language models (MLLMs) suffer from high computational costs due to excessive visual tokens, particularly in high-resolution and vide...
Vision-language models (VLMs) have become central to tasks such as visual question answering, image captioning, and text-to-image generation. However,...
Background: Achieving precise postoperative refractive outcomes remains a significant challenge in cataract surgery. While advanced intraocular lens (...
Spatial and activity-dependent gene regulation in the mammalian brain requires coordinated control of RNA synthesis and degradation, yet spatially res...
Spaceflight-associated neuro-ocular syndrome (SANS) threatens astronaut health during long-duration missions, yet its molecular pathology remains uncl...
Robust estimation of systemic human cognitive states is critical for many applications, from simply detecting inefficiencies in human task performance...
Vision foundation models trained on discretely sampled images achieve strong performance on classification benchmarks, yet whether their representatio...
The development of large vision language models drives the demand for managing, and applying massive amounts of multimodal data, making OCR technology...
Classifier-free guidance (CFG) is a widely used technique for controllable generation in diffusion and flow-based models. Despite its empirical succes...
Deep visual features are increasingly used as the interface in vision systems, motivating the need to describe feature characteristics and control fea...
Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity o...