Latest AI and machine learning research in ophthalmology for healthcare professionals.
Large-scale video generative models are trained on vast and diverse visual data, enabling them to internalize rich structural, semantic, and dynamic priors of the visual world. While these models have demonstrated impressive generative capability, their potential as general-purpose visual learners remains largely untapped. In this work, we introduce V-Bridge, a framework that bridges this latent c...
Vision-to-code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured representations with high visual fidelity. While recent Large Vision Language Models (LVLMs) achieve strong results via supervised fine-tuning, reinforcement learning remains challenging due to misaligned reward signals. Existing rewards either rely on textua...
Accurate classification of brain tumors from MRI is critical for guiding clinical decision-making; however, existing deep learning models are often hi...
Generalizing image classification across domains remains challenging in critical tasks such as fundus image-based diabetic retinopathy (DR) grading an...
Practical webcam gaze tracking is constrained not only by error, but also by calibration burden, robustness to head motion and session drift, runtime ...
Self-supervised and multimodal vision encoders learn strong visual representations that are widely adopted in downstream vision tasks and large vision...
Prevalent Computational Aberration Correction (CAC) methods are typically tailored to specific optical systems, leading to poor generalization and lab...
Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. Howev...
Breast cancer is one of the most common causes of death among women worldwide, with millions of fatalities annually. Magnetic Resonance Imaging (MRI) ...
Current training-free methods tackle MLLM hallucination with separate strategies: either enhancing visual signals or suppressing text inertia. However...
While large vision-language models (LVLMs) achieve strong performance on multimodal tasks, they frequently generate hallucinations -- unfaithful outpu...
Generative models are widely employed to enhance the photorealism of synthetic data for training computer vision algorithms. However, they often intro...
Vision-language pretraining has driven significant progress in medical image analysis. However, current methods typically supervise visual encoders us...
Vision-language models (VLMs) face significant computational inefficiencies caused by excessive generation of visual tokens. While prior work shows th...
Human vision achieves remarkable perceptual performance while operating under strict metabolic constraints. A key ingredient is the selective attentio...
Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA ...
Purpose: This study aimed to compare the reliability of myopia-related information from AI chatbots using a set of commonly asked questions by parents...
Medical vision-language pretraining (VLP) models have recently been investigated for their generalization to diverse downstream tasks. However, curren...
Vision-language pretraining has driven significant progress in medical image analysis. However, current methods typically supervise visual encoders us...
Rotation equivariance constitutes one of the most general and crucial structural priors for visual data, yet it remains notably absent from current Ma...