Latest AI and machine learning research in ophthalmology for healthcare professionals.
Replicating In-Context Learning (ICL) in computer vision remains challenging due to task heterogeneity. We propose \textbf{VIRAL}, a framework that elicits visual reasoning from a pre-trained image editing model by formulating ICL as conditional generation via visual analogy ($x_s : x_t :: x_q : y_q$). We adapt a frozen Diffusion Transformer (DiT) using role-aware multi-image conditioning and intr...
Optical coherence tomography (OCT) has revolutionized retinal disease diagnosis with its high-resolution and three-dimensional imaging nature, yet its full diagnostic automation in clinical practices remains constrained by multi-stage workflows and conventional single-slice single-task AI models. We present Full-process OCT-based Clinical Utility System (FOCUS), a foundation model-driven framework...
The proliferation of generative AI poses challenges for information integrity assurance, requiring systems that connect model governance with end-user...
Despite recent progress in vision-language models (VLMs), existing approaches often fail to generate personalized responses based on the user's specif...
Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a vid...
Retinal diseases spanning a broad spectrum can be effectively identified and diagnosed using complementary signals from multimodal data. However, mult...
Human vision is foveated, with variable resolution peaking at the center of a large field of view; this reflects an efficient trade-off for active sen...
Inverse scattering in optical coherence tomography (OCT) seeks to recover both structural images and intrinsic tissue optical properties, including re...
Flow matching learns a velocity field that transports a base distribution to data. We study how small latent perturbations propagate through these flo...
Deep learning systems in digital pathology are widely regarded as opaque, limiting clinical trust and interpretability. We present a framework for emp...
Visualization's design knowledge-effectiveness rankings, encoding guidelines, color models, preattentive processing rules -- derives from six decades ...
Logit Lens has been proposed for visualizing tokens that contribute most to LLM answers. Recently, Logit Lens was also shown to be applicable in autor...
Large vision-language models (VLMs) are vulnerable to transfer-based adversarial perturbations, enabling attackers to optimize on surrogate models and...
While deep learning models like Vision Transformer (ViT) have achieved significant advances, they typically require large datasets. With data privacy ...
Deploying ADAS and ADS across countries remains challenging due to differences in legislation, traffic infrastructure, and visual conventions, which i...
Large multimodal reasoning models solve challenging visual problems via explicit long-chain inference: they gather visual clues from images and decode...
Recent genomic foundation models largely adopt large language model architectures that treat DNA as a one-dimensional token sequence. However, exhaust...
Modern Vision-Language Models (VLMs) exhibit a critical flaw in compositional reasoning, often confusing "a red cube and a blue sphere" with "a blue c...
Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision-DeepResearch systems that use search engines for complex visual-text...
The integration of visual information into Large Language Models (LLMs) has enabled Multimodal LLMs (MLLMs), but the quadratic memory and computationa...