Latest AI and machine learning research in ophthalmology for healthcare professionals.
Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled datasets are rarely available. We propose answer-conditioned chain-of-thought (CoT) distillation for rapidly adapting small vision-language models (VLMs) to new industrial tasks using minimal labeled data. A frontier VLM receives each training image along with...
Medical image classifiers are often trained within one source population, yet clinical deployment requires robustness to patients whose appearance, acquisition style, and disease prevalence differ from the source cohort. Existing fairness and robustness methods often require group supervision or treat appearance variation as an undifferentiated nuisance, which is insufficient when population-corre...
Eye movements provide valuable insights into human cognition and are a critical variable in numerous functional magnetic resonance imaging (fMRI) stud...
A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single ...
We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transforme...
While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solv...
Medical images require comprehensive and accurate interpretation to support the diagnosis of diverse clincial conditions. Recent vision-language gener...
Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, s...
High-quality visual representation is a long-standing pursuit in computer vision. In the context of multimodal LLMs (MLLMs), feeding higher-resolution...
Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood. Exist...
Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial sca...
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-C...
Cancer genomics and diagnostics is a rapidly evolving field in which identifying which topics attract early citation prominence can inform laboratory ...
In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment th...
Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, t...
Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories. Although recent vision-language and...
Accurate and scalable assessment of quantitative neuroimaging biomarkers, such as white matter hyperintensities (WMH) and hippocampal (HIP) volumes, i...
Despite rapid advances in chest X-ray (CXR) foundation models, most radiology report generation (RRG) systems still rely on heavily downsampled inputs...
Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are ...
Squint and cataract are major ocular disorders that majorly affect visual perception and interaction capability. This paper proposes a real-time video...