Latest AI and machine learning research in ophthalmology for healthcare professionals.
We release NSD-Imagery, a benchmark dataset of human fMRI activity paired with mental images, to complement the existing Natural Scenes Dataset (NSD), a large-scale dataset of fMRI activity paired with seen images that enabled unprecedented improvements in fMRI-to-image reconstruction efforts. Recent models trained on NSD have been evaluated only on seen image reconstruction. Using NSD-Imagery, ...
Accurate Autism Spectrum Disorder (ASD) diagnosis is vital for early intervention. This study presents a hybrid deep learning framework combining Vision Transformers (ViT) and Vision Mamba to detect ASD using eye-tracking data. The model uses attention-based fusion to integrate visual, speech, and facial cues, capturing both spatial and temporal dynamics. Unlike traditional handcrafted methods, ...
Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event...
Time-series anomaly detection (TSAD) has played a vital role in a variety of fields, including healthcare, finance, and industrial monitoring. Prior...
Recent advancements in quantum machine learning have shown promise in enhancing classical neural network architectures, particularly in domains invo...
Visual parsing of images and videos is critical for a wide range of real-world applications. However, progress in this field is constrained by limit...
Recent advancements in Large Vision-Language Models built upon Large Language Models have established aligning visual features with LLM representati...
While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has sho...
Reconstructing visual information from brain activity bridges the gap between neuroscience and computer vision. Even though progress has been made i...
Humans are susceptible to optical illusions, which serve as valuable tools for investigating sensory and cognitive processes. Inspired by human visi...
Recent advances in vision-language models have enabled rich semantic understanding across modalities. However, these encoding methods lack the abili...
Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level a...
Image degradation is a prevalent issue in various real-world applications, affecting visual quality and downstream processing tasks. In this study, ...
Visual Language Models (VLMs) are now sufficiently advanced to support a broad range of applications, including answering complex visual questions, ...
Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in ...
Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehens...
Chain-of-Thought (CoT) has widely enhanced mathematical reasoning in Large Language Models (LLMs), but it still remains challenging for extending it...
Transparency is a paramount concern in the medical field, prompting researchers to delve into the realm of explainable AI (XAI). Among these XAI met...
This paper develops the sketching (i.e., randomized dimension reduction) theory for real algebraic varieties and images of polynomial maps, includin...
Image Quality Assessment (IQA) models are increasingly relied upon to evaluate image quality in real-world systems -- from compression and enhanceme...