Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5001-5020 of 9,853 articles

NSD-Imagery: A benchmark dataset for extending fMRI vision decoding methods to mental imagery

We release NSD-Imagery, a benchmark dataset of human fMRI activity paired with mental images, to complement the existing Natural Scenes Dataset (NSD), a large-scale dataset of fMRI activity paired with seen images that enabled unprecedented improvements in fMRI-to-image reconstruction efforts. Recent models trained on NSD have been evaluated only on seen image reconstruction. Using NSD-Imagery, ...

Hybrid Vision Transformer-Mamba Framework for Autism Diagnosis via Eye-Tracking Analysis

Accurate Autism Spectrum Disorder (ASD) diagnosis is vital for early intervention. This study presents a hybrid deep learning framework combining Vision Transformers (ViT) and Vision Mamba to detect ASD using eye-tracking data. The model uses attention-based fusion to integrate visual, speech, and facial cues, capturing both spatial and temporal dynamics. Unlike traditional handcrafted methods, ...

Multimodal Spatial Language Maps for Robot Navigation and Manipulation

Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event...

Harnessing Vision-Language Models for Time Series Anomaly Detection

Time-series anomaly detection (TSAD) has played a vital role in a variety of fields, including healthcare, finance, and industrial monitoring. Prior...

Vision-QRWKV: Exploring Quantum-Enhanced RWKV Models for Image Classification

Recent advancements in quantum machine learning have shown promise in enhancing classical neural network architectures, particularly in domains invo...

PhysLab: A Benchmark Dataset for Multi-Granularity Visual Parsing of Physics Experiments

Visual parsing of images and videos is critical for a wide range of real-world applications. However, progress in this field is constrained by limit...

CoMemo: LVLMs Need Image Context with Image Memory

Recent advancements in Large Vision-Language Models built upon Large Language Models have established aligning visual features with LLM representati...

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study

While traditional computer vision models have historically struggled to generalize to endoscopic domains, the emergence of foundation models has sho...

HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion

Reconstructing visual information from brain activity bridges the gap between neuroscience and computer vision. Even though progress has been made i...

Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?

Humans are susceptible to optical illusions, which serve as valuable tools for investigating sensory and cognitive processes. Inspired by human visi...

MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Recent advances in vision-language models have enabled rich semantic understanding across modalities. However, these encoding methods lack the abili...

DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models

Vision-Language-Action (VLA) models have advanced autonomous driving, but existing benchmarks still lack scenario diversity, reliable action-level a...

Degradation-Aware Image Enhancement via Vision-Language Classification

Image degradation is a prevalent issue in various real-world applications, affecting visual quality and downstream processing tasks. In this study, ...

BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models

Visual Language Models (VLMs) are now sufficiently advanced to support a broad range of applications, including answering complex visual questions, ...

Coordinated Robustness Evaluation Framework for Vision-Language Models

Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in ...

Refer to Anything with Vision-Language Prompts

Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehens...

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

Chain-of-Thought (CoT) has widely enhanced mathematical reasoning in Large Language Models (LLMs), but it still remains challenging for extending it...

Stable Vision Concept Transformers for Medical Diagnosis

Transparency is a paramount concern in the medical field, prompting researchers to delve into the realm of explainable AI (XAI). Among these XAI met...

Norming Sets for Tensor and Polynomial Sketching

This paper develops the sketching (i.e., randomized dimension reduction) theory for real algebraic varieties and images of polynomial maps, includin...

Robustness as Architecture: Designing IQA Models to Withstand Adversarial Perturbations

Image Quality Assessment (IQA) models are increasingly relied upon to evaluate image quality in real-world systems -- from compression and enhanceme...

Browse Categories