Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6061-6080 of 9,853 articles

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering

Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about images. Since its inception in 2015, VQA has rapidly evolved, driven by advances in deep learning, attention mechanisms, and transformer-based models. This survey traces the jou...

UNetVL: Enhancing 3D Medical Image Segmentation with Chebyshev KAN Powered Vision-LSTM

3D medical image segmentation has progressed considerably due to Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), yet these methods struggle to balance long-range dependency acquisition with computational efficiency. To address this challenge, we propose UNETVL (U-Net Vision-LSTM), a novel architecture that leverages recent advancements in temporal information processing. UNE...

LEO: Boosting Mixture of Vision Encoders for Multimodal Large Language Models

Enhanced visual understanding serves as a cornerstone for multimodal large language models (MLLMs). Recent hybrid MLLMs incorporate a mixture of vis...

Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing

This paper introduces Shake-VLA, a Vision-Language-Action (VLA) model-based system designed to enable bimanual robotic manipulation for automated co...

TWIX: Automatically Reconstructing Structured Data from Templatized Documents

Many documents, that we call templatized documents, are programmatically generated by populating fields in a visual template. Effective data extract...

CeViT: Copula-Enhanced Vision Transformer in multi-task learning and bi-group image covariates with an application to myopia screening

We aim to assist image-based myopia screening by resolving two longstanding problems, "how to integrate the information of ocular images of a pair o...

On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective

Graph Neural Networks (GNNs) have become the standard approach for learning and reasoning over relational data, leveraging the message-passing mecha...

Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs

Current multimodal large language models (MLLMs) often underperform on mathematical problem-solving tasks that require fine-grained visual understan...

Emergent Symbol-like Number Variables in Artificial Neural Networks

What types of numeric representations emerge in neural systems? What would a satisfying answer to this question look like? In this work, we interpre...

Generate, Transduct, Adapt: Iterative Transduction with VLMs

Transductive zero-shot learning with vision-language models leverages image-image similarities within the dataset to achieve better classification a...

An Attention-Guided Deep Learning Approach for Classifying 39 Skin Lesion Types

The skin, as the largest organ of the human body, is vulnerable to a diverse array of conditions collectively known as skin lesions, which encompass...

Language-Inspired Relation Transfer for Few-shot Class-Incremental Learning

Depicting novel classes with language descriptions by observing few-shot samples is inherent in human-learning systems. This lifelong learning capab...

AI-Driven Diabetic Retinopathy Screening: Multicentric Validation of AIDRSS in India

Purpose: Diabetic retinopathy (DR) is a major cause of vision loss, particularly in India, where access to retina specialists is limited in rural ar...

Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning

Infants develop complex visual understanding rapidly, even preceding the acquisition of linguistic skills. As computer vision seeks to replicate the...

V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer

Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a l...

Feedback-Driven Vision-Language Alignment with Minimal Human Supervision

Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is ofte...

On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis

Recently, Visual Autoregressive ($\mathsf{VAR}$) Models introduced a groundbreaking advancement in the field of image generation, offering a scalabl...

An Efficient Adaptive Compression Method for Human Perception and Machine Vision Tasks

While most existing neural image compression (NIC) and neural video compression (NVC) methodologies have achieved remarkable success, their optimiza...

Topology-based deep-learning segmentation method for deep anterior lamellar keratoplasty (DALK) surgical guidance using M-mode OCT data

Deep Anterior Lamellar Keratoplasty (DALK) is a partial-thickness corneal transplant procedure used to treat corneal stromal diseases. A crucial ste...

Deep Learning for Ophthalmology: The State-of-the-Art and Future Trends

The emergence of artificial intelligence (AI), particularly deep learning (DL), has marked a new era in the realm of ophthalmology, offering transfo...

Browse Categories