Latest AI and machine learning research in ophthalmology for healthcare professionals.
Abstract Purpose: Glaucoma, a leading cause of irreversible vision loss, often remains undiagnosed due to its asymptomatic progression and the limitations of existing screening methods. This study aimed to validate an artificial intelligence machine learning algorithm for the camera-agnostic detection of glaucomatous optic neuropathy using macula-centered fundus images. Methods: Data were collecte...
Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across downstream tasks, despite substantial computational investment. We postulate that this limitation arises from a mismatch between pretraining objectives and the deman...
Vision-Language Action (VLA) models have shown remarkable progress in robotic manipulation by leveraging the powerful perception abilities of Vision-L...
AI models for drug discovery and chemical literature mining must interpret molecular images and generate outputs consistent with 3D geometry and stere...
Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although CoT...
Vision-Language-Action (VLA) models have shown promise in robot manipulation but often struggle to generalize to new instructions or complex multi-tas...
Vision-language models (VLMs) lag behind text-only language models on mathematical reasoning when the same problems are presented as images rather tha...
Recent advancements in Large Vision-Language Models (LVLMs) have pushed them closer to becoming general-purpose assistants. Despite their strong perfo...
Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assess...
Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired, and if ...
Human-centric visual analysis plays a pivotal role in diverse applications, including surveillance, healthcare, and human-computer interaction. With t...
Current Vision-Language Models (VLMs) fail at quantitative spatial reasoning because their architectures destroy pixel-level information required for ...
Retinal fundus photography is indispensable for ophthalmic screening and diagnosis, yet image quality is often degraded by noise, artifacts, and uneve...
In this study, we propose a technique to improve the accuracy and reduce the size of convolutional neural networks (CNNs) running on edge devices for ...
In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model per...
Vision-as-inverse-graphics, the concept of reconstructing an image as an editable graphics program is a long-standing goal of computer vision. Yet eve...
Sociotechnical challenges of machine learning in healthcare and social welfare are mismatches between how a machine learning tool functions and the st...
Accurate segmentation of brain tumors is essential for clinical diagnosis and treatment planning. Deep learning is currently the state-of-the-art for ...
ImportanceVision-language models (VLMs) enable generalist multimodal reasoning, but their ability to resolve brief, low-contrast cues in surgical vide...
Vision-Language Pre-training (VLP) models demonstrate strong performance across various downstream tasks by learning from large-scale image-text pairs...