Latest AI and machine learning research in ophthalmology for healthcare professionals.
A key trend in Large Reasoning Models (e.g., OpenAI's o3) is the native agentic ability to use external tools such as web browsers for searching and writing/executing code for image manipulation to think with images. In the open-source research community, while significant progress has been made in language-only agentic abilities such as function calling and tool integration, the development of ...
We propose Visual-only Question Answering (VoQA), a novel multimodal task in which questions are visually embedded within images, without any accompanying textual input. This requires models to locate, recognize, and reason over visually embedded textual questions, posing challenges for existing large vision-language models (LVLMs), which show notable performance drops even with carefully design...
Vision Mamba has recently emerged as a promising alternative to Transformer-based architectures, offering linear complexity in sequence length while...
Early prevention and standardized management of refractory wounds in the elderly are very important for improving prognosis, reducing disability rate,...
Driver drowsiness is a leading cause of road accidents, resulting in significant societal, economic, and emotional losses. This paper introduces a nov...
PURPOSE OF REVIEW: Chronic pain significantly impacts quality of life for millions globally, with spinal cord stimulation (SCS) as an established trea...
Large language models are increasingly integrated into news recommendation systems, raising concerns about their role in spreading misinformation. I...
With the rapid improvement of machine learning (ML) models, cognitive scientists are increasingly asking about their alignment with how humans think...
Inspired by foveal vision, hard attention models promise interpretability and parameter economy. However, existing models like the Recurrent Model o...
The back-focal plane (BFP) of a high-numerical aperture objective contains the fluoro-phore radiation pattern, which encodes information about the a...
Integrating Large Language Models with symbolic planners is a promising direction for obtaining verifiable and grounded plans compared to planning i...
3D visual illusion is a perceptual phenomenon where a two-dimensional plane is manipulated to simulate three-dimensional spatial relationships, maki...
The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models fo...
Recent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency...
Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-com...
Recently, significant advances have been made in 3D object generation. Building upon the generated geometry, current pipelines typically employ imag...
As optical systems become increasingly complex, accurate and fast alignment is becoming more critical. Active alignment (AA) techniques dynamically op...
This study aims to develop and evaluate a convolutional neural network (CNN)-based architecture for detecting eye blink episodes in electroencephalogr...
We present \textit{kornia-rs}, a high-performance 3D computer vision library written entirely in native Rust, designed for safety-critical and real-...
Despite achieving significant progress in 2D image understanding, large multimodal models (LMMs) struggle in the physical world due to the lack of s...