Latest AI and machine learning research in ophthalmology for healthcare professionals.
This study presents a novel human-machine interface (HMI) based on both electrooculography (EOG) and electroencephalography (EEG). This hybrid interface works in two modes: an EOG mode recognizes eye movements such as blinks, and an EEG mode detects event related potentials (ERPs) like P300. While both eye movements and ERPs have been separately used for implementing assistive interfaces, which he...
The current version of the Human Disease Ontology (DO) (http://www.disease-ontology.org) database expands the utility of the ontology for the examination and comparison of genetic variation, phenotype, protein, drug and epitope data through the lens of human disease. DO is a biomedical resource of standardized common and rare disease concepts with stable identifiers organized by disease etiology. ...
There is an increasing interest in learning mappings from features to saliency maps based on human fixation data on natural images. These models have ...
Ocular complications reported after robotic-assisted laparoscopic radical prostatectomy (RALP) include corneal abrasion and ischemic optic neuropathy....
Magnetic tubular implantable micro-robots are batch fabricated by electroforming. These microdevices can be used in targeted drug delivery and minimal...
OBJECTIVE: Prediction of epileptic seizures can improve the living conditions for refractory epilepsy patients. We aimed to improve sensitivity and sp...
Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess fin...
Vision-language models (VLMs) have achieved remarkable performance by leveraging complementary information from large-scale image-text pairs. However,...
General-purpose vision-language models (VLMs) now support strong visual recognition, instruction following, and generation. However, most pretrained v...
Text-Video Retrieval (TVR) retrieves videos that match a natural-language query, but extending image-text models such as CLIP to videos is fundamental...
Mineral image classification is important for geological exploration and resource development, but it remains challenging due to substantial intra-cla...
Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separati...
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pat...
Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inpu...
Vision-language models (VLMs) have demonstrated strong performance in visual question answering with natural images. However, they continue to struggl...
Despite their strong multimodal understanding ability, multimodal large language models (MLLMs) incur substantial computational overhead when processi...
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders strugg...
Image retouching is commonly formulated as enhancing overall visual quality through color adjustment, but in practice, it also serves to emphasize vis...
While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, functional unification does not guar...
Speculative decoding accelerates generation without changing its output, yet on vision-language models (VLMs) it has been caught in a self-defeating c...