Latest AI and machine learning research in ophthalmology for healthcare professionals.
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leadin...
This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic or patterned contact lenses has been extensively studied under the presentation attack detection (PAD) paradigm, clear prescription contact lenses, that are typically transparent, have received comparatively less attention despite their widespread use. Unlike pa...
Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-worl...
Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR...
Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize ...
Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures ar...
Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and...
Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assis...
Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language ...
Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation...
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, t...
Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial...
Myopia-induced posterior-pole remodeling is frequently accompanied by Optic Disc (OD) deformation and Peripapillary Atrophy (PPA), both of which provi...
Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific b...
The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Mode...
Purpose: To evaluate whether fluorescence lifetime imaging ophthalmoscopy (FLIO) combined with deep learning can detect metabolic signatures for class...
The peripheral nervous system innervates the pancreatic ductal adenocarcinoma (PDAC) microenvironment, and perineural invasion (PNI), the invasion of ...
Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertai...
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approach...
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural...