Latest AI and machine learning research in ophthalmology for healthcare professionals.
The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) res...
With the popularity of foundational models, parameter efficient fine tuning has become the defacto approach to leverage pretrained models to perform downstream tasks. Taking inspiration from recent advances in large language models, Visual Prompt Tuning, and similar techniques, learn an additional prompt to efficiently finetune a pretrained vision foundational model. However, we observe that suc...
Zero-shot anomaly detection (ZSAD) identifies anomalies without needing training samples from the target dataset, essential for scenarios with priva...
To perform image editing based on single-view, inverse physically based rendering, we present a method combining a learning-based approach with prog...
Rapid Serial Visual Presentation (RSVP)-based Brain-Computer Interfaces (BCIs) facilitate high-throughput target image detection by identifying even...
Retinal vascular morphology is crucial for diagnosing diseases such as diabetes, glaucoma, and hypertension, making accurate segmentation of retinal...
X-ray image based medical report generation achieves significant progress in recent years with the help of the large language model, however, these ...
Compared to human vision, locust visual systems excel at rapid and precise collision detection, despite relying on only hundreds of thousands of neu...
Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks co...
Retinal image registration is vital for diagnostic therapeutic applications within the field of ophthalmology. Existing public datasets, focusing on...
Visual-language models (VLM) have emerged as a powerful tool for learning a unified embedding space for vision and language. Inspired by large langu...
Large language models and vision transformers have demonstrated impressive zero-shot capabilities, enabling significant transferability in downstrea...
While Vision Language Models (VLMs) are impressive in tasks such as visual question answering (VQA) and image captioning, their ability to apply mul...
Instructional cataract surgery videos are crucial for ophthalmologists and trainees to observe surgical details repeatedly. This paper presents a de...
Vision generation remains a challenging frontier in artificial intelligence, requiring seamless integration of visual understanding and generative c...
Previous visual object tracking methods employ image-feature regression models or coordinate autoregression models for bounding box prediction. Imag...
We present an approach to modifying Transformer architectures by integrating graph-aware relational reasoning into the attention mechanism, merging ...
While mainstream vision-language models (VLMs) have advanced rapidly in understanding image level information, they still lack the ability to focus ...
Diabetic Retinopathy (DR) is a major cause of blindness worldwide, caused by damage to the blood vessels in the retina due to diabetes. Early detect...
Spatial scheduling of electrode activation ("rastering") is essential for safely operating high-density retinal implants, yet its perceptual consequ...