Latest AI and machine learning research in ophthalmology for healthcare professionals.
Recently, non-convolutional models such as the Vision Transformer (ViT) and Vision Mamba (Vim) have achieved remarkable performance in computer vision tasks. However, their reliance on fixed-size patches often results in excessive encoding of background regions and omission of critical local details, especially when informative objects are sparsely distributed. To address this, we introduce a fu...
The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpret. Establishing a robust theoretical foundation to explain its remarkable success and limitations has therefore become an increasingly prominent focus in recent research. Some notable directions have explored understan...
Urban cultures and architectural styles vary significantly across cities due to geographical, chronological, historical, and socio-political factors...
Students' academic emotions significantly influence their social behavior and learning performance. Traditional approaches to automatically and accu...
Hallucinations pose a significant challenge to the reliability of large vision-language models, making their detection essential for ensuring accura...
Visual servoing technology has been well developed and applied in many automated manufacturing tasks, especially in tools' pose alignment. To access...
As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reaso...
One promise that Vision-Language-Action (VLA) models hold over traditional imitation learning for robotics is to leverage the broad generalization c...
Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in the...
Absolute localization, aiming to determine an agent's location with respect to a global reference, is crucial for unmanned aerial vehicles (UAVs) in...
Despite the rapid progress of multimodal large language models (MLLMs), they have largely overlooked the importance of visual processing. In a simpl...
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textua...
Automated 3D CT diagnosis empowers clinicians to make timely, evidence-based decisions by enhancing diagnostic accuracy and workflow efficiency. Whi...
Large Vision-Language Models (LVLMs) have achieved impressive progress across various applications but remain vulnerable to malicious queries that e...
Typical large vision-language models (LVLMs) apply autoregressive supervision solely to textual sequences, without fully incorporating the visual mo...
Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retriev...
Artificial intelligence (AI) has become a fundamental tool for assisting clinicians in analyzing ophthalmic images, such as optical coherence tomogr...
The rapid advancement of transformer-based language models has catalyzed breakthroughs in biomedical and clinical natural language processing; howev...
To meet the growing demand for systematic surgical training, wetlab environments have become indispensable platforms for hands-on practice in ophtha...
The parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research. While CLIP is ...