Latest AI and machine learning research in ophthalmology for healthcare professionals.
Purpose: To develop and evaluate deep learning models for automated detection of corneal perforation in microbial keratitis using anterior segment optical coherence tomography (ASOCT) images. Methods: We enrolled 150 patients with microbiologically confirmed keratitis. Contralateral healthy eyes served as controls. Four convolutional neural network models using ResNet architecture were trained and...
As a classic vision task, anomaly detection has been widely applied in industrial inspection and medical imaging. In this task, data scarcity is often a frequently-faced issue. To solve it, the few-shot anomaly detection (FSAD) scheme is attracting increasing attention. In recent years, beyond traditional visual paradigm, Vision-Language Model (VLM) has been extensively explored to boost this fiel...
Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remain...
Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems...
Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in v...
Phishing websites now rely heavily on visual imitation-copied logos, similar layouts, and matching colours-to avoid detection by text- and URL-based s...
Ultra-high-resolution (UHR) remote sensing imagery couples kilometer-scale context with query-critical evidence that may occupy only a few pixels. Suc...
Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign condition...
Physical reasoning over visual inputs demands tight integration of visual perception, domain knowledge, and multi-step symbolic inference. Yet even st...
We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanis...
Automated diagnosis based on color fundus photography is essential for large-scale glaucoma screening. However, existing deep learning models are typi...
We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanis...
Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. D...
Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-...
Illumination using correlated photon sources has been established as an approach to allowing high-fidelity images to be reconstructed from noisy camer...
Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classific...
Scene change detection (SCD) is crucial for urban monitoring and navigation but remains challenging in real-world environments due to lighting variati...
While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and do...
Vision-Language Models (VLM) have revolutionized multimodal learning by jointly processing visual and textual information. Yet, they face significant ...
Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieva...