Latest AI and machine learning research in ophthalmology for healthcare professionals.
In this paper, we present a self-calibrating framework that jointly optimizes camera parameters, lens distortion and 3D Gaussian representations, enabling accurate and efficient scene reconstruction. In particular, our technique enables high-quality scene reconstruction from Large field-of-view (FOV) imagery taken with wide-angle lenses, allowing the scene to be modeled from a smaller number of ...
The retinal afterimage is a widely known effect in the human visual system, which has been studied and used in the context of a number of major art movements. Therefore, when considering the general role of computation in the visual arts, this begs the question whether this effect, too, may be induced using partly automated techniques. If so, it may become a computationally controllable ingredie...
This study investigates whether vision-language models (VLMs) can perform pragmatic inference, focusing on ignorance implicatures, utterances that i...
We propose general visual inspection model using Vision-Language Model~(VLM) with few-shot images of non-defective or defective products, along with...
Explainability is a highly demanded requirement for applications in high-risk areas such as medicine. Vision Transformers have mainly been limited t...
Text-to-image diffusion models have recently enabled the creation of visually compelling, detailed images from textual prompts. However, their abili...
Time series analysis has witnessed the inspiring development from traditional autoregressive models, deep learning models, to recent Transformers an...
Purpose: Foundation models, trained on multitudes of public datasets, often require additional fine-tuning or re-prompting mechanisms to be applied ...
Automated analysis of surgical videos is crucial for improving surgical training, workflow optimization, and postoperative assessment. We introduce ...
Vision-language navigation (VLN) has emerged as a promising paradigm, enabling mobile robots to perform zero-shot inference and execute tasks withou...
Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hiera...
As a powerful all-weather Earth observation tool, synthetic aperture radar (SAR) remote sensing enables critical military reconnaissance, maritime s...
Omnidirectional image (ODI) data is captured with a field-of-view of 360x180, which is much wider than the pinhole cameras and captures richer surro...
Monocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, whic...
Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual ques...
Vision-Language Models (VLMs), such as GPT-4V and Llama 3.2 vision, have garnered significant research attention for their ability to leverage Large...
Zero-Shot Anomaly Detection (ZSAD) is an emerging AD paradigm. Unlike the traditional unsupervised AD setting that requires a large number of normal...
Personal space, also known as peripersonal space, is crucial in human social interaction, influencing comfort, communication, and social stress. Est...
Whole slide pathology image classification presents challenges due to gigapixel image sizes and limited annotation labels, hindering model generaliz...
3D vision is of paramount importance for numerous applications ranging from machine intelligence to precision metrology. Despite much recent progres...