Latest AI and machine learning research in ophthalmology for healthcare professionals.
Generating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is \textbf{not suitable for evaluating the distortion}. In this work, we first propose a distortion-specific CL...
Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing approaches primarily rely on parameter-efficient fine-tuning of image-text pre-trained models, yet they often suffer from limited interpretability and poor generalization due to inadequate temporal modeling. To address these, we propose a simple yet effective vid...
Deep learning has emerged as the predominant solution for classifying medical images. We intend to apply these developments to the ultra-widefield (...
Retrieval-augmented generation (RAG) has emerged as a pivotal technique in artificial intelligence (AI), particularly in enhancing the capabilities ...
Image compression technology eliminates redundant information to enable efficient transmission and storage of images, serving both machine vision an...
Recent advancements in ophthalmology foundation models such as RetFound have demonstrated remarkable diagnostic capabilities but require massive dat...
Parameter-efficient fine-tuning (PEFT) has emerged as a crucial approach for adapting large vision transformers to downstream tasks without the proh...
Well-being is a dynamic construct that evolves over time and fluctuates within individuals, presenting challenges for accurate quantification. Reduc...
Recent advancements demonstrated by DeepSeek-R1 have shown that complex reasoning abilities in large language models (LLMs), including sophisticated...
News outlets' competition for attention in news interfaces has highlighted the need for demographically-aware saliency prediction models. Despite re...
Vision-Language Models (VLMs) learn a shared feature space for text and images, enabling the comparison of inputs of different modalities. While pri...
Deep learning has become an essential part of computer vision, with deep neural networks (DNNs) excelling in predictive performance. However, they o...
In the computer vision community, the preference for pre-training visual models has largely shifted toward sRGB images due to their ease of acquisit...
Deep neural networks (DNNs) has shown great promise in computer vision tasks. However, machine vision achieved by DNNs cannot be as robust as human ...
Vision Language Models exhibited immense potential for embodied AI, yet they often lack the sophisticated situational reasoning required for complex...
Vision encoders typically generate a large number of visual tokens, providing information-rich representations but significantly increasing computat...
Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we s...
The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In part...
Bokeh rendering methods play a key role in creating the visually appealing, softly blurred backgrounds seen in professional photography. While recen...
Vision-language models for Earth observation (EO) typically rely on the visual spectrum of data as the only model input, thus failing to leverage th...