Latest AI and machine learning research in ophthalmology for healthcare professionals.
Diffusion models have recently emerged as powerful frameworks for generating high-quality images. While recent studies have explored their application to time series forecasting, these approaches face significant challenges in cross-modal modeling and transforming visual information effectively to capture temporal patterns. In this paper, we propose LDM4TS, a novel framework that leverages the p...
Single camera 3D perception for traffic monitoring faces significant challenges due to occlusion and limited field of view. Moreover, fusing information from multiple cameras at the image feature level is difficult because of different view angles. Further, the necessity for practical implementation and compatibility with existing traffic infrastructure compounds these challenges. To address the...
In this paper, we present a novel Left-Prompt-Guided (LPG) paradigm to address a diverse range of reference-based vision tasks. Inspired by the huma...
Large language models (LLMs) have significantly advanced the field of natural language generation. However, they frequently generate unverified outp...
Vision is a primary means of how humans perceive the environment, but Blind and Low-Vision (BLV) people need assistance understanding their surround...
Entity tracking is a fundamental challenge in natural language understanding, requiring models to maintain coherent representations of entities. Pre...
Lens flares arise from light reflection and refraction within sensor arrays, whose diverse types include glow, veiling glare, reflective flare and s...
Recent advancements in large language models (LLMs) have demonstrated extraordinary comprehension capabilities with remarkable breakthroughs on vari...
The Convolutional Neural Network (CNN) has shown impressive performance in image classification because of its strong learning capabilities. However...
This paper addresses the critical issue of burnout among cybersecurity professionals, a growing concern that threatens the effectiveness of digital ...
Vision-language models (VLMs) excel in various visual benchmarks but are often constrained by the lack of high-quality visual fine-tuning data. To a...
The Vision Transformer (ViT) has made significant strides in the field of computer vision. However, as the depth of the model and the resolution of ...
Variational Autoencoders (VAEs) and other generative models are widely employed in artificial intelligence to synthesize new data. However, current ...
We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, pa...
Multimodal conversational generative AI has shown impressive capabilities in various vision and language understanding through learning massive text...
We present HealthGPT, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabili...
Ocular surface and tear diseases are among the most common and significant ocular conditions affecting eye health. In recent years, research and clini...
With the continuous evolution of computer technology and the surging advent of the big data era, artificial intelligence (AI) has already manifested e...
This study reports the first steps toward establishing a computer vision system to help caregivers of bedridden patients detect pressure ulcers (PUs) ...
The continuous operation of Earth-orbiting satellites generates vast and ever-growing archives of Remote Sensing (RS) images. Natural language prese...