Latest AI and machine learning research in ophthalmology for healthcare professionals.
According to the EPA, only 25% of waste is recycled, and just 60% of U.S. municipalities offer curbside recycling. Plastics fare worse, with a recycling rate of only 8%; an additional 16% is incinerated, while the remaining 76% ends up in landfills. The low plastic recycling rate stems from contamination, poor economic incentives, and technical difficulties, making efficient recycling a challeng...
Recently, vision transformers (ViTs) have achieved excellent performance on vision tasks by measuring the global self-attention among the image patches. Given $n$ patches, they will have quadratic complexity such as $\mathcal{O}(n^2)$ and the time cost is high when splitting the input image with a small granularity. Meanwhile, the pivotal information is often randomly gathered in a few regions o...
Despite their remarkable progress in multimodal understanding tasks, large vision language models (LVLMs) often suffer from "hallucinations", genera...
This paper presents a method for generating dynamic caustic patterns by utilising dual-optimised holographic fields with Phased Array Transducer (PA...
Eye gaze can provide rich information on human psychological activities, and has garnered significant attention in the field of Human-Robot Interact...
Visual autoregressive (VAR) modeling has marked a paradigm shift in image generation from next-token prediction to next-scale prediction. VAR predic...
OBJECTIVE: This study aims to develop ultrasound biomicroscopy (UBM)-based artificial intelligence (AI) models for preoperative differentiation of acu...
This article introduces a novel deep-learning based framework, Super-resolution/Denoising network (SDNet), for simultaneous denoising and super-resolu...
Multimodal Large Language Models (MLLMs) perform well on tasks such as visual question answering, but it remains unclear whether their reasoning rel...
Pretrained generative models have opened new frontiers in brain decoding by enabling the synthesis of realistic texts and images from non-invasive b...
As generative AI tools become more integrated into creative workflows, questions of ownership in co-creative contexts have become increasingly urgen...
Chain-of-thought reasoning has significantly improved the performance of Large Language Models (LLMs) across various domains. However, this reasonin...
Large vision-language models (LVLMs) remain vulnerable to hallucination, often generating content misaligned with visual inputs. While recent approa...
With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for t...
Accurately estimating the refractive environment over multiple frequencies within the marine atmospheric boundary layer is crucial for the effective...
The intrication of brain signals drives research that leverages multimodal AI to align brain modalities with visual and textual data for explainable...
Fine-grained edited image detection of localized edits in images is crucial for assessing content authenticity, especially given that modern diffusi...
We revisit the classical problem of Bayesian ensembles and address the challenge of learning optimal combinations of Bayesian models in an online, c...
Generalization of deep-learning-based (DL) computer vision algorithms to various image perturbations is hard to establish and remains an active area...
Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to r...