Latest AI and machine learning research in identifying and reporting dependent adult abuse for healthcare professionals.
Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to ground-based VLN, aerial VLN requires the agent to decide the next action in both horizontal and vertical directions based on the first-person view observations. Previous methods struggle to perform well due to the longer n...
Current practices for reporting the level of differential privacy (DP) guarantees for machine learning (ML) algorithms provide an incomplete and potentially misleading picture of the guarantees and make it difficult to compare privacy levels across different settings. We argue for using Gaussian differential privacy (GDP) as the primary means of communicating DP guarantees in ML, with the full p...
Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either...
We introduce a novel visual tokenization framework that embeds a provable PCA-like structure into the latent token space. While existing visual toke...
Current 3D stylization techniques primarily focus on static scenes, while our world is inherently dynamic, filled with moving objects and changing e...
Unreadable code could be a breeding ground for errors. Thus, previous work defined approaches based on machine learning to automatically assess code...
When discussing the Aerial-Ground Person Re-identification (AGPReID) task, we face the main challenge of the significant appearance variations cause...
Multivariate Time Series Classification (MTSC) is crucial in extensive practical applications, such as environmental monitoring, medical EEG analysi...
Automated CT report generation plays a crucial role in improving diagnostic accuracy and clinical workflow efficiency. However, existing methods lac...
We present V$^2$Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conv...
Single Image Reflection Removal (SIRR) is a canonical blind source separation problem and refers to the issue of separating a reflection-contaminate...
Soybean leaf disease detection is critical for agricultural productivity but faces challenges due to visually similar symptoms and limited interpret...
Background Multimodal generative artificial intelligence (AI) technologies can produce preliminary radiology reports, and validation with reader studi...
Wearable sensing devices, such as Holter monitors, will play a crucial role in the future of digital health. Unsupervised learning frameworks such a...
Two-tower models are widely adopted in the industrial-scale matching stage across a broad range of application domains, such as content recommendati...
The accurate assessment of the brain's functional network is seen as crucial for the understanding of complex relationships between different brain re...
Neural Radiance Fields (NeRF) have been gaining attention as a significant form of 3D content representation. With the proliferation of NeRF-based c...
Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting i...
Medical image segmentation plays a crucial role in various clinical applications. A major challenge in medical image segmentation is achieving accur...
When an individual reports a negative interaction with some system, how can their personal experience be contextualized within broader patterns of s...