Latest AI and machine learning research in identifying and reporting child abuse for healthcare professionals.
Reinforcement Learning from Human/AI Feedback (RLHF/RLAIF) has been extensively utilized for preference alignment of text-to-image models. Existing methods face certain limitations in terms of both data and algorithm. For training data, most approaches rely on manual annotated preference data, either by directly fine-tuning the generators or by training reward models to provide training signals....
Brain decoding currently faces significant challenges in individual differences, modality alignment, and high-dimensional embeddings. To address individual differences, researchers often use source subject data, which leads to issues such as privacy leakage and heavy data storage burdens. In modality alignment, current works focus on aligning the softmax probability distribution but neglect the ...
As the demand for efficient, low-power computing in embedded and edge devices grows, traditional computing methods are becoming less effective for h...
Reference-based sketch colorization methods have garnered significant attention due to their potential applications in the animation production indu...
Recent studies focus on the Remote Sensing Image-Text Retrieval (RSITR), which aims at searching for the corresponding targets based on the given qu...
The proliferation of AI-generated content brings significant concerns on the forensic and security issues such as source tracing, copyright protecti...
Knowledge Distillation (KD) compresses neural networks by learning a small network (student) via transferring knowledge from a pre-trained large net...
Category-agnostic pose estimation aims to locate keypoints on query images according to a few annotated support images for arbitrary novel classes. ...
Painting textures for existing geometries is a critical yet labor-intensive process in 3D asset generation. Recent advancements in text-to-image (T2...
Deep Reinforcement Learning (DRL) has demonstrated strong performance in robotic control but remains susceptible to out-of-distribution (OOD) states...
Large Language Models (LLMs) encapsulate a surprising amount of factual world knowledge. However, their performance on temporal questions and histor...
With the continuous emergence of various social media platforms frequently used in daily life, the multimodal meme understanding (MMU) task has been...
Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following huma...
Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either...
We introduce a novel visual tokenization framework that embeds a provable PCA-like structure into the latent token space. While existing visual toke...
Current 3D stylization techniques primarily focus on static scenes, while our world is inherently dynamic, filled with moving objects and changing e...
Unreadable code could be a breeding ground for errors. Thus, previous work defined approaches based on machine learning to automatically assess code...
When discussing the Aerial-Ground Person Re-identification (AGPReID) task, we face the main challenge of the significant appearance variations cause...
Multivariate Time Series Classification (MTSC) is crucial in extensive practical applications, such as environmental monitoring, medical EEG analysi...
We present V$^2$Dial - a novel expert-based model specifically geared towards simultaneously handling image and video input data for multimodal conv...