Latest AI and machine learning research in identifying and reporting child abuse for healthcare professionals.
Despite significant progress in Unified Multimodal Retrieval (UMR) powered by Large Multimodal Models (LMMs), existing embedding methods primarily focus on sample-level objectives via contrastive learning while overlooking the crucial subject-level semantics. This limitation hinders the model's ability to group semantically coherent subjects in complex multimodal queries, manifesting as semantic a...
Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. Current multi-modal approaches mainly focus on integrating complementary visual modalities, yet neglect the incorporating of non-visual textual data - a rich source of knowledge that can bridge semantic gaps between visua...
With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-sh...
Multimodal Large Language Models (MLLMs) have been increasingly used as automatic evaluators-a paradigm known as MLLM-as-a-Judge. However, their relia...
Wideband spectrum sensing for low-altitude monitoring is critical yet challenging due to heterogeneous protocols,large bandwidths, and non-stationary ...
Cyclic peptides are recognized as versatile scaffolds for therapeutic and functional applications due to their structural stability and resistance to ...
Recent semi-dense image matching methods have achieved remarkable success, but two long-standing issues still impair their performance. At the coarse ...
Recent advances in open-vocabulary object detection focus primarily on two aspects: scaling up datasets and leveraging contrastive learning to align l...
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm. It aims to retrieve target images from large-scale image databases that are ...
In clinical practice, crossmodal information including medical images and tabular data is essential for disease diagnosis. There exists a significant ...
The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the...
Human Activity Recognition using wearable inertial sensors is foundational to healthcare monitoring, fitness analytics, and context-aware computing, y...
The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the...
Mechanical ventilation (MV) is a life-saving intervention for patients with acute respiratory failure (ARF) in the ICU. However, inappropriate ventila...
Large pretrained diffusion models have significantly enhanced the quality of generated videos, and yet their use in real-time streaming remains limite...
Large pretrained diffusion models have significantly enhanced the quality of generated videos, and yet their use in real-time streaming remains limite...
Existing concept customization methods have achieved remarkable outcomes in high-fidelity and multi-concept customization. However, they often neglect...
Accurate estimation of enzyme kinetic parameters is essential for enzyme engineering and industrial biocatalysis, yet their experimental measurement r...
Infrared image super-resolution (IISR) under real-world conditions is a practically significant yet rarely addressed task. Pioneering works are often ...
Safe visual navigation is critical for indoor mobile robots operating in cluttered environments. Existing benchmarks, however, often neglect collision...