Latest AI and machine learning research in identifying and reporting child abuse for healthcare professionals.
Recent advancements in Large Vision-Language Models built upon Large Language Models have established aligning visual features with LLM representations as the dominant paradigm. However, inherited LLM architectural designs introduce suboptimal characteristics for multimodal processing. First, LVLMs exhibit a bimodal distribution in attention allocation, leading to the progressive neglect of midd...
Under extreme low-light conditions, traditional frame-based cameras, due to their limited dynamic range and temporal resolution, face detail loss and motion blur in captured images. To overcome this bottleneck, researchers have introduced event cameras and proposed event-guided low-light image enhancement algorithms. However, these methods neglect the influence of global low-frequency noise caus...
Intrinsically disordered regions (IDRs) account for one-third of the human proteome and play essential biological roles. However, predicting the fun...
Video anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities i...
In this paper, we investigate the challenges associated with using egocentric devices to photorealistic reconstruct the scene in high dynamic range....
Multimodal image fusion effectively aggregates information from diverse modalities, with fused images playing a crucial role in vision systems. Howe...
The detection of ligand binding sites for proteins is a fundamental step in Structure-Based Drug Design. Despite notable advances in recent years, e...
Recent studies, including DeepSeek-R1 and Kimi-k1.5, have demonstrated that reinforcement learning with rule-based, binary-valued reward functions c...
Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and a...
Industrial Internet of Things environments increasingly rely on advanced Anomaly Detection and explanation techniques to rapidly detect and mitigate...
Urban segregation refers to the physical and social division of people, often driving inequalities within cities and exacerbating socioeconomic and ra...
As an important atmospheric pollutant causing serious harm to human health and the natural environment, monitoring of surface NO (SNO) level is of cri...
Blockchain protocols incentivize participation through monetary rewards, assuming rational actors behave honestly to maximize their gains. However, ...
Domain-specific instruction-tuning has become the defacto standard for improving the performance of large language models (LLMs) in specialized appl...
White Light Imaging (WLI) and Narrow Band Imaging (NBI) are the two main colonoscopic modalities for polyp classification. While NBI, as optical chr...
Recent advances in Multimodal Large Language Models (MLLMs) have shown promising results in integrating diverse modalities such as texts and images....
Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries ...
High-resolution remote sensing (HRRS) image segmentation is challenging due to complex spatial layouts and diverse object appearances. While CNNs ex...
Domain adaptation remains a challenge when there is significant manifold discrepancy between source and target domains. Although recent methods leve...
Omni-domain infrared small target detection (IRSTD) poses formidable challenges, as a single model must seamlessly adapt to diverse imaging systems,...