Latest AI and machine learning research in work force for healthcare professionals.
In text-to-image person retrieval tasks, the diversity of natural language expressions and the implicitness of visual semantics often lead to the problem of Expression Drift, where semantically equivalent texts exhibit significant feature discrepancies in the embedding space due to phrasing variations, thereby degrading the robustness of image-text alignment. This paper proposes a semantic compens...
Open-vocabulary remote sensing image segmentation (OVRSIS) remains underexplored due to fragmented datasets, limited training diversity, and the lack of evaluation benchmarks that reflect realistic geospatial application demands. Our previous \textit{OVRSISBenchV1} established an initial cross-dataset evaluation protocol, but its limited scope is insufficient for assessing realistic open-world gen...
Masked image modeling (MIM) is a highly effective self-supervised learning (SSL) approach to extract useful feature representations from unannotated d...
Composed Image Retrieval (CIR) aims to retrieve target images by integrating a reference image with a corresponding modification text. CIR requires jo...
Background: Datasets related to infectious diseases are essential for public health decision-making, yet their reuse remains limited by persistent bar...
Flow matching has emerged as a powerful generative framework, with recent few-step methods achieving remarkable inference acceleration. However, we id...
We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input...
While the shortage of explicit action data limits Vision-Language-Action (VLA) models, human action videos offer a scalable yet unlabeled data source....
Deep learning underpins a wide range of applications in MRI, including reconstruction, artifact removal, and segmentation. However, progress has been ...
Existing defect/anomaly generation methods often rely on few-shot learning, which overfits to specific defect categories due to the lack of large-scal...
Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse down...
Cell state diversity drives tissue adaptability, repair, and disease resilience, but fully capturing this cellular complexity remains a challenge. Mos...
Ask a frontier model how to taper six milligrams of alprazolam (psychiatrist retired, ten days of pills left, abrupt cessation causes seizures) and it...
Ship detection for navigation is a fundamental perception task in intelligent waterway transportation systems. However, existing public ship detection...
RL training of multi-turn LLM agents is inherently unstable, and reasoning quality directly determines task performance. Entropy is widely used to tra...
Data heterogeneity hinders clinical deployment of medical image analysis models, and generative data augmentation helps mitigate this issue. However, ...
Background: Clinical documentation and information retrieval consume over half of physicians working hours, contributing to cognitive overload and bur...
Generating realistic single-cell transcriptomic profiles from structured biological descriptions would enable controlled simulation, data augmentation...
Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, con...
Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. W...