Latest AI and machine learning research in identifying and reporting dependent adult abuse for healthcare professionals.
Medical visual question answering (MedVQA) plays a vital role in clinical decision-making by providing contextually rich answers to image-based queries. Although vision-language models (VLMs) are widely used for this task, they often generate factually incorrect answers. Retrieval-augmented generation addresses this challenge by providing information from external sources, but risks retrieving i...
Garment manipulation is a significant challenge for robots due to the complex dynamics and potential self-occlusion of garments. Most existing methods of efficient garment unfolding overlook the crucial role of standardization of flattened garments, which could significantly simplify downstream tasks like folding, ironing, and packing. This paper presents APS-Net, a novel approach to garment man...
Text on historical maps provides valuable information for studies in history, economics, geography, and other related fields. Unlike structured or s...
Few-shot industrial anomaly detection (FS-IAD) presents a critical challenge for practical automated inspection systems operating in data-scarce env...
We develop a mass-conserving, adaptive-rank solver for the 1D1V Wigner-Poisson system. Our work is motivated by applications to the study of the sto...
Estimating single-cell responses across various perturbations facilitates the identification of key genes and enhances drug screening, significantly...
Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key chal...
Conducting disparity assessments at regular time intervals is critical for surfacing potential biases in decision-making and improving outcomes acro...
Patent images are technical drawings that convey information about a patent's innovation. Patent image retrieval systems aim to search in vast colle...
Incomplete multi-modal medical image segmentation faces critical challenges from modality imbalance, including imbalanced modality missing rates and...
Medical images are usually collected from multiple domains, leading to domain shifts that impair the performance of medical image segmentation model...
Cutting-edge works have demonstrated that text-to-image (T2I) diffusion models can generate adversarial patches that mislead state-of-the-art object...
Peptides can serve as building blocks for supramolecular materials because of their unique ability to self-assemble, offering potential applications i...
Driven by advancements in motion capture and generative artificial intelligence, leveraging large-scale MoCap datasets to train generative models fo...
The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologi...
Modern data marketplaces and data sharing consortia increasingly rely on incentive mechanisms to encourage agents to contribute data. However, schem...
Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event...
Recent advancements in Large Vision-Language Models built upon Large Language Models have established aligning visual features with LLM representati...
Under extreme low-light conditions, traditional frame-based cameras, due to their limited dynamic range and temporal resolution, face detail loss an...
Intrinsically disordered regions (IDRs) account for one-third of the human proteome and play essential biological roles. However, predicting the fun...