Latest AI and machine learning research in covid-19 for healthcare professionals.
The recent Segment Anything Model 2 (SAM2) has demonstrated exceptional capabilities in interactive object segmentation for both images and videos. However, as a foundational model on interactive segmentation, SAM2 performs segmentation directly based on mask memory from the past six frames, leading to two significant challenges. Firstly, during inference in videos, objects may disappear since S...
The development of powerful user representations is a key factor in the success of recommender systems (RecSys). Online platforms employ a range of RecSys techniques to personalize user experience across diverse in-app surfaces. User representations are often learned individually through user's historical interactions within each surface and user representations across different surfaces can be ...
Contemporary diffusion models built upon U-Net or Diffusion Transformer (DiT) architectures have revolutionized image generation through transformer...
Unsupervised novelty detection (UND), aimed at identifying novel samples, is essential in fields like medical diagnosis, cybersecurity, and industri...
Medical image segmentation is a crucial and time-consuming task in clinical care, where mask precision is extremely important. The Segment Anything ...
The perfect alignment of 3D echocardiographic images captured from various angles has improved image quality and broadened the field of view. This s...
Human image animation aims to generate human videos of given characters and backgrounds that adhere to the desired pose sequence. However, existing ...
Face registration deforms a template mesh to closely fit a 3D face scan, the quality of which commonly degrades in non-skin regions (e.g., hair, bea...
Industrial Anomaly Detection (IAD) is critical for ensuring product quality by identifying defects. Traditional methods such as feature embedding an...
The rapid evolution of malware variants requires robust classification methods to enhance cybersecurity. While Large Language Models (LLMs) offer po...
Among the new techniques of Versatile Video Coding (VVC), the quadtree with nested multi-type tree (QT+MTT) block structure yields significant codin...
The hematology analytics used for detection and classification of small blood components is a significant challenge. In particular, when objects exi...
To jointly tackle the challenges of data and node heterogeneity in decentralized learning, we propose a distributed strong lottery ticket hypothesis...
The RGB-Depth (RGB-D) Video Object Segmentation (VOS) aims to integrate the fine-grained texture information of RGB with the spatial geometric clues...
This study performs a comprehensive evaluation of quantitative measurements as extracted from automated deep-learning-based segmentation methods, be...
Multi-label requirements classification is a challenging task, especially when dealing with numerous classes at varying levels of abstraction. The d...
In today's age of social media and marketing, copyright issues can be a major roadblock to the free sharing of images. Generative AI models have mad...
This paper focuses on a key challenge in visual art understanding: given an art image, the model pinpoints pixel regions that trigger a specific hum...
Zero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of a...
We introduce CrossWKV, a novel cross-attention mechanism for the state-based RWKV-7 model, designed to enhance the expressive power of text-to-image...