Latest AI and machine learning research in universal precautions for healthcare professionals.
Navigating peripersonal space requires reaching targets in both horizontal (e.g., desks) and vertical (e.g., shelves) layouts with high precision. We developed a haptic glove to aid peri-personal target navigation and investigated the effectiveness of different feedback delivery methods. Twenty-two participants completed target navigation tasks under various conditions, including scene layout (h...
Accurate segmentation of nodules in both 2D breast ultrasound (BUS) and 3D automated breast ultrasound (ABUS) is crucial for clinical diagnosis and treatment planning. Therefore, developing an automated system for nodule segmentation can enhance user independence and expedite clinical analysis. Unlike fully-supervised learning, weakly-supervised segmentation (WSS) can streamline the laborious an...
Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challeng...
Segmentation is a fundamental task in computer vision, with prompt-driven methods gaining prominence due to their flexibility. The Segment Anything ...
The acquisition of annotated datasets with paired images and segmentation masks is a critical challenge in domains such as medical imaging, remote s...
We present a target-aware video diffusion model that generates videos from an input image in which an actor interacts with a specified target while ...
Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face signi...
High-resolution semantic segmentation is essential for applications such as image editing, bokeh imaging, AR/VR, etc. Unfortunately, existing datase...
Recent 3D face editing methods using masks have produced high-quality edited images by leveraging Neural Radiance Fields (NeRF). Despite their impre...
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text p...
While virtual try-on for clothes and shoes with diffusion models has gained attraction, virtual try-on for ornaments, such as bracelets, rings, earr...
Traditional transformer-based semantic segmentation relies on quantized embeddings. However, our analysis reveals that autoencoder accuracy on segme...
Video object segmentation is crucial for the efficient analysis of complex medical video data, yet it faces significant challenges in data availabil...
This paper focuses on a typical uplink transmission scenario over multiple-input multiple-output multiple access channel (MIMO-MAC) and thus propose...
Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with v...
Rectified flow models have achieved remarkable performance in image and video generation tasks. However, existing numerical solvers face a trade-off...
The remarkable performance of large multimodal models (LMMs) has attracted significant interest from the image segmentation community. To align with...
Entity Segmentation (ES) aims at identifying and segmenting distinct entities within an image without the need for predefined class labels. This cha...
Pre-trained segmentation models are a powerful and flexible tool for segmenting images. Recently, this trend has extended to medical imaging. Yet, o...
Solving medical imaging data scarcity through semantic image generation has attracted significant attention in recent years. However, existing metho...