Latest AI and machine learning research in care of terminally ill / palliative care for healthcare professionals.
We present Sparsh-X, the first multisensory touch representations across four tactile modalities: image, audio, motion, and pressure. Trained on ~1M contact-rich interactions collected with the Digit 360 sensor, Sparsh-X captures complementary touch signals at diverse temporal and spatial scales. By leveraging self-supervised learning, Sparsh-X fuses these modalities into a unified representatio...
In the era of large models and big data, the security of optical fiber communication backbone networks has garnered significant attention. Quantum noise stream cipher (QNSC) stands as a crucial method for safeguarding the physical layer security of optical fiber communications, yet the current schemes lag behind the rate capabilities of existing 400G optical fiber backbone networks. In this paper,...
A novel self-adaptive secure end-to-end (E2E) transmission approach is proposed for a radio-over-fiber (RoF) system. The system integrates deep learni...
Autoregressive Transformers are increasingly being deployed as end-to-end robot and autonomous vehicle (AV) policy architectures, owing to their sca...
Multimodal chatbots have become one of the major topics for dialogue systems in both research community and industry. Recently, researchers have she...
Infrared small target detection (IRSTD) remains a long-standing challenge in complex backgrounds due to low signal-to-clutter ratios (SCR), diverse ...
End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. Howev...
We propose a feed-forward Gaussian Splatting model that unifies 3D scene and semantic field reconstruction. Combining 3D scenes with semantic fields...
The continuous improvements on image compression with variational autoencoders have lead to learned codecs competitive with conventional approaches ...
End-to-end autonomous driving has emerged as a dominant paradigm, yet its highly entangled black-box models pose significant challenges in terms of ...
End-to-end multi-modal planning is a promising paradigm in autonomous driving, enabling decision-making with diverse trajectory candidates. A key co...
In complex driving environments, autonomous vehicles must navigate safely. Relying on a single predicted path, as in regression-based approaches, us...
A central challenge in modern language models (LMs) is intrinsic hallucination: the generation of information that is plausible but unsubstantiated ...
Image matching, which establishes correspondences between two-view images to recover 3D structure and camera geometry, serves as a cornerstone in co...
The detection of ligand binding sites for proteins is a fundamental step in Structure-Based Drug Design. Despite notable advances in recent years, e...
When developing control laws for robotic systems, the principle factor when examining their performance is choosing inputs that allow smooth trackin...
Pre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in co...
In recent years, federated learning (FL) has made significant advance in privacy-sensitive applications. However, it can be hard to ensure that FL p...
Spatial intelligence, encompassing 3D reconstruction, perception, and reasoning, is fundamental to applications such as robotics, aerial imaging, an...
Spatial intelligence, encompassing 3D reconstruction, perception, and reasoning, is fundamental to applications such as robotics, aerial imaging, an...