Latest AI and machine learning research in care of terminally ill / palliative care for healthcare professionals.
Normalizing Flows (NFs) are a classical family of likelihood-based methods that have received revived attention. Recent efforts such as TARFlow have shown that NFs are capable of achieving promising performance on image modeling tasks, making them viable alternatives to other methods such as diffusion models. In this work, we further advance the state of Normalizing Flow generative models by i...
Vision-Language Models(VLMs) excel at autoregressive text generation, yet end-to-end autonomous driving requires multi-task learning with structured outputs and heterogeneous decoding behaviors, such as autoregressive language generation, parallel object detection and trajectory regression. To accommodate these differences, existing systems typically introduce separate or cascaded decoders, result...
Grounded Multimodal Named Entity Recognition (GMNER) aims to jointly identify named entity mentions in text, predict their semantic types, and ground ...
Quantitative analysis of animal behavior is fundamental to neuroscience and ethology but remains constrained by the scalability, subjectivity, and lim...
DNA-based storage has emerged as a promising approach to the global data crisis, offering molecular-scale density and millennial-scale stability at lo...
Fall detection in elderly care requires not only accurate classification but also reliable explanations that clinicians can trust. However, existing p...
Proteins are fundamental macromolecules involved in virtually all biological processes. Their physiological roles are tightly linked to their three-di...
Task-adapted compressed sensing magnetic resonance imaging (CS-MRI) is emerging to address the specific demands of downstream clinical tasks with sign...
Traditional Active Noise Control (ANC) systems are mostly based on FxLMS algorithms, but such algorithms rely on linear assumptions and are often limi...
Reliable recognition of standard cine cardiac MRI views is essential because each view determines which cardiac anatomy is visualized and which quanti...
Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or refer...
Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ...
Video depth estimation is essential for providing 3D scene structure in applications ranging from autonomous driving to mixed reality. Current end-to-...
Developing optical systems for free-space applications requires simulation tools that accurately capture turbulence-induced wavefront distortions and ...
Decreasing sequence length is a common way to accelerate transformers, but prior token reduction work often targets classification and reports proxy m...
World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the...
Deep learning models utilizing longitudinal healthcare data have significantly advanced epidemiological research. However, contemporary transformer-ba...
End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies thr...
Most of the recent generative image super-resolution (SR) methods rely on adapting large text-to-image (T2I) diffusion models pretrained on web-scale ...
This paper proposes an end-to-end shared attention estimation method via group detection. Most previous methods estimate shared attention (SA) without...