Latest AI and machine learning research in care of terminally ill / palliative care for healthcare professionals.
Understanding pain-related facial behaviors is essential for digital healthcare in terms of effective monitoring, assisted diagnostics, and treatment planning, particularly for patients unable to communicate verbally. Existing data-driven methods of detecting pain from facial expressions are limited due to interpretability and severity quantification. To this end, we propose GraphAU-Pain, levera...
Text Image Machine Translation (TIMT)-the task of translating textual content embedded in images-is critical for applications in accessibility, cross-lingual information access, and real-world document understanding. However, TIMT remains a complex challenge due to the need for accurate optical character recognition (OCR), robust visual-text reasoning, and high-quality translation, often requiri...
Modality fusion is a cornerstone of multimodal learning, enabling information integration from diverse data sources. However, vanilla fusion methods...
Early identification of weeds is essential for effective management and control, and there is growing interest in automating the process using compu...
Current care in multiple sclerosis (MS) primarily relies on infrequently obtained data such as magnetic resonance imaging, clinical laboratory tests o...
Brain network analysis plays a crucial role in diagnosing and monitoring neurodegenerative disorders such as Alzheimer's disease (AD). Existing appr...
Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language mode...
Front-end engineering involves a complex workflow where engineers conceptualize designs, translate them into code, and iteratively refine the implem...
Front-end engineering involves a complex workflow where engineers conceptualize designs, translate them into code, and iteratively refine the implem...
Gesture recognition presents a promising avenue for interfacing with unmanned aerial vehicles (UAVs) due to its intuitive nature and potential for p...
To encode point clouds containing both geometry and attributes, most learning-based compression schemes treat geometry and attribute coding separate...
We tackle the problem of monocular-to-stereo video conversion and propose a novel architecture for inpainting and refinement of the warped right vie...
We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for ...
With recent breakthroughs in large-scale modeling, the Segment Anything Model (SAM) has demonstrated significant potential in a variety of visual ap...
Integer quantization has emerged as a critical technique to facilitate deployment on resource-constrained devices. Although they do reduce the compl...
End-to-end (E2E) autonomous driving systems offer a promising alternative to traditional modular pipelines by reducing information loss and error ac...
Data engineering pipelines are essential - albeit costly - components of predictive analytics frameworks requiring significant engineering time and ...
In recent years, large-scale pre-trained diffusion transformer models have made significant progress in video generation. While current DiT models c...
Effective road infrastructure management is crucial for modern society. Traditional manual inspection techniques remain constrained by cost, efficie...
This paper introduces a comprehensive end-to-end pipeline for Optical Character Recognition (OCR) on Urdu newspapers. In our approach, we address th...