Latest AI and machine learning research in adhd/add for healthcare professionals.
Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and limited temporally aware datasets. Existing methods for temporal expression either conflate time with text-based numerical values, add a series of dedicated temporal tokens, or regress time using specialized temporal g...
Recent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods. As TTS systems continue to evolve, detection models must be able to efficiently adapt to previously unseen generation models with minimal data. This paper introduces ADD-GP, a few-shot adaptive framework based on a Gaussian Process...
Prompt learning is a crucial technique for adapting pre-trained multimodal language models (MLLMs) to user tasks. Federated prompt personalization (...
Software systems have grown as an indispensable commodity used across various industries, and almost all essential services depend on them for effec...
Despite rapid advances in large language models (LLMs), their integration with traditional supervised machine learning (ML) techniques that have pro...
Embedding-based collaborative filtering, often coupled with nearest neighbor search, is widely deployed in large-scale recommender systems for perso...
Embedding-based collaborative filtering, often coupled with nearest neighbor search, is widely deployed in large-scale recommender systems for perso...
Text-image-to-video (TI2V) generation is a critical problem for controllable video generation using both semantic and visual conditions. Most existi...
Recent advancements in Neural Radiance Fields (NeRF) and 3D Gaussian-based Simultaneous Localization and Mapping (SLAM) methods have demonstrated ex...
Large language models are designed to encode general purpose knowledge about the world from Internet data. Yet, a wealth of information falls outsid...
White Light Imaging (WLI) and Narrow Band Imaging (NBI) are the two main colonoscopic modalities for polyp classification. While NBI, as optical chr...
Performativity of predictions refers to the phenomena that prediction-informed decisions may influence the target they aim to predict, which is wide...
Conformal prediction (CP) provides sets of candidate classes with a guaranteed probability of containing the true class. However, it typically relie...
The extraction of visual features is an essential step in Visual Question Answering (VQA). Building a good visual representation of the analyzed sce...
This technical report presents a natural language processing (NLP)-based approach for systematically classifying scientific literature on childhood ...
Contemporary Quranic Orthography (CQO) relies on a precise system of phonetic notation that can be traced back to the early stages of Islam, when th...
This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical C...
According to the World Health Organization (WHO), approximately 5% of children and 2.5% of adults suffer from attention deficit hyperactivity disorde...
Visually impaired individuals face daily challenges in social engagement and routine activities due to limited access to real-time environmental infor...
The aim of this study was to establish a nomogram based on clinical, radiomics, and deep transfer learning (DTL) features to predict meningioma grade....