Latest AI and machine learning research in schizophrenia for healthcare professionals.
This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalFscore, a novel metric built upon the language graph and is designed to evaluate both the accuracy a...
Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-language model (VLM) that solves the Chinese Calligraphy Contextualization (CC$^2$) problem through thr...
Vision-Language Models (VLMs) have advanced multi-modal tasks like image captioning, visual question answering, and reasoning. However, they often g...
Large Visual Language Models (LVLMs) increasingly rely on preference alignment to ensure reliability, which steers the model behavior via preference...
Vision-language models (VLMs) have achieved remarkable advancements, capitalizing on the impressive capabilities of large language models (LLMs) acr...
Score-based diffusion models have achieved incredible performance in generating realistic images, audio, and video data. While these models produce ...
Large multimodal models (LMMs) "see" images by leveraging the attention mechanism between text and visual tokens in the transformer decoder. Ideally...
Object Goal Navigation-requiring an agent to locate a specific object in an unseen environment-remains a core challenge in embodied AI. Although rec...
Large Vision Language Models (LVLMs) are becoming increasingly important in the medical domain, yet Medical LVLMs (Med-LVLMs) frequently generate ha...
Background:Speech patterns have emerged as potential diagnostic markers for conditions with varying etiologies. Machine learning (ML) presents an op...
Depression is a widespread mental health disorder, and clinical interviews are the gold standard for assessment. However, their reliance on scarce p...
Hallucinations are spurious structures not present in the ground truth, posing a critical challenge in medical image reconstruction, especially for ...
In the dynamic landscape of artificial intelligence, the exploration of hallucinations within vision-language (VL) models emerges as a critical fron...
Large Language Models (LLMs) excel in language comprehension and generation but are prone to hallucinations, producing factually incorrect or unsupp...
The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the mod...
Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffe...
Advancements in Large Vision-Language Models (LVLMs) have demonstrated promising performance in a variety of vision-language tasks involving image-c...
With the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also ...
Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their perf...
In this study, we investigate the feasibility of using a human-centered artificial intelligence (AI) chat platform where medical specialists collabo...