Latest AI and machine learning research in schizophrenia for healthcare professionals.
Large Visual Language Models (LVLMs) increasingly rely on preference alignment to ensure reliability, which steers the model behavior via preference fine-tuning on preference data structured as ``image - winner text - loser text'' triplets. However, existing approaches often suffer from limited diversity and high costs associated with human-annotated preference data, hindering LVLMs from fully a...
Vision-language models (VLMs) have achieved remarkable advancements, capitalizing on the impressive capabilities of large language models (LLMs) across diverse tasks. Despite this, a critical challenge known as hallucination occurs when models overconfidently describe objects or attributes absent from the image, a problem exacerbated by the tendency of VLMs to rely on linguistic priors. This lim...
Score-based diffusion models have achieved incredible performance in generating realistic images, audio, and video data. While these models produce ...
Large multimodal models (LMMs) "see" images by leveraging the attention mechanism between text and visual tokens in the transformer decoder. Ideally...
Object Goal Navigation-requiring an agent to locate a specific object in an unseen environment-remains a core challenge in embodied AI. Although rec...
Large Vision Language Models (LVLMs) are becoming increasingly important in the medical domain, yet Medical LVLMs (Med-LVLMs) frequently generate ha...
Background:Speech patterns have emerged as potential diagnostic markers for conditions with varying etiologies. Machine learning (ML) presents an op...
Depression is a widespread mental health disorder, and clinical interviews are the gold standard for assessment. However, their reliance on scarce p...
Hallucinations are spurious structures not present in the ground truth, posing a critical challenge in medical image reconstruction, especially for ...
In the dynamic landscape of artificial intelligence, the exploration of hallucinations within vision-language (VL) models emerges as a critical fron...
Large Language Models (LLMs) excel in language comprehension and generation but are prone to hallucinations, producing factually incorrect or unsupp...
The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the mod...
Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffe...
Advancements in Large Vision-Language Models (LVLMs) have demonstrated promising performance in a variety of vision-language tasks involving image-c...
With the rise of generative models, there is a growing interest in unifying all tasks within a generative framework. Anomaly detection methods also ...
Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their perf...
In this study, we investigate the feasibility of using a human-centered artificial intelligence (AI) chat platform where medical specialists collabo...
Foundation Models that are capable of processing and generating multi-modal data have transformed AI's role in medicine. However, a key limitation o...
Vision-language models in pathology enable multimodal case retrieval and automated report generation. Many of the models developed so far, however, ...
Self-supervised learning (SSL) vision encoders learn high-quality image representations and thus have become a vital part of developing vision modal...