Latest AI and machine learning research in schizophrenia for healthcare professionals.
In standard large vision-language models (LVLMs) pre-training, the model typically maximizes the joint probability of the caption conditioned on the image via next-token prediction (NTP); however, since only a small subset of caption tokens directly relates to the visual content, this naive NTP unintentionally fits the model to noise and increases the risk of hallucination. We present PRIOR, a s...
Visual planning, by offering a sequence of intermediate visual subgoals to a goal-conditioned low-level policy, achieves promising performance on long-horizon manipulation tasks. To obtain the subgoals, existing methods typically resort to video generation models but suffer from model hallucination and computational cost. We present Vis2Plan, an efficient, explainable and white-box visual planni...
The Cancer Genome Atlas (TCGA) has enabled novel discoveries and served as a large-scale reference through its harmonized genomics, clinical, and im...
Despite significant advancements in multimodal reasoning tasks, existing Large Vision-Language Models (LVLMs) are prone to producing visually ungrou...
In the age of social media, the rapid spread of misinformation and rumors has led to the emergence of infodemics, where false information poses a si...
Vision-Language Models (VLMs) are becoming increasingly popular in the medical domain, bridging the gap between medical images and clinical language...
The scarcity of high-quality multimodal biomedical data limits the ability to effectively fine-tune pretrained Large Language Models (LLMs) for specia...
BACKGROUND: Electroencephalography (EEG) is a noninvasive, cost-effective, and robust tool, which directly measures in vivo neuronal mass activity wit...
BACKGROUND AND OBJECTIVE: Accurate detection of schizophrenia poses a grand challenge as a complex and heterogeneous mental disorder. Current diagnost...
BACKGROUND: Cognitive deficits are a central feature of schizophrenia for which there are not any established pharmacological treatments. Antipsychoti...
Retrieval-Augmented Generation (RAG) systems have gained widespread adoption by application builders because they leverage sources of truth to enabl...
Hallucinations in vision-language models (VLMs) hinder reliability and real-world applicability, usually stemming from distribution shifts between p...
Zero-shot learning (ZSL) aims to recognize unseen classes by aligning images with intermediate class semantics, like human-annotated concepts or cla...
The collaborative paradigm of large and small language models (LMs) effectively balances performance and cost, yet its pivotal challenge lies in pre...
The acquisition of information-rich images within a limited time budget is crucial in medical imaging. Medical image translation (MIT) can help enha...
Accurate extraction of key information from 2D engineering drawings is crucial for high-precision manufacturing. Manual extraction is time-consuming...
Hallucinations in large language models (LLMs) present a growing challenge across real-world applications, from healthcare to law, where factual rel...
IMPORTANCE: The diagnosis of schizophrenia and bipolar disorder is often delayed several years despite illness typically emerging in late adolescence ...
Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently g...
Mental and cognitive representations are believed to reside on low-dimensional, non-linear manifolds embedded within high-dimensional brain activity...