Latest AI and machine learning research in schizophrenia for healthcare professionals.
Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, content hallucination, safety concerns, and bias. Addressing these limitations, we introduce MJ-BENCH-VIDEO, a large-scale video preference benchmark designed to evaluate video ge...
Magnetic Resonance Imaging generally requires long exposure times, while being sensitive to patient motion, resulting in artifacts in the acquired images, which may hinder their diagnostic relevance. Despite research efforts to decrease the acquisition time, and designing efficient acquisition sequences, motion artifacts are still a persistent problem, pushing toward the need for the development...
Hallucination has been a long-standing and inevitable problem that hinders the application of Large Vision-Language Models (LVLMs) in domains that r...
Mental illness is a widespread and debilitating condition with substantial societal and personal costs. Traditional diagnostic and treatment approac...
Multimodal Large Language Models (MLLMs) still struggle with hallucinations despite their impressive capabilities. Recent studies have attempted to ...
Large language models (LLMs) have demonstrated immense capabilities in understanding textual data and are increasingly being adopted to help researc...
Fusing visual understanding into language generation, Multi-modal Large Language Models (MLLMs) are revolutionizing visual-language applications. Ye...
Despite their impressive performance on multi-modal tasks, large vision-language models (LVLMs) tend to suffer from hallucinations. An important typ...
Vision language models have achieved impressive results across various fields. However, adoption in remote sensing remains limited, largely due to t...
The updated recommendations on diagnostic procedures and treatment pathways for a medical condition are documented as graphical flows in Clinical Pr...
Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities in understanding and describing visual content, achieving state-of-th...
This paper introduces an approach to question answering over knowledge bases like Wikipedia and Wikidata by performing "question-to-question" matchi...
Hallucination remains a major challenge for Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) has gained increasing attenti...
Recently, large language models have shown great potential to transform online medical consultation. Despite this, most research targets improving d...
Effective chart summary can significantly reduce the time and effort decision makers spend interpreting charts, enabling precise and efficient commu...
Traditional similarity-based schema matching methods are incapable of resolving semantic ambiguities and conflicts in domain-specific complex mappin...
Despite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: stat...
We introduce the world's first clinical terminology for the Chinese healthcare community, namely MedCT, accompanied by a clinical foundation model M...
The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abili...
In assistive robotics serving people with disabilities (PWD), accurate place recognition in built environments is crucial to ensure that robots navi...