Latest AI and machine learning research in medicare for healthcare professionals.
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward l...
Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source. While effective at scale, this paradigm underexplores an important alternative source of supervision: a range of existing multi-label classification datasets, which provide cleaner and more explicit disease signals than ...
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the...
Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengt...
Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents t...
Objective: To develop and validate a supervised text-embedded transformer matching model to identify fall injuries in Medicare data, and evaluate the ...
Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual...
DNA encodes biological function across a continuum of sequence scales, from single-nucleotide and motif-level grammar to regulatory neighborhoods, chr...
Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods sti...
Detecting when the statistical behavior of an engineered system changes, and identifying which component is responsible, are core problems in the moni...
Many high-resolution imaging systems face the same fundamental question: when have enough measurements been collected to reconstruct an image accurate...
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under s...
Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness...
Efficient border control is becoming a significant global challenge, mainly due to severe congestion and extended passenger waiting times. To mitigate...
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize hi...
Objective This study was designed to evaluate the diagnostic performance of fluorescence-based rapid on-site specimen evaluation (F-ROSE) for antimicr...
Background: In Japan, acute inpatient care is divided into approximately 335 secondary medical care areas, which serve as the basic units for planning...
People with Down syndrome have higher age-specific mortality rates compared to the general population as well as peers with other intellectual and dev...
Background. Predicting conversion from mild cognitive impairment (MCI) to Alzheimer's disease (AD) is central to trial enrichment and care planning, y...
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limi...