Latest AI and machine learning research in medicare for healthcare professionals.
Wildfire monitoring from UAVs requires reliable reasoning over complex aerial scenes, where smoke, scale variation, and occlusions often limit RGB-only interpretation. We introduce FlameVQA, a multiple-choice visual question answering benchmark for UAV-based wildfire intelligence built on FLAME 3, leveraging paired RGB imagery and radiometric thermal TIFFs for temperature-grounded, safety-critical...
Recent image generation models achieve impressive quality in single-image synthesis, but often fail to maintain consistency across sequential outputs, as required in comics, storyboards, and visual narratives. We propose Long-Context Generation (LCG), a framework for long-context multi-image text-to-image generation, to improve consistency and scalability in long-context multi-image generation. LC...
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object...
Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by c...
Purpose: To investigate how artificial intelligence (AI) systems detect referrable diabetic retinopathy (DR) from retinal photographs by analysing hea...
Reliable navigation in underwater environments remains a key challenge in marine robotics. In such scenarios, forward-looking sonars are a natural cho...
Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates, multichann...
Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting t...
Nickel has been studied for a long time as an environmental contaminant but less so in its connection to population health. It does not announce itsel...
Composed image retrieval (CIR) uses a reference image and a text modification to search for a target image. However, such queries often describe sever...
Alternative isoform usage can alter gene function independently of total gene expression, creating a need to resolve transcript isoforms at single-cel...
Fingerprint recognition is still dominated by task-specific pipelines, where enhancement, structural parsing, alignment, and matching are optimized in...
Coastal algal bloom monitoring requires frequent, spatially detailed, and globally consistent observations, provided by Landsat-8/9 and Sentinel-2 A/B...
Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven ...
Large vision-language models (LVLMs) hallucinate: they assert visual details that the image does not support. A principled remedy is selective predict...
Background and rationale: Knee osteoarthritis (KOA) is a leading cause of lower limb disability worldwide, characterized by functional limitations, st...
Remote photoplethysmography (rPPG) transformers achieve low heart-rate error on benchmarks, yet their decisions remain opaque--a growing concern as rP...
Diffusion models have emerged as state-of-the-art generative models for high-fidelity image synthesis, particularly in their classifier-free guided an...
Background and purpose: oART enables daily plan adaptation to interfraction anatomical variations, but cumulative dose estimation remains limited by D...
Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large mul...