Latest AI and machine learning research in work force for healthcare professionals.
The rapid proliferation of AI-generated images, powered by generative adversarial networks (GANs), diffusion models, and other synthesis techniques, has raised serious concerns about misinformation, copyright violations, and digital security. However, detecting such images in a generalized and robust manner remains a major challenge due to the vast diversity of generative models and data distribut...
Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to handle this limitation by injecting geometry tokens from pretrained 3D foundation models into VLMs. Nevertheless, we observe that naive token fusion followed by ...
Pediatric bipolar disorder is challenging to diagnose accurately due to symptom heterogeneity. More standardized and data-driven approaches are needed...
We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasonin...
Recent multimodal large language models are computationally expensive because Transformers must process a large number of visual tokens. We present \t...
Recent advances in multimodal vision-language models (VLMs) have enabled joint reasoning over visual and textual information, yet their application to...
Clinical documentation is a critical factor for patient safety, diagnosis, and continuity of care. The administrative burden of EHRs is a significant ...
Long-tail distributions in driving datasets pose a fundamental challenge for 3D perception, as rare classes exhibit substantial intra-class diversity ...
The expansion of generative AI and LLM services underscores the growing need for adaptive mechanisms to select an appropriate available model to respo...
Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimizat...
The expansion of generative AI and LLM services underscores the growing need for adaptive mechanisms to select an appropriate available model to respo...
tRNA are adapter molecules with an integral role in translation and further roles in stress adaptation. Processing of tRNA is tightly regulated and in...
Predicting future states in uncertain environments, such as wildfire spread, medical diagnosis, or autonomous driving, requires models that can consid...
Frailty is a prevalent geriatric syndrome, and the shortage of objective biomarkers restricts its early diagnosis and intervention. This study aimed t...
Background: The administrative burden of clinical documentation is a recognised contributor to clinician burnout and diminished care quality. Ambient ...
Collecting and annotating datasets for pixel-level semantic segmentation tasks are highly labor-intensive. Data augmentation provides a viable solutio...
YOLO detectors are known for their fast inference speed, yet training them remains unexpectedly time-consuming due to their exhaustive pipeline that p...
In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with accurate l...
Creating high-fidelity, animatable 3D dog avatars remains a formidable challenge in computer vision. Unlike human digital doubles, animal reconstructi...
While hyperspectral images (HSI) benefit from numerous spectral channels that provide rich information for classification, the increased dimensionalit...