Latest AI and machine learning research in covid-19 for healthcare professionals.
With the increasing size of frontier LLMs, post-training quantization has become the standard for memory-efficient deployment. Recent work has shown that basic rounding-based quantization schemes pose security risks, as they can be exploited to inject malicious behaviors into quantized models that remain hidden in full precision. However, existing attacks cannot be applied to more complex quanti...
Multimodal large language models (MLLMs) have recently achieved significant progress in visual tasks, including semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on complex tasks involving mathematics and logic. However, their capacity for reasoning tasks involving fine-grained visual understanding remains insufficiently evaluated. To address th...
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hin...
Reasoning Video Object Segmentation is a challenging task, which generates a mask sequence from an input video and an implicit, complex text query. ...
Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange posi...
Open-set image segmentation poses a significant challenge because existing methods often demand extensive training or fine-tuning and generally stru...
Tabular data have been playing a vital role in diverse real-world fields, including healthcare, finance, etc. With the recent success of Large Langu...
Despite the advances in Referring Expression Segmentation (RES) benchmarks, their evaluation protocols remain constrained, primarily focusing on eit...
High-resolution remote sensing (HRRS) image segmentation is challenging due to complex spatial layouts and diverse object appearances. While CNNs ex...
Images are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods...
Medical image classification is critical for clinical decision-making, yet demands for accuracy, interpretability, and generalizability remain chall...
Medical image classification is critical for clinical decision-making, yet demands for accuracy, interpretability, and generalizability remain chall...
BACKGROUND: This study aimed to develop and validate machine learning models to predict pathological complete response (pCR) after neoadjuvant therapy...
Recent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they ...
Rotary Position Embedding (RoPE) is a widely adopted technique for encoding relative positional information in large language models (LLMs). However...
With recent breakthroughs in large-scale modeling, the Segment Anything Model (SAM) has demonstrated significant potential in a variety of visual ap...
While eXplainable AI (XAI) has advanced significantly, few methods address interpretability in embedded vector spaces where dimensions represent com...
The spread of infectious diseases is often influenced by human mobility across different geographical regions. Although numerous studies have invest...
This study presents a robust framework that leverages advanced imaging techniques and machine learning for feature extraction and classification of ...
Autonomous driving systems (ADS) require extensive testing and validation before deployment. However, it is tedious and time-consuming to construct ...