Latest AI and machine learning research in covid-19 for healthcare professionals.
We tackle the challenge of open-vocabulary segmentation, where we need to identify objects from a wide range of categories in different environments, using text prompts as our input. To overcome this challenge, existing methods often use multi-modal models like CLIP, which combine image and text features in a shared embedding space to bridge the gap between limited and extensive vocabulary recog...
Post-training quantization (PTQ) reduces excessive hardware cost by quantizing full-precision models into lower bit representations on a tiny calibration set, without retraining. Despite the remarkable progress made through recent efforts, traditional PTQ methods typically encounter failure in dynamic and ever-changing real-world scenarios, involving unpredictable data streams and continual doma...
The accurate prediction of antigen-antibody structures is essential for advancing immunology and therapeutic development, as it helps elucidate mole...
We introduce Neptune, a benchmark for long video understanding that requires reasoning over long time horizons and across different modalities. Many...
We explore a novel video creation experience, namely Video Creation by Demonstration. Given a demonstration video and a context image from a differe...
Open-vocabulary panoptic segmentation aims to segment and classify everything in diverse scenes across an unbounded vocabulary. Existing methods typ...
In medical imaging, precise annotation of lesions or organs is often required. However, 3D volumetric images typically consist of hundreds or thousa...
Flare, an optical phenomenon resulting from unwanted scattering and reflections within a lens system, presents a significant challenge in imaging. T...
Since pioneering work of Hinton et al., knowledge distillation based on Kullback-Leibler Divergence (KL-Div) has been predominant, and recently its ...
Neural View Synthesis (NVS) has demonstrated efficacy in generating high-fidelity dense viewpoint videos using a image set with sparse views. Howeve...
Diffusion models generate samples by incrementally reversing a process that turns data into noise. We show that when the step size goes to zero, the...
We present a novel machine learning (ML) method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabul...
We present ACDiT, a novel Autoregressive blockwise Conditional Diffusion Transformer, that innovatively combines autoregressive and diffusion paradi...
Low-rank factorization is a popular model compression technique that minimizes the error $\delta$ between approximated and original weight matrices....
In generative models, two paradigms have gained attraction in various applications: next-set prediction-based Masked Generative Models and next-nois...
Linear Transformers have gained attention as efficient alternatives to standard Transformers, but their performance in retrieval and long-context ta...
Deep Learning (DL) libraries, such as PyTorch, are widely used for building and deploying DL models on various hardware platforms. Meanwhile, they a...
We focus on tertiary lymphoid structure (TLS) semantic segmentation in whole slide image (WSI). Unlike TLS binary segmentation, TLS semantic segment...
The Segment Anything Model (SAM), originally built on a 2D Vision Transformer (ViT), excels at capturing global patterns in 2D natural images but st...
Understanding the effects of quarantine policies in populations with underlying social networks is crucial for public health, yet most causal infere...