Latest AI and machine learning research in work force for healthcare professionals.
Visual Grounding is a task that aims to localize a target region in an image based on a free-form natural language description. With the rise of Transformer architectures, there is an increasing need for larger datasets to boost performance. However, the high cost of manual annotation poses a challenge, hindering the scale of data and the ability of large models to enhance their effectiveness. P...
Plant DNA methylation changes occur hundreds to thousands of times faster than DNA mutations and can be transmitted transgenerationally, making them useful for studying population-scale patterns in clonal or selfing species. However, a state-of-the-art approach to use them for inferring population genetic processes and demographic histories is lacking. To address this, we compare evolutionary sign...
The swift evolution of telehealth has revolutionized how medical professionals deliver healthcare services and boost convenience and accessibility. ...
While pre-trained multimodal representations (e.g., CLIP) have shown impressive capabilities, they exhibit significant compositional vulnerabilities...
Recent advances in diffusion models have led to impressive image generation capabilities, but aligning these models with human preferences remains c...
We aim at using Energy-based Model (EBM) framework to better understand adversarial training (AT) in classifiers, and additionally to analyze the in...
Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference spee...
Perceptual voice quality assessment is essential for diagnosing and monitoring voice disorders by providing standardized evaluations of vocal functi...
Classifier-Free Guidance (CFG) is a widely used technique for improving conditional diffusion models by linearly combining the outputs of conditiona...
Temporal Logic (TL), especially Signal Temporal Logic (STL), enables precise formal specification, making it widely used in cyber-physical systems s...
The quality of training data is critical to the performance of machine learning applications in domains like transportation, healthcare, and robotic...
Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has ...
Natural images exhibit label diversity (clean vs. noisy) in noisy-labeled image classification and prevalence diversity (abundant vs. sparse) in lon...
The substantial training cost of diffusion models hinders their deployment. Immiscible Diffusion recently showed that reducing diffusion trajectory ...
As learned image compression (LIC) methods become increasingly computationally demanding, enhancing their training efficiency is crucial. This paper...
Nanomaterial research is becoming a vital area for energy, medicine, and materials science, and accurate analysis of the nanoparticle topology is es...
As the marginal cost of scaling computation (data and parameters) during model pre-training continues to increase substantially, test-time scaling (...
Large-scale Vision Language Models (LVLMs) are increasingly being applied to a wide range of real-world multimodal applications, involving complex v...
BACKGROUND: The English National Health Service (NHS) strives for a fair, diverse, and inclusive workplace, but Black and Minority Ethnic (BME) repres...
We present an efficient and reliable large-scale non-adiabatic dynamics simulation method based on machine learning Hamiltonian and force field. The q...