Latest AI and machine learning research in ophthalmology for healthcare professionals.
Medical ultrasonography is an essential imaging technique for examining superficial organs and tissues, including lymph nodes, breast, and thyroid. It employs high-frequency ultrasound waves to generate detailed images of the internal structures of the human body. However, manually contouring regions of interest in these images is a labor-intensive task that demands expertise and often results i...
Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still struggles with fine-grained, region-level understanding, hindering its effectiveness on these dense prediction tasks. We identify two pivotal factors required to address this limitation: semantic coherence and fine-gra...
Rapid advancements in RISC-V hardware development shift the focus from low-level optimizations to higher-level parallelization. Recent RISC-V proces...
Time series classification is a fundamental task in healthcare and industry, yet the development of time series foundation models (TSFMs) remains li...
Neural disorders refer to any condition affecting the nervous system and that influence how individuals perceive and interact with the world. Tradit...
Vision Transformers (ViTs) exhibit superior performance in computer vision tasks but face deployment challenges on resource-constrained devices due ...
The application of visual instruction tuning and other post-training techniques has significantly enhanced the capabilities of Large Language Models...
Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a crit...
Different medical imaging modalities capture diagnostic information at varying spatial resolutions, from coarse global patterns to fine-grained loca...
Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach ...
Reasoning Segmentation (RS) is a multimodal vision-text task that requires segmenting objects based on implicit text queries, demanding both precise...
Counterfactual image generation presents significant challenges, including preserving identity, maintaining perceptual quality, and ensuring faithfu...
Vision encoders are increasingly used in modern applications, from vision-only models to multimodal systems such as vision-language models. Despite ...
Architectural cultures across regions are characterized by stylistic diversity, shaped by historical, social, and technological contexts in addition...
Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract v...
Optical Coherence Tomography (OCT) provides high-resolution, 3D, and non-invasive visualization of retinal layers in vivo, serving as a critical too...
Multimodal Large Language Models (MLLMs) promise advanced vision language capabilities, yet their effectiveness in visually presented mathematics re...
The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologi...
Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual diffe...
Recent generations of language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answe...