Latest AI and machine learning research in ophthalmology for healthcare professionals.
Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-grained editing. To address these issues, we propose FireEdit, an innovative Fine-grained Instruction...
End-to-end (E2E) autonomous driving methods still struggle to make correct decisions in interactive closed-loop evaluation due to limited causal reasoning capability. Current methods attempt to leverage the powerful understanding and reasoning abilities of Vision-Language Models (VLMs) to resolve this dilemma. However, the problem is still open that few VLMs for E2E methods perform well in the c...
Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as i...
We introduce RGB-Th-Bench, the first benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to comprehend RGB-Thermal image pai...
Recent advancements in vision-language models have achieved remarkable results in making language models understand vision inputs. However, a unifie...
The synergy between generative and discriminative models receives growing attention. While discriminative Contrastive Language-Image Pre-Training (C...
Skin lesion segmentation is a critical challenge in computer vision, and it is essential to separate pathological features from healthy skin for dia...
Recent years have witnessed remarkable advances in Large Vision-Language Models (LVLMs), which have achieved human-level performance across various ...
Parkinson's disease (PD) is a prevalent neurodegenerative disorder globally. The eye's retina is an extension of the brain and has great potential i...
Feature Coding for Machines (FCM) aims to compress intermediate features effectively for remote intelligent analytics, which is crucial for future i...
Semantic segmentation has made significant strides in pixel-level image understanding, yet it remains limited in capturing contextual and semantic r...
Purpose: To characterize the 3D structural phenotypes of the optic nerve head (ONH) in patients with glaucoma, high myopia, and concurrent high myop...
Food image recognition is a challenging task in computer vision due to the high variability and complexity of food images. In this study, we investi...
Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience...
Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and ap...
Efficiently searching for relevant case studies is critical in architectural design, as designers rely on precedent examples to guide or inspire the...
Glaucoma is an incurable ophthalmic disease that damages the optic nerve, leads to vision loss, and ranks among the leading causes of blindness worl...
Image fusion is a crucial technique in the field of computer vision, and its goal is to generate high-quality fused images and improve the performan...
Despite the significant success of Large Vision-Language models(LVLMs), these models still suffer hallucinations when describing images, generating ...
Existing multilingual benchmarks for Large Vision Language Models (LVLMs) suffer from limitations including language-specific content biases, disjoi...