Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5601-5620 of 9,853 articles

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-grained editing. To address these issues, we propose FireEdit, an innovative Fine-grained Instruction...

ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

End-to-end (E2E) autonomous driving methods still struggle to make correct decisions in interactive closed-loop evaluation due to limited causal reasoning capability. Current methods attempt to leverage the powerful understanding and reasoning abilities of Vision-Language Models (VLMs) to resolve this dilemma. However, the problem is still open that few VLMs for E2E methods perform well in the c...

Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as i...

RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models

We introduce RGB-Th-Bench, the first benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to comprehend RGB-Thermal image pai...

Improved Alignment of Modalities in Large Vision Language Models

Recent advancements in vision-language models have achieved remarkable results in making language models understand vision inputs. However, a unifie...

GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers

The synergy between generative and discriminative models receives growing attention. While discriminative Contrastive Language-Image Pre-Training (C...

ASP-VMUNet: Atrous Shifted Parallel Vision Mamba U-Net for Skin Lesion Segmentation

Skin lesion segmentation is a critical challenge in computer vision, and it is essential to separate pathological features from healthy skin for dia...

LangBridge: Interpreting Image as a Combination of Language Embeddings

Recent years have witnessed remarkable advances in Large Vision-Language Models (LVLMs), which have achieved human-level performance across various ...

Adaptive Wavelet Filters as Practical Texture Feature Amplifiers for Parkinson's Disease Screening in OCT

Parkinson's disease (PD) is a prevalent neurodegenerative disorder globally. The eye's retina is an extension of the brain and has great potential i...

Multiscale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines

Feature Coding for Machines (FCM) aims to compress intermediate features effectively for remote intelligent analytics, which is crucial for future i...

Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications

Semantic segmentation has made significant strides in pixel-level image understanding, yet it remains limited in capturing contextual and semantic r...

3D Structural Phenotype of the Optic Nerve Head at the Intersection of Glaucoma and Myopia -- A Key to Improving Glaucoma Diagnosis in Myopic Populations

Purpose: To characterize the 3D structural phenotypes of the optic nerve head (ONH) in patients with glaucoma, high myopia, and concurrent high myop...

Improving Food Image Recognition with Noisy Vision Transformer

Food image recognition is a challenging task in computer vision due to the high variability and complexity of food images. In this study, we investi...

MC-LLaVA: Multi-Concept Personalized Vision-Language Model

Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience...

Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and ap...

ArchSeek: Retrieving Architectural Case Studies Using Vision-Language Models

Efficiently searching for relevant case studies is critical in architectural design, as designers rely on precedent examples to guide or inspire the...

Rethinking Glaucoma Calibration: Voting-Based Binocular and Metadata Integration

Glaucoma is an incurable ophthalmic disease that damages the optic nerve, leads to vision loss, and ranks among the leading causes of blindness worl...

OCCO: LVM-guided Infrared and Visible Image Fusion Framework based on Object-aware and Contextual COntrastive Learning

Image fusion is a crucial technique in the field of computer vision, and its goal is to generate high-quality fused images and improve the performan...

Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models

Despite the significant success of Large Vision-Language models(LVLMs), these models still suffer hallucinations when describing images, generating ...

PM4Bench: A Parallel Multilingual Multi-Modal Multi-task Benchmark for Large Vision Language Model

Existing multilingual benchmarks for Large Vision Language Models (LVLMs) suffer from limitations including language-specific content biases, disjoi...

Browse Categories