Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6381-6400 of 9,853 articles

Agtech Framework for Cranberry-Ripening Analysis Using Vision Foundation Models

Agricultural domains are being transformed by recent advances in AI and computer vision that support quantitative visual evaluation. Using aerial and ground imaging over a time series, we develop a framework for characterizing the ripening process of cranberry crops, a crucial component for precision agriculture tasks such as comparing crop breeds (high-throughput phenotyping) and detecting dise...

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

Vision-Language Models (VLMs) have shown promising capabilities in handling various multimodal tasks, yet they struggle in long-context scenarios, particularly in tasks involving videos, high-resolution images, or lengthy image-text documents. In our work, we first conduct an empirical analysis of the long-context capabilities of VLMs using our augmented long-context multimodal datasets. Our fin...

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Large Vision-Language Models (VLMs) have been extended to understand both images and videos. Visual token compression is leveraged to reduce the con...

MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images

Existing multi-modal learning methods on fundus and OCT images mostly require both modalities to be available and strictly paired for training and t...

Causal Graphical Models for Vision-Language Compositional Understanding

Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language...

Disentanglement and Compositionality of Letter Identity and Letter Position in Variational Auto-Encoder Vision Models

Human readers can accurately count how many letters are in a word (e.g., 7 in ``buffalo''), remove a letter from a given position (e.g., ``bufflo'')...

Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs...

POINTS1.5: Building a Vision-Language Model towards Real World Applications

Vision-language models have made significant strides recently, demonstrating superior performance across a range of tasks, e.g. optical character re...

Embedding and Enriching Explicit Semantics for Visible-Infrared Person Re-Identification

Visible-infrared person re-identification (VIReID) retrieves pedestrian images with the same identity across different modalities. Existing methods ...

LOMA: Language-assisted Semantic Occupancy Network via Triplane Mamba

Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usu...

FILA: Fine-Grained Vision Language Models

Recently, there has been growing interest in the capability of multimodal large language models (MLLMs) to process high-resolution images. A common ...

Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment

The high-performance computing (HPC) landscape is undergoing rapid transformation, with an increasing emphasis on energy-efficient and heterogeneous...

Generate Any Scene: Evaluating and Improving Text-to-Vision Generation with Scene Graph Programming

DALL-E and Sora have gained attention by producing implausible images, such as "astronauts riding a horse in space." Despite the proliferation of te...

GN-FR:Generalizable Neural Radiance Fields for Flare Removal

Flare, an optical phenomenon resulting from unwanted scattering and reflections within a lens system, presents a significant challenge in imaging. T...

Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics

Textured meshes significantly enhance the realism and detail of objects by mapping intricate texture details onto the geometric structure of 3D mode...

TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning

Despite the efficiency of prompt learning in transferring vision-language models (VLMs) to downstream tasks, existing methods mainly learn the promp...

Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions

In recent years, Visual Question Answering (VQA) has made significant strides, particularly with the advent of multimodal models that integrate visi...

A Review of Intelligent Device Fault Diagnosis Technologies Based on Machine Vision

This paper provides a comprehensive review of mechanical equipment fault diagnosis methods, focusing on the advancements brought by Transformer-base...

Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models

Existing Large Vision-Language Models (LVLMs) excel at matching concepts across multi-modal inputs but struggle with compositional concepts and high...

Dense Depth from Event Focal Stack

We propose a method for dense depth estimation from an event stream generated when sweeping the focal plane of the driving lens attached to an event...

Browse Categories