Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5481-5500 of 9,853 articles

FastRSR: Efficient and Accurate Road Surface Reconstruction from Bird's Eye View

Road Surface Reconstruction (RSR) is crucial for autonomous driving, enabling the understanding of road surface conditions. Recently, RSR from the Bird's Eye View (BEV) has gained attention for its potential to enhance performance. However, existing methods for transforming perspective views to BEV face challenges such as information loss and representation sparsity. Moreover, stereo matching in...

Enhancing Wide-Angle Image Using Narrow-Angle View of the Same Scene

A common dilemma while photographing a scene is whether to capture it at a wider angle, allowing more of the scene to be covered but in less detail or to click in a narrow angle that captures better details but leaves out portions of the scene. We propose a novel method in this paper that infuses wider shots with finer quality details that is usually associated with an image captured by the prim...

Identity-Aware Vision-Language Model for Explainable Face Forgery Detection

Recent advances in generative artificial intelligence have enabled the creation of highly realistic image forgeries, raising significant concerns ab...

Visual moral inference and communication

Humans can make moral inferences from multiple sources of input. In contrast, automated moral inference in artificial intelligence typically relies ...

Breaking the Lens of the Telescope: Online Relevance Estimation over Large Retrieval Sets

Advanced relevance models, such as those that use large language models (LLMs), provide highly accurate relevance estimations. However, their comput...

Evolved Hierarchical Masking for Self-Supervised Learning

Existing Masked Image Modeling methods apply fixed mask patterns to guide the self-supervised training. As those mask patterns resort to different c...

Multi-modal and Multi-view Fundus Image Fusion for Retinopathy Diagnosis via Multi-scale Cross-attention and Shifted Window Self-attention

The joint interpretation of multi-modal and multi-view fundus images is critical for retinopathy prevention, as different views can show the complet...

Using Vision Language Models for Safety Hazard Identification in Construction

Safety hazard identification and prevention are the key elements of proactive safety management. Previous research has extensively explored the appl...

Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict

Vision-language models (VLMs) have demonstrated impressive performance by effectively integrating visual and textual information to solve complex ta...

LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping

Anamorphosis refers to a category of images that are intentionally distorted, making them unrecognizable when viewed directly. Their true form only ...

Steering CLIP's vision transformer with sparse autoencoders

While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have help...

Hypergraph Vision Transformers: Images are More than Nodes, More than Edges

Recent advancements in computer vision have highlighted the scalability of Vision Transformers (ViTs) across various tasks, yet challenges remain in...

Hardware, Algorithms, and Applications of the Neuromorphic Vision Sensor: a Review

Neuromorphic, or event, cameras represent a transformation in the classical approach to visual sensing encodes detected instantaneous per-pixel illu...

FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations

Visual understanding is inherently contextual -- what we focus on in an image depends on the task at hand. For instance, given an image of a person ...

VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering

The increasing availability of multimodal data across text, tables, and images presents new challenges for developing models capable of complex cros...

EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models

Vision models are increasingly deployed in critical applications such as autonomous driving and CCTV monitoring, yet they remain susceptible to reso...

[Advancements in machine learning applications in refractive surgery].

Refractive error is a significant factor contributing to visual impairment, imposing a relatively large burden on the social economy. Although refract...

Apr 11 2025 40189889
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs)...

TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs

While text-to-image (T2I) generation models have achieved remarkable progress in recent years, existing evaluation methodologies for vision-language...

Kimi-VL Technical Report

We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-co...

Browse Categories