Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 4961-4980 of 9,853 articles

DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Transformer and Mamba

Recently, non-convolutional models such as the Vision Transformer (ViT) and Vision Mamba (Vim) have achieved remarkable performance in computer vision tasks. However, their reliance on fixed-size patches often results in excessive encoding of background regions and omission of critical local details, especially when informative objects are sparsely distributed. To address this, we introduce a fu...

Revisiting Transformers with Insights from Image Filtering

The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpret. Establishing a robust theoretical foundation to explain its remarkable success and limitations has therefore become an increasingly prominent focus in recent research. Some notable directions have explored understan...

UrbanSense:AFramework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models

Urban cultures and architectural styles vary significantly across cities due to geographical, chronological, historical, and socio-political factors...

Using Vision Language Models to Detect Students' Academic Emotion through Facial Expressions

Students' academic emotions significantly influence their social behavior and learning performance. Traditional approaches to automatically and accu...

HalLoc: Token-level Localization of Hallucinations for Vision Language Models

Hallucinations pose a significant challenge to the reliability of large vision-language models, making their detection essential for ensuring accura...

Energy Aware Camera Location Search Algorithm for Increasing Precision of Observation in Automated Manufacturing

Visual servoing technology has been well developed and applied in many automated manufacturing tasks, especially in tools' pose alignment. To access...

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reaso...

From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models

One promise that Vision-Language-Action (VLA) models hold over traditional imitation learning for robotics is to leverage the broad generalization c...

3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation

Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in the...

Hierarchical Image Matching for UAV Absolute Visual Localization via Semantic and Structural Constraints

Absolute localization, aiming to determine an agent's location with respect to a global reference, is crucial for unmanned aerial vehicles (UAVs) in...

Vision Matters: Simple Visual Perturbations Can Boost Multimodal Math Reasoning

Despite the rapid progress of multimodal large language models (MLLMs), they have largely overlooked the importance of visual processing. In a simpl...

FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textua...

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding

Automated 3D CT diagnosis empowers clinicians to make timely, evidence-based decisions by enhancing diagnostic accuracy and workflow efficiency. Whi...

DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt

Large Vision-Language Models (LVLMs) have achieved impressive progress across various applications but remain vulnerable to malicious queries that e...

Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better

Typical large vision-language models (LVLMs) apply autoregressive supervision solely to textual sequences, without fully incorporating the visual mo...

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models

Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retriev...

MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis

Artificial intelligence (AI) has become a fundamental tool for assisting clinicians in analyzing ophthalmic images, such as optical coherence tomogr...

PlantBert: An Open Source Language Model for Plant Science

The rapid advancement of transformer-based language models has catalyzed breakthroughs in biomedical and clinical natural language processing; howev...

WetCat: Automating Skill Assessment in Wetlab Cataract Surgery Videos

To meet the growing demand for systematic surgical training, wetlab environments have become indispensable platforms for hands-on practice in ophtha...

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval

The parameter-efficient adaptation of the image-text pretraining model CLIP for video-text retrieval is a prominent area of research. While CLIP is ...

Browse Categories