Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6081-6100 of 9,853 articles

LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) res...

MedFocusCLIP : Improving few shot classification in medical datasets using pixel wise attention

With the popularity of foundational models, parameter efficient fine tuning has become the defacto approach to leverage pretrained models to perform downstream tasks. Taking inspiration from recent advances in large language models, Visual Prompt Tuning, and similar techniques, learn an additional prompt to efficiently finetune a pretrained vision foundational model. However, we observe that suc...

KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration

Zero-shot anomaly detection (ZSAD) identifies anomalies without needing training samples from the target dataset, essential for scenarios with priva...

Materialist: Physically Based Editing Using Single-Image Inverse Rendering

To perform image editing based on single-view, inverse physically based rendering, we present a method combining a learning-based approach with prog...

Exploring EEG and Eye Movement Fusion for Multi-Class Target RSVP-BCI

Rapid Serial Visual Presentation (RSVP)-based Brain-Computer Interfaces (BCIs) facilitate high-throughput target image detection by identifying even...

DGSSA: Domain generalization with structural and stylistic augmentation for retinal vessel segmentation

Retinal vascular morphology is crucial for diagnosing diseases such as diabetes, glaucoma, and hypertension, making accurate segmentation of retinal...

Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-ray Report Generation

X-ray image based medical report generation achieves significant progress in recent years with the help of the large language model, however, these ...

A Bio-Inspired Research Paradigm of Collision Perception Neurons Enabling Neuro-Robotic Integration: The LGMD Case

Compared to human vision, locust visual systems excel at rapid and precise collision detection, despite relying on only hundreds of thousands of neu...

Human Gaze Boosts Object-Centered Representation Learning

Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks co...

COph100: A comprehensive fundus image registration dataset from infants constituting the "RIDIRP" database

Retinal image registration is vital for diagnostic therapeutic applications within the field of ophthalmology. Existing public datasets, focusing on...

Visual Large Language Models for Generalized and Specialized Applications

Visual-language models (VLM) have emerged as a powerful tool for learning a unified embedding space for vision and language. Inspired by large langu...

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Large language models and vision transformers have demonstrated impressive zero-shot capabilities, enabling significant transferability in downstrea...

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

While Vision Language Models (VLMs) are impressive in tasks such as visual question answering (VQA) and image captioning, their ability to apply mul...

Identifying Surgical Instruments in Pedagogical Cataract Surgery Videos through an Optimized Aggregation Network

Instructional cataract surgery videos are crucial for ophthalmologists and trainees to observe surgical details repeatedly. This paper presents a de...

Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks

Vision generation remains a challenging frontier in artificial intelligence, requiring seamless integration of visual understanding and generative c...

DeTrack: In-model Latent Denoising Learning for Visual Object Tracking

Previous visual object tracking methods employ image-feature regression models or coordinate autoregression models for bounding box prediction. Imag...

Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers

We present an approach to modifying Transformer architectures by integrating graph-aware relational reasoning into the attention mechanism, merging ...

Guiding Medical Vision-Language Models with Explicit Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations

While mainstream vision-language models (VLMs) have advanced rapidly in understanding image level information, they still lack the ability to focus ...

Diabetic Retinopathy Detection Using CNN with Residual Block with DCGAN

Diabetic Retinopathy (DR) is a major cause of blindness worldwide, caused by damage to the blood vessels in the retina due to diabetes. Early detect...

Simulated prosthetic vision confirms checkerboard as an effective raster pattern for epiretinal implants

Spatial scheduling of electrode activation ("rastering") is essential for safely operating high-density retinal implants, yet its perceptual consequ...

Browse Categories