Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6361-6380 of 9,853 articles

High-speed and High-quality Vision Reconstruction of Spike Camera with Spike Stability Theorem

Neuromorphic vision sensors, such as the dynamic vision sensor (DVS) and spike camera, have gained increasing attention in recent years. The spike camera can detect fine textures by mimicking the fovea in the human visual system, and output a high-frequency spike stream. Real-time high-quality vision reconstruction from the spike stream can build a bridge to high-level vision task applications o...

ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy Prediction

Inferring the 3D structure of a scene from a single image is an ill-posed and challenging problem in the field of vision-centric autonomous driving. Existing methods usually employ neural radiance fields to produce voxelized 3D occupancy, lacking instance-level semantic reasoning and temporal photometric consistency. In this paper, we propose ViPOcc, which leverages the visual priors from vision...

Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval

Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a...

Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track

One of the key goals of artificial intelligence (AI) is the development of a multimodal system that facilitates communication with the visual world ...

SightGlow: A Web Extension to Enhance Color Perception and Interaction for Vision Deficiency

SightGlow is a web extension tailored to improve color perception accuracy for individuals with red-green color blindness. The research was focused ...

RapidNet: Multi-Level Dilated Convolution Based Mobile Backbone

Vision transformers (ViTs) have dominated computer vision in recent years. However, ViTs are computationally expensive and not well suited for mobil...

Do large language vision models understand 3D shapes?

Large vision language models (LVLM) are the leading A.I approach for achieving a general visual understanding of the world. Models such as GPT, Clau...

Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels

Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical ta...

Optimizing Vision-Language Interactions Through Decoder-Only Models

Vision-Language Models (VLMs) have emerged as key enablers for multimodal tasks, but their reliance on separate visual encoders introduces challenge...

MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance

The Long Short-Term Memory (LSTM) networks have traditionally faced challenges in scaling and effectively capturing complex dependencies in visual t...

EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing

Editing complex visual content based on ambiguous instructions remains a challenging problem in vision-language modeling. While existing models can ...

XYScanNet: A State Space Model for Single Image Deblurring

Deep state-space models (SSMs), like recent Mamba architectures, are emerging as a promising alternative to CNN and Transformer networks. Existing M...

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

We present DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecesso...

ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?

Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. T...

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their...

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models

The obesity phenomenon, known as the heavy issue, is a leading cause of preventable chronic diseases worldwide. Traditional calorie estimation tools...

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens

Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of ...

Acquisition of Spatially-Varying Reflectance and Surface Normals via Polarized Reflectance Fields

Accurately measuring the geometry and spatially-varying reflectance of real-world objects is a complex task due to their intricate shapes formed by ...

Co-creating Humanistic AI AgeTech to Support Dynamic Care Ecosystems: A Preliminary Guiding Model.

As society rapidly digitizes, successful aging necessitates using technology for health and social care and social engagement. Technologies aimed to s...

Dec 13 2024 39094095
FovealNet: Advancing AI-Driven Gaze Tracking Solutions for Optimized Foveated Rendering System Performance in Virtual Reality

Leveraging real-time eye-tracking, foveated rendering optimizes hardware efficiency and enhances visual quality virtual reality (VR). This approach ...

Browse Categories