Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6501-6520 of 9,853 articles

VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration

Vision-Language Models (VLMs) have demonstrated impressive performance across a versatile set of tasks. A key challenge in accelerating VLMs is storing and accessing the large Key-Value (KV) cache that encodes long visual contexts, such as images or videos. While existing KV cache compression methods are effective for Large Language Models (LLMs), directly migrating them to VLMs yields suboptima...

ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives

Multimedia file fragment classification (MFFC) aims to identify file fragment types, e.g., image/video, audio, and text without system metadata. It is of vital importance in multimedia storage and communication. Existing MFFC methods typically treat fragments as 1D byte sequences and emphasize the relations between separate bytes (interbytes) for classification. However, the more informative rel...

Context-Based Visual-Language Place Recognition

In vision-based robot localization and SLAM, Visual Place Recognition (VPR) is essential. This paper addresses the problem of VPR, which involves ac...

Learning Precise, Contact-Rich Manipulation through Uncalibrated Tactile Skins

While visuomotor policy learning has advanced robotic manipulation, precisely executing contact-rich tasks remains challenging due to the limitation...

Recent consumer OLED monitors can be suitable for vision science

Vision science imposes rigorous requirements for the design and execution of psychophysical studies and experiments. These requirements ensure preci...

Vernacularizing Taxonomies of Harm is Essential for Operationalizing Holistic AI Safety

Operationalizing AI ethics and safety principles and frameworks is essential to realizing the potential benefits and mitigating potential harms caus...

Robust Visual Representation Learning with Multi-modal Prior Knowledge for Image Classification Under Distribution Shift

Despite the remarkable success of deep neural networks (DNNs) in computer vision, they fail to remain high-performing when facing distribution shift...

Reducing Hallucinations in Vision-Language Models via Latent Space Steering

Hallucination poses a challenge to the deployment of large vision-language models (LVLMs) in applications. Unlike in large language models (LLMs), h...

Reimagining partial thickness keratoplasty: An eye mountable robot for autonomous big bubble needle insertion

Autonomous surgical robots have demonstrated significant potential to standardize surgical outcomes, driving innovations that enhance safety and con...

Continuous Pupillography: A Case for Visual Health Ecosystem

This article aims to cover pupillography, and its potential use in a number of ophthalmological diagnostic applications in biomedical space. With th...

Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models

Federated prompt learning benefits federated learning with CLIP-like Vision-Language Model's (VLM's) robust representation learning ability through ...

When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries

In an era where societal narratives are increasingly shaped by algorithmic curation, investigating the political neutrality of LLMs is an important ...

EG-SpikeFormer: Eye-Gaze Guided Transformer on Spiking Neural Networks for Medical Image Analysis

Neuromorphic computing has emerged as a promising energy-efficient alternative to traditional artificial intelligence, predominantly utilizing spiki...

Zero-Shot Pupil Segmentation with SAM 2: A Case Study of Over 14 Million Images

We explore the transformative potential of SAM 2, a vision foundation model, in advancing gaze estimation and eye tracking technologies. By signific...

Artificial intelligence techniques in inherited retinal diseases: A review

Inherited retinal diseases (IRDs) are a diverse group of genetic disorders that lead to progressive vision loss and are a major cause of blindness i...

Mapping Hong Kong's Financial Ecosystem: A Network Analysis of the SFC's Licensed Professionals and Institutions

We present the first study of the Public Register of Licensed Persons and Registered Institutions maintained by the Hong Kong Securities and Futures...

Pair-VPR: Place-Aware Pre-training and Contrastive Pair Classification for Visual Place Recognition with Vision Transformers

In this work we propose a novel joint training method for Visual Place Recognition (VPR), which simultaneously learns a global descriptor and a pair...

Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing

Recent developments of vision large language models (LLMs) have seen remarkable progress, yet still encounter challenges towards multimodal generali...

BEVLoc: Cross-View Localization and Matching via Birds-Eye-View Synthesis

Ground to aerial matching is a crucial and challenging task in outdoor robotics, particularly when GPS is absent or unreliable. Structures like buil...

Resolution limit of the eye: how many pixels can we see?

As large engineering efforts go towards improving the resolution of mobile, AR and VR displays, it is important to know the maximum resolution at wh...

Browse Categories