Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5901-5920 of 9,853 articles

Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction

In this paper, we present a self-calibrating framework that jointly optimizes camera parameters, lens distortion and 3D Gaussian representations, enabling accurate and efficient scene reconstruction. In particular, our technique enables high-quality scene reconstruction from Large field-of-view (FOV) imagery taken with wide-angle lenses, allowing the scene to be modeled from a smaller number of ...

Computational techniques enabling the perception of virtual images exclusive to the retinal afterimage

The retinal afterimage is a widely known effect in the human visual system, which has been studied and used in the context of a number of major art movements. Therefore, when considering the general role of computation in the visual arts, this begs the question whether this effect, too, may be induced using partly automated techniques. If so, it may become a computationally controllable ingredie...

Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues

This study investigates whether vision-language models (VLMs) can perform pragmatic inference, focusing on ignorance implicatures, utterances that i...

Vision-Language In-Context Learning Driven Few-Shot Visual Inspection Model

We propose general visual inspection model using Vision-Language Model~(VLM) with few-shot images of non-defective or defective products, along with...

Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification

Explainability is a highly demanded requirement for applications in high-risk areas such as medicine. Vision Transformers have mainly been limited t...

Diffusion Models Through a Global Lens: Are They Culturally Inclusive?

Text-to-image diffusion models have recently enabled the creation of visually compelling, detailed images from textual prompts. However, their abili...

Harnessing Vision Models for Time Series Analysis: A Survey

Time series analysis has witnessed the inspiring development from traditional autoregressive models, deep learning models, to recent Transformers an...

SASVi -- Segment Any Surgical Video

Purpose: Foundation models, trained on multitudes of public datasets, often require additional fine-tuning or re-prompting mechanisms to be applied ...

$\mathsf{CSMAE~}$:~Cataract Surgical Masked Autoencoder (MAE) based Pre-training

Automated analysis of surgical videos is crucial for improving surgical training, workflow optimization, and postoperative assessment. We introduce ...

ClipRover: Zero-shot Vision-Language Exploration and Target Discovery by Mobile Robots

Vision-language navigation (VLN) has emerged as a promising paradigm, enabling mobile robots to perform zero-shot inference and execute tasks withou...

ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification

Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hiera...

SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation

As a powerful all-weather Earth observation tool, synthetic aperture radar (SAR) remote sensing enables critical military reconnaissance, maritime s...

A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision

Omnidirectional image (ODI) data is captured with a field-of-view of 360x180, which is much wider than the pinhole cameras and captures richer surro...

EventEgo3D++: 3D Human Motion Capture from a Head-Mounted Event Camera

Monocular egocentric 3D human motion capture remains a significant challenge, particularly under conditions of low lighting and fast movements, whic...

Vision-Language Models for Edge Networks: A Comprehensive Survey

Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual ques...

NanoVLMs: How small can we go and still make coherent Vision Language Models?

Vision-Language Models (VLMs), such as GPT-4V and Llama 3.2 vision, have garnered significant research attention for their ability to leverage Large...

Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models

Zero-Shot Anomaly Detection (ZSAD) is an emerging AD paradigm. Unlike the traditional unsupervised AD setting that requires a large number of normal...

SensPS: Sensing Personal Space Comfortable Distance between Human-Human Using Multimodal Sensors

Personal space, also known as peripersonal space, is crucial in human social interaction, influencing comfort, communication, and social stress. Est...

MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification

Whole slide pathology image classification presents challenges due to gigapixel image sizes and limited annotation labels, hindering model generaliz...

Extended monocular 3D imaging

3D vision is of paramount importance for numerous applications ranging from machine intelligence to precision metrology. Despite much recent progres...

Browse Categories