Latest AI and machine learning research in ophthalmology for healthcare professionals.
Multimodal foundation models, such as GPT-4o, have recently made remarkable progress, but it is not clear where exactly these models stand in terms of understanding vision. In this paper, we benchmark the performance of popular multimodal foundation models (GPT-4o, o4-mini, Gemini 1.5 Pro and Gemini 2.0 Flash, Claude 3.5 Sonnet, Qwen2-VL, Llama 3.2) on standard computer vision tasks (semantic se...
Biological function arises through the dynamical interactions of multiple subsystems, including those between brain areas, within gene regulatory networks, and more. A common approach to understanding these systems is to model the dynamics of each subsystem and characterize communication between them. An alternative approach is through the lens of control theory: how the subsystems control one a...
Deep neural networks have achieved remarkable results in computer vision tasks. In the early days, Convolutional Neural Networks (CNNs) were the mai...
Pruning is widely used to reduce the complexity of deep learning models, but its effects on interpretability and representation learning remain poor...
Emerging low-altitude economy networks (LAENets) require agile and privacy-preserving resource control under dynamic agent mobility and limited infr...
We introduce Generalized Test-Time Augmentation (GTTA), a highly effective method for improving the performance of a trained model, which unlike oth...
Large language models (LLMs) can simulate clinical reasoning based on natural language prompts, but their utility in ophthalmology is largely unexpl...
Traditional simulator-based training for maritime professionals is critical for ensuring safety at sea but often depends on subjective trainer asses...
In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks durin...
Early detection and diagnosis of diabetic retinopathy is one of the current research focuses in ophthalmology. However, due to the subtle features o...
Ophthalmic surgical robots offer superior stability and precision by reducing the natural hand tremors of human surgeons, enabling delicate operatio...
Vision State Space Models (SSMs), particularly architectures like Vision Mamba (ViM), have emerged as promising alternatives to Vision Transformers ...
The architecture of multimodal large language models (MLLMs) commonly connects a vision encoder, often based on CLIP-ViT, to a large language model....
Humans are able to recognize objects based on both local texture cues and the configuration of object parts, yet contemporary vision models primaril...
Just noticeable difference (JND), the minimum change that the human visual system (HVS) can perceive, has been studied for decades. Although recent ...
Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and actin...
PURPOSE: The purpose of this study was to investigate the potential of a novel anatomical metric of ametropia-fundus refraction offset (FRO)-in strati...
The classification and analysis of coal are crucial for energy production and resource management. Shadowgraphy, leveraging variations in air refracti...
Sensitive and cost-effective detection methods utilizing portable equipment are crucial for applications in food safety inspection, environmental moni...
BACKGROUND: Grading fluorescein angiography (FA) for uveitis is complex, often leading to the oversight of retinal inflammation in clinical studies. T...