Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5941-5960 of 9,853 articles

HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robotics. One fundamental challenge is the lack of robotic data, which are typically obtained through expensive on-robot operation. A promising remedy is to leverage cheaper, off-domain data such as action-free videos, han...

Stochastic Forward-Backward Deconvolution: Training Diffusion Models with Finite Noisy Datasets

Recent diffusion-based generative models achieve remarkable results by training on massive datasets, yet this practice raises concerns about memorization and copyright infringement. A proposed remedy is to train exclusively on noisy data with potential copyright issues, ensuring the model never observes original content. However, through the lens of deconvolution theory, we show that although it...

Vision-in-the-loop Simulation for Deep Monocular Pose Estimation of UAV in Ocean Environment

This paper proposes a vision-in-the-loop simulation environment for deep monocular pose estimation of a UAV operating in an ocean environment. Recen...

DCFormer: Efficient 3D Vision-Language Modeling with Decomposed Convolutions

Vision-language models (VLMs) have been widely applied to 2D medical image analysis due to their ability to align visual and textual representations...

Cached Multi-Lora Composition for Multi-Concept Image Generation

Low-Rank Adaptation (LoRA) has emerged as a widely adopted technique in text-to-image models, enabling precise rendering of multiple distinct elemen...

Mechanistic Understandings of Representation Vulnerabilities and Engineering Robust Vision Transformers

While transformer-based models dominate NLP and vision applications, their underlying mechanisms to map the input space to the label space semantica...

An Optimized YOLOv5 Based Approach For Real-time Vehicle Detection At Road Intersections Using Fisheye Cameras

Real time vehicle detection is a challenging task for urban traffic surveillance. Increase in urbanization leads to increase in accidents and traffi...

Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting

Recent advancements in time series forecasting have explored augmenting models with text or vision modalities to improve accuracy. While text provid...

Memory-dependent abstractions of stochastic systems through the lens of transfer operators

With the increasing ubiquity of safety-critical autonomous systems operating in uncertain environments, there is a need for mathematical methods for...

Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More

Since the introduction of Vision Transformer (ViT), patchification has long been regarded as a de facto image tokenization approach for plain visual...

A Beam's Eye View to Fluence Maps 3D Network for Ultra Fast VMAT Radiotherapy Planning

Volumetric Modulated Arc Therapy (VMAT) revolutionizes cancer treatment by precisely delivering radiation while sparing healthy tissues. Fluence map...

RadVLM: A Multitask Conversational Vision-Language Model for Radiology

The widespread use of chest X-rays (CXRs), coupled with a shortage of radiologists, has driven growing interest in automated CXR analysis and AI-ass...

Diffusion Instruction Tuning

We introduce Lavender, a simple supervised fine-tuning (SFT) method that boosts the performance of advanced vision-language models (VLMs) by leverag...

Personalization Toolkit: Training Free Personalization of Large Vision Language Models

Large Vision Language Models (LVLMs) have significant potential to provide personalized assistance by adapting to the unique needs and preferences o...

Iterative Refinement and Flexible Iteratively Reweighed Solvers for Linear Inverse Problems with Sparse Solutions

This paper presents a new algorithmic framework for computing sparse solutions to large-scale linear discrete ill-posed problems. The approach is mo...

Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration

Large Vision-Language Models (LVLMs) exhibit impressive multimodal reasoning capabilities but remain highly susceptible to object hallucination, whe...

Rethinking Homogeneity of Vision and Text Tokens in Large Vision-and-Language Models

Large vision-and-language models (LVLMs) typically treat visual and textual embeddings as homogeneous inputs to a large language model (LLM). Howeve...

AquaticCLIP: A Vision-Language Foundation Model for Underwater Scene Analysis

The preservation of aquatic biodiversity is critical in mitigating the effects of climate change. Aquatic scene understanding plays a pivotal role i...

The Shape of Generalization through the Lens of Norm-based Capacity Control

Understanding how the test risk scales with model complexity is a central question in machine learning. Classical theory is challenged by the learni...

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce h...

Browse Categories