Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 6341-6360 of 9,853 articles

Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation

Direct Preference Optimization (DPO) has been demonstrated to be highly effective in mitigating hallucinations in Large Vision Language Models (LVLMs) by aligning their outputs more closely with human preferences. Despite the recent progress, existing methods suffer from two drawbacks: 1) Lack of scalable token-level rewards; and 2) Neglect of visual-anchored tokens. To this end, we propose a no...

VISA: Retrieval Augmented Generation with Visual Source Attribution

Generation with source attribution is important for enhancing the verifiability of retrieval-augmented generation (RAG) systems. However, existing approaches in RAG primarily link generated content to document-level references, making it challenging for users to locate evidence among multiple content-rich retrieved documents. To address this challenge, we propose Retrieval-Augmented Generation w...

TopView: Vectorising road users in a bird's eye view from uncalibrated street-level imagery with deep learning

Generating a bird's eye view of road users is beneficial for a variety of applications, including navigation, detecting agent conflicts, and measuri...

A Unifying Information-theoretic Perspective on Evaluating Generative Models

Considering the difficulty of interpreting generative model output, there is significant current research focused on determining meaningful evaluati...

Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs

Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language mo...

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

In this work, we propose Visual-Predictive Instruction Tuning (VPiT) - a simple and effective extension to visual instruction tuning that enables a ...

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation

The visual understanding are often approached from 3 granular levels: image, patch and pixel. Visual Tokenization, trained by self-supervised recons...

A Black-Box Evaluation Framework for Semantic Robustness in Bird's Eye View Detection

Camera-based Bird's Eye View (BEV) perception models receive increasing attention for their crucial role in autonomous driving, a domain where conce...

Physics-Based Adversarial Attack on Near-Infrared Human Detector for Nighttime Surveillance Camera Systems

Many surveillance cameras switch between daytime and nighttime modes based on illuminance levels. During the day, the camera records ordinary RGB im...

Transmit What You Need: Task-Adaptive Semantic Communications for Visual Information

Recently, semantic communications have drawn great attention as the groundbreaking concept surpasses the limited capacity of Shannon's theory. Speci...

Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval

Existing cross-modal retrieval methods typically rely on large-scale vision-language pair data. This makes it challenging to efficiently develop a c...

Bringing Multimodality to Amazon Visual Search System

Image to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning mod...

FastVLM: Efficient Vision Encoding for Vision Language Models

Scaling the input image resolution is essential for enhancing the performance of Vision Language Models (VLMs), particularly in text-rich image unde...

Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Recent works on accelerating Vision-Language Models show that strong performance can be maintained across a variety of vision-language tasks despite...

A Knowledge-enhanced Pathology Vision-language Foundation Model for Cancer Diagnosis

Deep learning has enabled the development of highly robust foundation models for various pathological tasks across diverse diseases and patient coho...

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporat...

ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance Segmentation

Instance segmentation algorithms in remote sensing are typically based on conventional methods, limiting their application to seen scenarios and clo...

Rethinking Diffusion-Based Image Generators for Fundus Fluorescein Angiography Synthesis on Limited Data

Fundus imaging is a critical tool in ophthalmology, with different imaging modalities offering unique advantages. For instance, fundus fluorescein a...

Accelerating lensed quasar discovery and modeling with physics-informed variational autoencoders

Strongly lensed quasars provide valuable insights into the rate of cosmic expansion, the distribution of dark matter in foreground deflectors, and t...

Multimodal LLM for Intelligent Transportation Systems

In the evolving landscape of transportation systems, integrating Large Language Models (LLMs) offers a promising frontier for advancing intelligent ...

Browse Categories