Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5101-5120 of 9,853 articles

Adversarially Robust AI-Generated Image Detection for Free: An Information Theoretic Perspective

Rapid advances in Artificial Intelligence Generated Images (AIGI) have facilitated malicious use, such as forgery and misinformation. Therefore, numerous methods have been proposed to detect fake images. Although such detectors have been proven to be universally vulnerable to adversarial attacks, defenses in this field are scarce. In this paper, we first identify that adversarial training (AT), ...

One Rank at a Time: Cascading Error Dynamics in Sequential Learning

Sequential learning -- where complex tasks are broken down into simpler, hierarchical components -- has emerged as a paradigm in AI. This paper views sequential learning through the lens of low-rank linear regression, focusing specifically on how errors propagate when learning rank-1 subspaces sequentially. We present an analysis framework that decomposes the learning process into a series of ra...

Thinking with Generated Images

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reas...

Understanding Adversarial Training with Energy-based Models

We aim at using Energy-based Model (EBM) framework to better understand adversarial training (AT) in classifiers, and additionally to analyze the in...

Zero-Shot 3D Visual Grounding from Vision-Language Models

3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as...

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

Recent advancements in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, particularly in...

S2AFormer: Strip Self-Attention for Efficient Vision Transformer

Vision Transformer (ViT) has made significant advancements in computer vision, thanks to its token mixer's sophisticated ability to capture global d...

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Effectively retrieving, reasoning and understanding visually rich information remains a challenge for RAG methods. Traditional text-based methods ca...

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Effectively retrieving, reasoning and understanding visually rich information remains a challenge for RAG methods. Traditional text-based methods ca...

Broadening Our View: Assistive Technology for Cerebral Visual Impairment

Over the past decade, considerable research has been directed towards assistive technologies to support people with vision impairments using machine...

MMTBENCH: A Unified Benchmark for Complex Multimodal Table Reasoning

Multimodal tables those that integrate semi structured data with visual elements such as charts and maps are ubiquitous across real world domains, y...

Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI

From uncertainty quantification to real-world object detection, we recognize the importance of machine learning algorithms, particularly in safety-c...

Are Language Models Consequentialist or Deontological Moral Reasoners?

As AI systems increasingly navigate applications in healthcare, law, and governance, understanding how they handle ethically complex scenarios becom...

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration

Large vision-language models (LVLMs) achieve impressive performance on multimodal tasks but often suffer from hallucination, and confidently describ...

Visual Product Graph: Bridging Visual Products And Composite Images For End-to-End Style Recommendations

Retrieving semantically similar but visually distinct contents has been a critical capability in visual search systems. In this work, we aim to tack...

Dynamic Vision from EEG Brain Recordings: How much does EEG know?

Reconstructing and understanding dynamic visual information (video) from brain EEG recordings is challenging due to the non-stationary nature of EEG...

DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response

Large vision-language models (VLMs) have made great achievements in Earth vision. However, complex disaster scenes with diverse disaster types, geog...

RefAV: Towards Planning-Centric Scenario Mining

Autonomous Vehicles (AVs) collect and pseudo-label terabytes of multi-modal data localized to HD maps during normal fleet testing. However, identify...

The Role of AI in Early Detection of Life-Threatening Diseases: A Retinal Imaging Perspective

Retinal imaging has emerged as a powerful, non-invasive modality for detecting and quantifying biomarkers of systemic diseases-ranging from diabetes...

Not All Thats Rare Is Lost: Causal Paths to Rare Concept Synthesis

Diffusion models have shown strong capabilities in high-fidelity image generation but often falter when synthesizing rare concepts, i.e., prompts th...

Browse Categories