Ophthalmology

Latest AI and machine learning research in ophthalmology for healthcare professionals.

9,853 articles
Stay Ahead - Weekly Ophthalmology research updates
Subscribe
Browse Categories
Showing 5161-5180 of 9,853 articles

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM

Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been developed to generate not only text but also visual outputs from multimodal inputs. This advancement requires efficient image tokens that LLMs can process effectively both in input and output. However, existing image tokeni...

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM

Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been developed to generate not only text but also visual outputs from multimodal inputs. This advancement requires efficient image tokens that LLMs can process effectively both in input and output. However, existing image tokeni...

Automating Versatile Time-Series Analysis with Tiny Transformers on Embedded FPGAs

Transformer-based models have shown strong performance across diverse time-series tasks, but their deployment on resource-constrained devices remain...

Automating Versatile Time-Series Analysis with Tiny Transformers on Embedded FPGAs

Transformer-based models have shown strong performance across diverse time-series tasks, but their deployment on resource-constrained devices remain...

Decoupled Visual Interpretation and Linguistic Reasoning for Math Problem Solving

Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language mode...

Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding

Recent advancements in Large Vision-Language Models (LVLMs) have significantly expanded their utility in tasks like image captioning and visual ques...

The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts

The detection and grounding of multimedia manipulation has emerged as a critical challenge in combating AI-generated disinformation. While existing ...

Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies

Large-scale Vision Language Models (LVLMs) are increasingly being applied to a wide range of real-world multimodal applications, involving complex v...

VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding and generation, yet their vulnerability t...

Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning

To advance biomedical vison-language model capabilities through scaling up, fine-tuning, and instruction tuning, develop vision-language models with...

A deep learning model integrating domain-specific features for enhanced glaucoma diagnosis.

Glaucoma is a group of serious eye diseases that can cause incurable blindness. Despite the critical need for early detection, over 60% of cases remai...

May 23 2025 40410768
Predicting and interpreting key features of refractory Mycoplasma pneumoniae pneumonia using multiple machine learning methods.

In recent years, the incidence of refractory Mycoplasma pneumoniae pneumonia (RMPP) has significantly risen, posing severe pulmonary and extrapulmonar...

May 23 2025 40410245
Seeing through arthropod eyes: An AI-assisted, biomimetic approach for high-resolution, multi-task imaging.

Arthropods have intricate compound eyes and optic neuropils, exhibiting exceptional visual capabilities. Combining the strengths of digital imaging wi...

May 23 2025 40397741
Ocular Authentication: Fusion of Gaze and Periocular Modalities

This paper investigates the feasibility of fusing two eye-centric authentication modalities-eye movements and periocular images-within a calibration...

Optimizing Image Capture for Computer Vision-Powered Taxonomic Identification and Trait Recognition of Biodiversity Specimens

Biological collections house millions of specimens documenting Earth's biodiversity, with digital images increasingly available through open-access ...

Robustifying Vision-Language Models via Dynamic Token Reweighting

Large vision-language models (VLMs) are highly vulnerable to jailbreak attacks that exploit visual-textual interactions to bypass safety guardrails....

Let Androids Dream of Electric Sheep: A Human-like Image Implication Understanding and Reasoning Framework

Metaphorical comprehension in images remains a critical challenge for AI systems, as existing models struggle to grasp the nuanced cultural, emotion...

Action2Dialogue: Generating Character-Centric Narratives from Scene-Level Prompts

Recent advances in scene-based video generation have enabled systems to synthesize coherent visual narratives from structured prompts. However, a cr...

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models p...

A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical NLP

We present a Japanese domain-specific language model for the pharmaceutical field, developed through continual pretraining on 2 billion Japanese pha...

Browse Categories