Latest AI and machine learning research in ophthalmology for healthcare professionals.
Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been developed to generate not only text but also visual outputs from multimodal inputs. This advancement requires efficient image tokens that LLMs can process effectively both in input and output. However, existing image tokeni...
Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been developed to generate not only text but also visual outputs from multimodal inputs. This advancement requires efficient image tokens that LLMs can process effectively both in input and output. However, existing image tokeni...
Transformer-based models have shown strong performance across diverse time-series tasks, but their deployment on resource-constrained devices remain...
Transformer-based models have shown strong performance across diverse time-series tasks, but their deployment on resource-constrained devices remain...
Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language mode...
Recent advancements in Large Vision-Language Models (LVLMs) have significantly expanded their utility in tasks like image captioning and visual ques...
The detection and grounding of multimedia manipulation has emerged as a critical challenge in combating AI-generated disinformation. While existing ...
Large-scale Vision Language Models (LVLMs) are increasingly being applied to a wide range of real-world multimodal applications, involving complex v...
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding and generation, yet their vulnerability t...
To advance biomedical vison-language model capabilities through scaling up, fine-tuning, and instruction tuning, develop vision-language models with...
Glaucoma is a group of serious eye diseases that can cause incurable blindness. Despite the critical need for early detection, over 60% of cases remai...
In recent years, the incidence of refractory Mycoplasma pneumoniae pneumonia (RMPP) has significantly risen, posing severe pulmonary and extrapulmonar...
Arthropods have intricate compound eyes and optic neuropils, exhibiting exceptional visual capabilities. Combining the strengths of digital imaging wi...
This paper investigates the feasibility of fusing two eye-centric authentication modalities-eye movements and periocular images-within a calibration...
Biological collections house millions of specimens documenting Earth's biodiversity, with digital images increasingly available through open-access ...
Large vision-language models (VLMs) are highly vulnerable to jailbreak attacks that exploit visual-textual interactions to bypass safety guardrails....
Metaphorical comprehension in images remains a critical challenge for AI systems, as existing models struggle to grasp the nuanced cultural, emotion...
Recent advances in scene-based video generation have enabled systems to synthesize coherent visual narratives from structured prompts. However, a cr...
Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models p...
We present a Japanese domain-specific language model for the pharmaceutical field, developed through continual pretraining on 2 billion Japanese pha...