Latest AI and machine learning research in ophthalmology for healthcare professionals.
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared embedding space with the LLM while preserving semantic similarity. Existing connectors, such as multilayer perceptrons (MLPs), often produce out-of-distribution o...
Large language models (LLMs) have shown significant promise across various medical applications, with ophthalmology being a notable area of focus. Many ophthalmic tasks have shown substantial improvement through the integration of LLMs. However, before these models can be widely adopted in clinical practice, evaluating their capabilities and identifying their limitations is crucial. To address t...
Lensless cameras disregard the conventional design that imaging should mimic the human eye. This is done by replacing the lens with a thin mask, and...
The integration of language instructions with robotic control, particularly through Vision Language Action (VLA) models, has shown significant poten...
The ability of tumors to evolve and adapt by developing subclones in different genetic and epigenetic states is a major challenge in oncology. Tradi...
PURPOSE: The purpose of this study was to develop an artificial intelligence (AI)-based intraocular lens (IOLs) power calculation formula for improvin...
PURPOSE: To evaluate the performance of various approaches of processing three-dimensional (3D) optical coherence tomography (OCT) images for deep lea...
PURPOSE: This study aims to develop an automated pipeline to detect retinal detachment from B-scan ocular ultrasonography (USG) images by using deep l...
Existing methods for analyzing linguistic content from picture descriptions for assessment of cognitive-linguistic impairment often overlook the par...
Vision-language navigation in unknown environments is crucial for mobile robots. In scenarios such as household assistance and rescue, mobile robots...
Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficie...
Real-world applications are stretching context windows to hundreds of thousand of tokens while Large Language Models (LLMs) swell from billions to t...
Vision classifiers are often trained on proprietary datasets containing sensitive information, yet the models themselves are frequently shared openl...
In this 2022 work we argued that, despite claims about successful modeling of the visual brain using artificial nets, the problem is far from being ...
Segment Anything Model (SAM) represents a large-scale segmentation model that enables powerful zero-shot capabilities with flexible prompts. While S...
Hallucination has been a long-standing and inevitable problem that hinders the application of Large Vision-Language Models (LVLMs) in domains that r...
Visual reasoning refers to the task of solving questions about visual information. Current visual reasoning methods typically employ pre-trained vis...
State Space Models (SSMs) with selective scan (Mamba) have been adapted into efficient vision models. Mamba, unlike Vision Transformers, achieves li...
Although backpropagation is widely accepted as a training algorithm for artificial neural networks, researchers are always looking for inspiration f...
Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potenti...