Latest AI and machine learning research in ophthalmology for healthcare professionals.
Supervised convolutional neural networks (CNNs) are widely used to solve imaging inverse problems, achieving state-of-the-art performance in numerous applications. However, despite their empirical success, these methods are poorly understood from a theoretical perspective and often treated as black boxes. To bridge this gap, we analyze trained neural networks through the lens of the Minimum Mean S...
Recent achievements of vision-language models in end-to-end OCR point to a new avenue for low-loss compression of textual information. This motivates earlier works that render the Transformer's input into images for prefilling, which effectively reduces the number of tokens through visual encoding, thereby alleviating the quadratically increased Attention computations. However, this partial compre...
Image segmentation plays a central role in computer vision. However, widely used evaluation metrics, whether pixel-wise, region-based, or boundary-foc...
Precise cell type targeting is critical for both clinical and experimental applications of adeno-associated viral (AAV) vectors, yet engineering vecto...
Humans can readily recognize words even when they are misspelled, though with slower responses, demonstrating remarkable robustness in reading. The co...
General-purpose Large Vision-Language Models (LVLMs), despite their massive scale, often falter in dermatology due to "diffuse attention" - the inabil...
Zero-Shot Anomaly Detection (ZSAD) leverages Vision-Language Models (VLMs) to enable supervision-free industrial inspection. However, existing ZSAD pa...
Reconstructing natural visual scenes from neural activity is a key challenge in neuroscience and computer vision. We present SpikeVAEDiff, a novel two...
Despite the remarkable progress of Vision-Language Models (VLMs) in adopting "Thinking-with-Images" capabilities, accurately evaluating the authentici...
Autonomous agents such as cars, robots and drones need to precisely localize themselves in diverse environments, including in GPS-denied indoor enviro...
Object detectors often perform well in-distribution, yet degrade sharply on a different benchmark. We study cross-dataset object detection (CD-OD) thr...
Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation a...
OBJECTIVE: Ulcerative colitis (UC) is a chronic inflammatory bowel disease for which remission is dependent on corticosteroid (CS) treatment. The dive...
Due to its efficiency, Post-Training Quantization (PTQ) has been widely adopted for compressing Vision Transformers (ViTs). However, when quantized in...
INTRODUCTION: Advances in artificial intelligence offer the promise of automated analysis of optical coherence tomography (OCT) scans to detect ocular...
How is visual stability maintained across saccades? One theory poses the visual system has an underlying assumption that the visual world has not chan...
Refractory epilepsy is an intractable neurological disorder that can be associated with oligogenic/polygenic etiologies. Through trio-based whole-exom...
Multi-view camera-based 3D perception can be conducted using bird's eye view (BEV) features obtained through perspective view-to-BEV transformations...
While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of ...
Leveraging the powerful representations of pre-trained vision foundation models -- traditionally used for visual comprehension -- we explore a novel...