Latest AI and machine learning research in ophthalmology for healthcare professionals.
Genomic foundation models capture sequence regularities, yet existing interpretability tools rarely ask where in the layer stack a biological grammar becomes stable. We introduce GDTR, the Genomic Deep-Thinking Ratio, a training-free residual-stream lens that assigns each nucleotide token a settling depth : the first layer at which its representation stabilizes against the post-final-norm referenc...
High-density probes record from thousands of neurons simultaneously, yet resolving single-neuron identity remains an ill-posed inverse problem. While detailed simulations precisely characterize the biophysical forward process, their utility for interpreting brain signal remains unclear. Here we show that biophysical simulations of population neuronal electrical signals serve as an effective bridge...
Objective: To develop a low-cost automated cataract severity classification system operating on standard consumer-grade colour photographs of the eye,...
Volumetric segmentation of optical coherence tomography (OCT) images is essential for diagnosing ocular diseases but requires labor-intensive voxel-wi...
Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the geometry separates it. Nearest-neighbor...
Background: Retinal fundus imaging is central to the early diagnosis of sight-threatening conditions including diabetic retinopathy, glaucoma, and ret...
Volumetric segmentation of optical coherence tomography (OCT) images is essential for diagnosing ocular diseases but requires labor-intensive voxel-wi...
End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, b...
Vision-only occupancy prediction requires recovering a semantic 3D occupancy field from calibrated surround-view images, where each view provides obse...
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image m...
Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial q...
Deep brain stimulation (DBS) is effective for treatment-refractory obsessive-compulsive disorder (OCD), but outcomes are heterogeneous and non-respond...
Physical simulations that predict the behavior of urban disasters, such as climate-related flooding, play a crucial role in disaster prevention and th...
Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked shoul...
Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models a...
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and r...
A surveillance camera is an image sensor whose silent physical degradation invalidates every downstream consumer of its data. In-situ integrity alarms...
With the proliferation of immersive Head-Mounted Displays (HMDs) for Virtual and Augmented Reality (VR/AR), reliable and high-precision eye tracking h...
Most single-cell foundation models are adapted from language models, representing each cell as a sequence of gene tokens. This discards the relationsh...
In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike p...