Latest AI and machine learning research in ophthalmology for healthcare professionals.
Road Surface Reconstruction (RSR) is crucial for autonomous driving, enabling the understanding of road surface conditions. Recently, RSR from the Bird's Eye View (BEV) has gained attention for its potential to enhance performance. However, existing methods for transforming perspective views to BEV face challenges such as information loss and representation sparsity. Moreover, stereo matching in...
A common dilemma while photographing a scene is whether to capture it at a wider angle, allowing more of the scene to be covered but in less detail or to click in a narrow angle that captures better details but leaves out portions of the scene. We propose a novel method in this paper that infuses wider shots with finer quality details that is usually associated with an image captured by the prim...
Recent advances in generative artificial intelligence have enabled the creation of highly realistic image forgeries, raising significant concerns ab...
Humans can make moral inferences from multiple sources of input. In contrast, automated moral inference in artificial intelligence typically relies ...
Advanced relevance models, such as those that use large language models (LLMs), provide highly accurate relevance estimations. However, their comput...
Existing Masked Image Modeling methods apply fixed mask patterns to guide the self-supervised training. As those mask patterns resort to different c...
The joint interpretation of multi-modal and multi-view fundus images is critical for retinopathy prevention, as different views can show the complet...
Safety hazard identification and prevention are the key elements of proactive safety management. Previous research has extensively explored the appl...
Vision-language models (VLMs) have demonstrated impressive performance by effectively integrating visual and textual information to solve complex ta...
Anamorphosis refers to a category of images that are intentionally distorted, making them unrecognizable when viewed directly. Their true form only ...
While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have help...
Recent advancements in computer vision have highlighted the scalability of Vision Transformers (ViTs) across various tasks, yet challenges remain in...
Neuromorphic, or event, cameras represent a transformation in the classical approach to visual sensing encodes detected instantaneous per-pixel illu...
Visual understanding is inherently contextual -- what we focus on in an image depends on the task at hand. For instance, given an image of a person ...
The increasing availability of multimodal data across text, tables, and images presents new challenges for developing models capable of complex cros...
Vision models are increasingly deployed in critical applications such as autonomous driving and CCTV monitoring, yet they remain susceptible to reso...
Refractive error is a significant factor contributing to visual impairment, imposing a relatively large burden on the social economy. Although refract...
Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs)...
While text-to-image (T2I) generation models have achieved remarkable progress in recent years, existing evaluation methodologies for vision-language...
We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-co...