Latest AI and machine learning research in ophthalmology for healthcare professionals.
Plant diseases significantly impact our food supply, causing problems for farmers, economies reliant on agriculture, and global food security. Accurate and timely plant disease diagnosis is crucial for effective treatment and minimizing yield losses. Despite advancements in agricultural technology, a precise and early diagnosis remains a challenge, especially in underdeveloped regions where agri...
Remote Sensing Vision-Language Models (RS VLMs) have made much progress in the tasks of remote sensing (RS) image comprehension. While performing well in multi-modal reasoning and multi-turn conversations, the existing models lack pixel-level understanding and struggle with multi-image inputs. In this work, we propose RSUniVLM, a unified, end-to-end RS VLM designed for comprehensive vision under...
Camera-based photoplethysmography (PPG) obtained from smartphones has shown great promise for personalized healthcare and secure authentication. Thi...
Current image generation models can effortlessly produce high-quality, highly realistic images, but this also increases the risk of misuse. In vario...
We introduce OCULAR, an innovative hardware and software solution for three-dimensional dynamic image analysis of fine particles. Current state-of-t...
Object detection is a fundamental task in computer vision and image understanding, with the goal of identifying and localizing objects of interest w...
Transformers, a groundbreaking architecture proposed for Natural Language Processing (NLP), have also achieved remarkable success in Computer Vision...
Superposition or Neuron Polysemanticity are important concepts in the field of interpretability and one might say they are these most intricately be...
How well are unimodal vision and language models aligned? Although prior work have approached answering this question, their assessment methods do n...
Metalens is an emerging optical system with an irreplaceable merit in that it can be manufactured in ultra-thin and compact sizes, which shows great...
LMMs have shown impressive visual understanding capabilities, with the potential to be applied in agents, which demand strong reasoning and planning...
Recent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them much longer than text...
We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2, a gene...
Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation le...
Applying pseudo labeling techniques has been found to be advantageous in semi-supervised 3D object detection (SSOD) in Bird's-Eye-View (BEV) for aut...
We present Liquid, an auto-regressive generation paradigm that seamlessly integrates visual comprehension and generation by tokenizing images into d...
Image Compression for Machines (ICM) aims to compress images for machine vision tasks rather than human viewing. Current works predominantly concent...
Machine learning-based embedded systems employed in safety-critical applications such as aerospace and autonomous driving need to be robust against ...
Glaucoma is a progressive optic neuropathy characterized by structural damage to the optic nerve head and functional changes in the visual field. De...
CLIP has shown impressive results in aligning images and texts at scale. However, its ability to capture detailed visual features remains limited be...