j-IR-vis: Vision model for infrared spectroscopy embeddings.
Journal:
The Journal of chemical physics
Published Date:
Sep 7, 2026
Abstract
Infrared (IR) spectroscopy provides rich structural information, but interpreting spectra at scale remains challenging. Here, we introduce j-IR-vis, a vision model trained on IR spectra for functional-group prediction and downstream molecular characterization. Trained separately on simulated (Dsim) and experimental (Dexp) datasets, j-IR-vis achieves strong functional group classification accuracy, with exact-match ratios of 0.77-0.81 on Dsim and 0.67-0.76 on Dexp, and macro-F1 scores up to 0.93. Grad-CAM saliency maps reveal that the model tends to focus on the fingerprint region (1500-400 cm-1), consistent with the model being sensitive to vibrational features in this region. Embedding-space analyses show that the inclusion of this region improves all geometric metrics-RankMe, Silhouette, Separation, and Lift-indicating a richer and more structured latent manifold. Fixed embeddings further enable accurate prediction of aromatic ring counts (weighted F1 = 0.69) and octanol-water partition coefficients (RMSE = 1.18, Spearman's ρ = 0.59), demonstrating representation reuse for structural and physicochemical tasks without retraining the vision backbone. Finally, cosine-similarity analyses reveal that j-IR-vis captures complementary chemical relationships compared to traditional fingerprint-based representations. Together, these results establish j-IR-vis as a reusable spectral encoder that bridges experimental spectroscopy and molecular machine learning, offering a spectra informed route to multi-modal chemical representations. j-IR-vis is openly accessible through the GitHub repository https://github.com/ChemAI-Lab/jirvis.
Authors
Keywords
No keywords available for this article.