- Home
- Research
- Photonic Data Science
- Publications
- Raman spectra comparison: cautions and pitfalls of similarity metrics
Raman spectra comparison: cautions and pitfalls of similarity metrics
in: Spectrochimica Acta Part A-Molecular and Biomolecular Spectroscopy (2026)
Raman spectroscopy detects indirectly vibrations of molecular compounds that can be used as fingerprints of molecules and chemical/biological compounds. This endorses Raman spectroscopy as a powerful tool for molecular identification and characterization, which can be straightforwardly achieved via database indexing, i.e., to compare a measured spectrum with databases to determine the probable compounds contained in the sample under measurement. Therein the comparison between two spectra plays an essential role, for which various similarity metrics are being commonly used, including Pearson's correlation coefficients, Euclidean distances, cosine distances, etc. However, spectral comparison is not as straightforward as it sounds, as real-world spectral measurement is often accompanied by artefacts and/or variability arising from instruments, samples, and experimental conditions that can distort spectral features and hence compromise the reliability of similarity assessments. It is important to understand the robustness and reliability of different similarity metrics encountering spectral artefacts and variations. In this study, we investigate different spectral similarity metrics to characterize their behavior against different types and strengths of artefacts. We generated a synthetic Raman spectral dataset to simulate real-world variability. Specifically, we added commonly observed spectral artefacts into the synthetic data, including Gaussian noise, baseline profiles, and wavenumber shifts. The results indicate that the root mean square error (RMSE), angle value, and Euclidean distance are more robust to Gaussian noise. The impact of baseline profiles, however, is dependent on the normalization method. More complicated behaviors were observed encountering wavenumber shifts, which depends on both the amount of wavenumber shifts and the band patterns of the spectra under comparison. These results highlight the importance of careful preprocessing before calculating the spectral similarity.