6 Downloads Updated 3 weeks ago
ollama run onnyx/Prism
Updated 3 weeks ago
3 weeks ago
de06bb7214c5 · 2.0GB
PRISM is not just a machine learning model; it is a comprehensive, end-to-end software pipeline designed to bridge the gap between raw mass spectrometry data and structural elucidation. By combining Retrieval-Augmented Generation (RAG) with Deep Sequence Translation, PRISM translates complex LC-MS/MS spectra into chemically valid 2D structures (SMILES and SELFIES) for both known library compounds and entirely novel molecules. Unlike traditional spectral matching tools that are limited to reference libraries, or pure deep learning models that frequently generate chemically invalid structures, PRISM integrates a robust cheminformatics validation layer and a scalable vector retrieval database to ensure high accuracy, chemical validity, and interpretability.
Key Features End-to-End Pipeline: Ingests raw spectral data (.mzML, .mgf, .mzXML) and outputs structured chemical data (.csv, .sdf, .json). Hybrid Retrieval-Augmented Architecture: Combines fast nearest-neighbor spectral retrieval with deep learning inference to handle both known and novel compounds seamlessly. Guaranteed Chemical Validity: Built-in RDKit integration sanitizes and validates all generated SMILES/SELFIES, filtering out chemically impossible structures. Dual Representation Support: Natively outputs both SMILES (for standard cheminformatics) and SELFIES (for robust, 100% valid generative molecular representation). Flexible Deployment: Use PRISM via a powerful Command Line Interface (CLI) for batch processing, or integrate it directly into your Python workflows via the API. Scalable Database: Built-in utilities to build, manage, and query custom spectral reference databases using FAISS.
Architecture & Workflow PRISM operates as a multi-stage software pipeline: Spectral Preprocessing: Raw spectra are binned, normalized, and noise-filtered. Retrieval Module: The query spectrum is embedded and searched against a local vector database (FAISS) to retrieve the top- k k most similar reference spectra. Inference Module: A Transformer-based sequence-to-sequence model takes the preprocessed spectrum and the retrieved contextual embeddings to predict the molecular sequence (SMILES/SELFIES). Cheminformatics Validation: Predicted sequences are parsed via RDKit. Invalid structures are discarded or re-sampled, and valid structures are converted into 2D coordinates for visualization.
Input / Output Formats Supported Inputs .mgf (Mascot Generic Format) .mzML (HUPO-PSI standard) .mzXML .msp (NIST format) Supported Outputs .csv / .tsv: Tabular data containing Spectrum ID, SMILES, SELFIES, Confidence Score, and Retrieval Metadata. .json: Nested JSON format including full spectral arrays and all top- k k predictions. .sdf: Structural Data File containing the generated 2D molecular structures, ready for visualization in tools like PyMOL, ChemDraw, or KNIME.