Frequently Asked Questions

What Molytica is, how to use it, and what it is not.

General

1. What is Molytica?

Molytica is a web application for exploring small molecules. You can draw a molecule's 2D structure from its SMILES string or common name, inspect its physicochemical and drug-likeness properties through a bioavailability radar, and run bioactivity predictions against specific cancer cell-line datasets. Under the hood, predictions come from a sparse dictionary learning model that represents each molecule as a sparse combination of a small number of learned substructure patterns derived from Weisfeiler–Leman graph kernels.

2. Who is this application for?

Molytica is designed for chemistry and bioinformatics students, researchers exploring cheminformatics pipelines, and anyone interested in how graph-based machine learning can be applied to molecular property prediction. It is also intended as a demonstration of the sparse dictionary learning methodology developed in this Final Year Project. Familiarity with SMILES notation is helpful but not strictly required — you can search molecules by common name via the PubChem lookup.

3. Is Molytica a drug discovery tool?

No. Molytica is developed as an academic Final Year Project for the purpose of mathematical and algorithmic verification of sparse dictionary learning applied to molecular graph data. While it implements a functional cheminformatics pipeline, it has not undergone the rigorous validation, clinical trials, or regulatory review required of a drug discovery platform. Predictions should be treated as illustrative outputs of the underlying algorithm, not as actionable pharmacological conclusions.

Using the Application

4. How do I visualize a molecule?

Navigate to the Analyze tab. Enter a SMILES string directly into the input field, or type a common compound name (e.g. "aspirin") to look it up via the PubChem PUG REST API. Once submitted, the application renders the molecule's 2D structure and calculates a set of physicochemical descriptors including molecular weight, LogP, hydrogen bond donors/acceptors, topological polar surface area (TPSA), and rotatable bond count.

5. What does the bioavailability radar show?

The bioavailability radar is a hexagonal chart displaying six key physicochemical axes: lipophilicity, size, polarity, solubility, saturation, and flexibility. Each axis is scaled to a range that characterizes orally bioavailable drug-like molecules, following the methodology described by SwissADME. A molecule whose radar profile falls mostly within the shaded drug-like zone is more likely to exhibit favorable oral bioavailability. Values that exceed the zone boundary indicate properties outside the typical drug-like range for that axis.

6. How do I run a bioactivity prediction?

To use the built-in reference models, go to the Analyze tab, switch to Predict mode, choose a cancer type, and enter a SMILES string or compound name. To use a model you trained yourself with the local trainer, go to the My Models tab, select the model, and enter a molecule there.

The backend loads the corresponding WL kernel → FDDL → classifier pipeline, computes the sparse representation of the input molecule, and returns a predicted bioactivity class. Predictions made with the reference models on the Analyze tab also include substructure-level atom heatmaps showing which parts of the molecule contributed most to the prediction.

7. What do the atom heatmaps mean?

After a reference-model prediction, Molytica overlays a color gradient on the molecule's 2D structure. Red atoms support the predicted class and blue atoms oppose it; darker shades indicate a stronger influence. The scores come from the learned dictionary atoms associated with those substructures and their coefficients in the sparse representation. This provides interpretability into which molecular substructures the model considers most relevant for the predicted class. Note that these attributions reflect the model's learned patterns and are not validated chemical explanations of bioactivity.

Training & Models

8. What is the local trainer?

The local trainer (molytica-trainer) is a Python package distributed via PyPI that runs a FastAPI server on your own machine at localhost:5000. It allows you to train sparse dictionary learning models on your own datasets locally, meaning your molecular data never leaves your computer during the training process. The trainer handles the full pipeline: Weisfeiler–Leman kernel computation, Fisher Discriminant Dictionary Learning (FDDL), and classifier fitting.

9. How do I install and run molytica-trainer?

You will need Python 3.11+ and a conda environment with RDKit installed from conda-forge (conda install -c conda-forge rdkit) — RDKit must come from conda, not pip. Then install the trainer from PyPI with pip install molytica-trainer; its remaining dependencies (scikit-learn, scipy, gensim, networkx, joblib) are installed automatically.

Launch the local server with python -m molytica_trainer.cli (on macOS/Linux you can also run molytica-train). The server starts on http://localhost:5000 and the Molytica web app detects it automatically. The Train page walks through the same steps.

10. How do I upload a trained model?

Once training completes, the Train page shows the run's results with an option to publish the model. Publishing uploads the resulting model bundle (containing the learned dictionary, classifier weights, thresholds, and a manifest) to your Molytica cloud account, where it appears under the My Models tab. Bundles are stored in your private Supabase storage bucket and become available for inference via the cloud backend. Only you can access your uploaded models.

11. What classifiers are available for training?

The training pipeline currently supports Logistic Regression, Gradient Boosting, Linear SVM (Support Vector Classification with a linear kernel), and Random Forest as the final classification stage after dictionary learning. These were selected for their compatibility with the sparse representation framework. The choice of classifier is specified when you configure a training job on the Train page.

Data & Privacy

12. Where is my data stored?

Account information and uploaded model bundles are stored in Supabase (authentication and cloud storage). Datasets you upload on the Datasets page are stored in a private Supabase bucket tied to your account. When using the local trainer, all data — including molecular datasets, intermediate kernel matrices, and trained model artifacts — remain on your local machine and are never transmitted to any external server unless you choose to publish a model.

13. What does “privacy-preserving” mean in this context?

Privacy-preserving refers specifically to the local training architecture. Because molytica-trainer runs entirely on your machine, sensitive or proprietary molecular datasets do not need to leave your local environment for model training. The only data that crosses the network is the final trained model bundle when you explicitly choose to upload it for cloud-based inference. This design is intentional for research settings where molecular data may be confidential or proprietary.

Limitations & Disclaimer

14. Are the predictions clinically validated?

No. The bioactivity predictions are outputs of a sparse dictionary learning model trained on NCI cancer cell-line screening datasets (NCI-1, NCI-33, NCI-41). These datasets are standard graph classification benchmarks used in machine learning research. The predictions have not been validated against clinical outcomes, and the model's accuracy is bounded by the size and representativeness of the training data. This application exists to demonstrate and verify the mathematical properties of the WL kernel and FDDL classification pipeline, not to provide medically actionable predictions.

15. Can I use Molytica results in a research paper?

You may reference Molytica as a demonstration tool and cite the underlying methodology (Weisfeiler–Leman kernels, Fisher Discriminant Dictionary Learning). However, any published use should clearly state that the results are from an unvalidated academic prototype and should not be presented as evidence of pharmacological activity. We recommend citing the Final Year Project report and the original algorithmic references.

16. What are the known limitations?

The current implementation has several known limitations: (a) predictions are limited to three cancer types based on the available NCI benchmark datasets; (b) the bioavailability radar uses clamped normalization, meaning extreme out-of-range values are clipped to the boundary rather than displayed beyond it; (c) atom heatmap attributions reflect learned statistical patterns, not causal chemical mechanisms; (d) the local trainer requires a moderately capable machine for kernel computation on large datasets; (e) the CORS configuration of the local trainer permits connections from any *.vercel.app subdomain, which is a known security limitation in the current release.