You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
GlycanBench is a full-stack integrated web platform for glycan analysis — spanning structure building, visualization, analysis, alignment, clustering, property prediction, and AI-powered literature exploration in glycobiology.
Click-to-build glycan sequence constructor with grammar enforcement. Simultaneously generates SNFG 2D image, format conversions (IUPAC/SMILES/GlycoCT/WURCS), and a 3D conformer (ETKDGv3 + MMFF94s + UFF pipeline) rendered in 3Dmol.js.
Interconvert glycan representations: IUPAC ↔ WURCS ↔ GlycoCT ↔ SMILES. Unsupported paths (IUPAC → GlycoCT/WURCS) are surfaced explicitly as orange notices rather than silent failures.
👁️ Visualize
Tool
Description
2D Draw
SNFG-style 2D glycan structure rendering with optional per-motif color highlighting via glycowork GlycoDraw.
3D Representation
On-demand 3D conformer from any IUPAC string using a tiered optimization pipeline: ETKDGv3 (chirality enforcement, small-ring torsion corrections) → MMFF94s (up to 2×2000 iterations) → UFF fallback. Rendered in 3Dmol.js with four display styles and screenshot export.
KEGG Pathway View
Live KEGG pathway maps with server-side CORS proxy, debounced autocomplete search (FastAPI Query parameter, min 3 chars), and direct PNG download. Ten curated glycan pathway examples provided.
🔬 Analyse
Tool
Description
Monosaccharide
Taxonomic-rank-stratified occurrence charts for any monosaccharide, with configurable rank, focus, modification toggle, and frequency threshold. Powered by glycowork.motif.analysis.characterize_monosaccharide().
Glycan Insight
Retrieve full biological context — species, phyla, motifs, cell lines, disease associations, glycan class, GlyTouCan ID — from a single IUPAC string or GlyTouCan accession. Dashboard with Chart.js Doughnut (phyla), species tag cloud, disease table.
Molecule Descriptors
17 physicochemical descriptors, elemental composition with O/N ratio, 5 glycan-specific SMARTS motif counts (pyranose, furanose, N-acetyl, carboxyl, sulfate), and 5 × 2048-bit fingerprints (Morgan R2/R3, Atom Pair, Torsion, RDKit). CSV export and PubChem link.
Motif Mutation
Stochastic in-silico glycan mutagenesis with three intensity modes (normal / moderate / extreme). Generates configurable numbers of mutant sequences, extracts pentamer glycoword motifs, and plots frequency distribution as a Chart.js bar chart.
🔗 Compare
Tool
Description
Two Glycans
Side-by-side Tanimoto similarity across five 2048-bit fingerprint types (Morgan R2/R3, Atom Pair, Torsion, RDKit). Per-fingerprint hover tooltips explain what each encodes.
📐 Align
Tool
Description
Glycan Sequences
Global Needleman-Wunsch alignment using the GLYSUM glycan-specific substitution matrix or custom match/mismatch/gap scoring. Fuzzy token resolution (cutoff 0.85) handles near-exact monosaccharide names. Outputs columnar alignment, score, and percent identity. TXT export.
📊 Cluster
Tool
Description
Cluster Glycans
Cluster ≥3 glycans using agglomerative (hierarchical) or K-means methods. Configurable fingerprint type, distance metric (Tanimoto / Dice / Cosine / Euclidean / glycowork graph similarity), and linkage method. Outputs dendrogram, heatmap, and CSV of assignments.
Optimize Clusters
Threshold sweep for agglomerative clustering produces an elbow plot (cluster count vs. distance threshold) to guide parameter selection.
Detect Outliers
Singleton-cluster detection with mean pairwise distance scoring. Singletons with distance > 0.45 are flagged as STRONG OUTLIERS.
🔮 Predict
Tool
Description
Immunogenicity
MPNN (Message Passing Neural Network) predicts immunogenicity from IUPAC-condensed sequences. Glycoword-graph representation, vocabulary analysis, rule-based motif flags (AlphaGal, Neu5Gc, complex N-glycan core). Animated probability bar, confidence badge, JSON download.
💬 Chat
Tool
Description
GlycomicsChat
Glycomics-domain-enforced AI assistant (Groq openai/gpt-oss-120b). LLM-based tool router selects PubMed, ArXiv, GlyTouCan DB, Structure Analysis, or Synthesis tools per query, with keyword-heuristic fallback. Question-type analysis and heuristic confidence scoring. Structured panels for resolved GlyTouCan accession data (WURCS, IUPAC, mass, formula).
matplotlib, scipy, scikit-learn, seaborn, and openpyxl (needed for pandas.read_excel on GLYSUM.xlsx) are imported at runtime but not pinned in requirements.txt — currently satisfied transitively. Pin them explicitly if setting up a clean environment.
Frontend
Package
Purpose
React 19 + TypeScript
UI framework
Vite (rolldown-vite fork)
Build tool
Tailwind CSS v4
Styling (CSS-first config, no tailwind.config.js)
Framer Motion
Animations
3Dmol.js + NGL
3D molecular visualization
Cytoscape.js
Biosynthetic network graph
Chart.js + react-chartjs-2
Doughnut & bar charts
react-zoom-pan-pinch
KEGG pathway interactive viewer
React Router v7
Client-side routing
Axios
HTTP requests
React Icons / lucide-react
Icon libraries
Node.js requirement: rolldown-vite needs Node ^20.19.0 or >=22.12.0. On Windows, if you hit Cannot find native binding from rolldown after npm install, it's almost always an old Node version silently skipping the platform-specific optional dependency — upgrade Node (e.g. via nvm-windows) and reinstall (rm -rf node_modules package-lock.json && npm install), don't just retry the install.
IUPAC sequences are tokenized into alternating [sugar, linkage] arrays. Consecutive 5-token windows (sugar–bond–sugar–bond–sugar, step 2) form glycowords that are indexed against glycoword_vocab.json. Tokens are nodes in a bidirectional sequential chain graph.
Available Model Files
File
Architecture
Status
Models_MPNN_immunoClassifier_final.pt
MPNN
Active (deployed)
GAT_immunoClassifier_large.pt
Graph Attention Network
Stored
GIN_immunoClassifier_large.pt
Graph Isomorphism Network
Stored
LSTM_immunoClassifier_large.pt
LSTM
Stored
Datasets
File
Description
GLYSUM.xlsx
Glycan substitution matrix (Alocci et al., Glycobiology, 2015) — used for glycan sequence alignment
merged_glycan_dataset.csv
Main annotated glycan dataset — 1,356 rows, 35 columns; see schema note below
monosaccharides_counts.csv
Monosaccharide frequency table used as replacement pool during motif mutation
species_data.csv
Glycan–species associations
Dataset schema and sources
merged_glycan_dataset.csv is the union of two source files:
Source tag (source column)
Rows
Label logic
glycobase.csv
1,320
Immunogenicity label assigned by the glycowork / SugarBase pipeline (0 = non-immunogenic, 1 = immunogenic)
immunogenic_glycans_clean.csv
36
Hand-curated positives (label = 1) for well-known tumor-associated and pathogen carbohydrate antigens
The extended column set — glytoucan_id, glycan_type, disease_association, disease_id, tissue_sample, tissue_id, Species, Genus, Family, Order, Class, Phylum, Kingdom, Domain — matches the SugarBase schema distributed with the glycowork package (v12_sugarbase.json), not NIBRT GlycoBase. Citations should reference the glycowork SugarBase accordingly.
Known data-integrity issue: 12 contradictory labels
Twelve glycan strings appear in both source files with opposite labels (0 from glycobase.csv, 1 from immunogenic_glycans_clean.csv). The model trains on identical inputs with conflicting targets, which corrupts the loss surface for these structures. The affected glycans are all clinically or biologically significant:
Glycan (IUPAC-condensed)
Common name
NeuNAc(a2-3)Gal(b1-3)[Fuc(a1-4)]GlcNAc
Sialyl-Lewis A (SLe^a / CA19-9 epitope)
Fuc(a1-2)Gal(b1-4)[Fuc(a1-3)]GlcNAc
Lewis Y (Le^y)
NeuNAc(a2-3)Gal(b1-3)[NeuNAc(a2-6)]GalNAc
Disialyl-T / sialyl core-1
NeuNAc(a2-3)Gal(b1-3)GalNAc
Sialyl-T antigen
NeuNAc(a2-3)Gal(b1-4)Glc
3′-sialyllactose (3-SL)
NeuNAc(a2-3)Gal(b1-4)GlcNAc
Sialyl-LacNAc
NeuNAc(a2-6)Gal(b1-4)GlcNAc(b1-3)Gal(b1-4)Glc
6-SL-LacNAc
NeuNAc(a2-3)Gal(b1-3)GlcNAc(b1-3)Gal(b1-4)Glc
Sialyl-LNT
Gal(b1-4)GlcNAc(b1-6)GalNAc
Core-2 trisaccharide
Fuc(a1-2)Gal(b1-3)GalNAc
T antigen (core-1 disaccharide)
Man(a1-2)Man(a1-2)Man
Trimannosyl (high-mannose fragment)
Gal(b1-3)GlcNAc(b1-3)Gal(b1-4)Glc
Lacto-N-tetraose derivative
Resolution: For each conflict, the immunogenic_glycans_clean.csv label (1) is authoritative — these are experimentally confirmed antigens. The deduplicated dataset resolves each collision by keeping the row from immunogenic_glycans_clean.csv and dropping the conflicting glycobase.csv row, yielding 1,344 unique glycans with consistent labels. Run python Backend/dataset/deduplicate_dataset.py to regenerate the clean file.
Scientific Reference
See STATE_OF_THE_ART.md (and STATE_OF_THE_ART.docx) for a tool-by-tool comparison of GlycanBench against the existing state of the art in glycomics informatics, including precise algorithmic details and UX contributions for all 17 implemented tools.
Citation
Vigneshwaran CJ & Ashok Palaniappan. GlycanBench: a unified resource for working with glycans, 2026 [submitted]
Authors
Vigneshwaran CJ1 & Ashok Palaniappan1,2*
1 Systems Computational Biology Lab 2 Bioinformatics Center
School of Chemical & Biotechnology, SASTRA Deemed University