Releases · BojarLab/glycowork · GitHub
Skip to content

Releases: BojarLab/glycowork

v1.9.0

Choose a tag to compare

@Bribak Bribak released this 16 Jun 11:07
f0e3afc

Changelog

[1.9.0]

  • Bumped Python version to 3.11 (3.10 reaches end of life in October 2026; https://devguide.python.org/versions/) (6a94ea9)
  • Bumped minimum pandas version to 2.1 (faf68f0)
  • Fixed scipy version as >=1.16 to guarantee Games-Howell test in the ANOVA for biodiversity tests (cf5e7bd)
  • The dev optional install set (only used for testing) now also includes glycontact>=0.3.3
  • Automated versioning in __init__.py (b72d113)

glycan_data

  • Curated datasets are now stored in a dedicated folder .glycan_data.datasets, for tidiness (c91ee8d)
  • Added new curated glycomics datasets: human_interstitialfluid_N_GPST000651, celllines_colorectal_O_PMID32236654, celllines_colorectal_GSL_PMID35489554, human_serum_ovarian_N_PMID41366864, mouse_platelet_N_101161ATVBAHA126324757 (ce499ae, 7ddd6c2, c91ee8d, 086d589, c29ceeb, e6f59e6)
  • Renamed and re-curated glycomics_mouse_gastric_O_GPST000464 to glycomics_mouse_gastric_O_PMID40667878 (c91ee8d)
  • Added new contrasts file to catalog which sample belongs to which comparison group for the curated glycomics datasets (aad159e)
  • df_species now contains a new taxon_id column with TaxIDs (3d304a3)

loader

Added ✨
  • Added the GlycoList class that equips List with graph-based matching capabilities (overriding index, remove, count, and in with compare_glycans isomorphisms) (49c9aeb)
  • Added the glycans, abundance, group1, group2, groups, paired, and name properties to GlycoDataFrame to benefit from GlycoList capabilities in glycans, get the abundance values via abundance and provide functions in glycowork.motif.analysis with group labels and information whether data are paired or not (aad159e, c91ee8d, cd1b9aa)

stats

Added ✨
  • MissForest now has a circadian API (also exposed in impute_and_normalize, to fine-tune data imputation of rhythmic data (starting from the same-phase median, instead of the overall median) (4305377)
Fixed 🐛
  • Fixed behavior of hotellings_t2 when paired=True (e419fc4)
  • Fixed get_glycoform_diff mutating its input (e419fc4)

motif

graph

Added ✨
  • All graph functions (e.g., compare_glycans, subgraph_isomorphism) now support narrow monosaccharide wildcards (e.g., Gal/Man instead of Hex) (f85ab80)
Changed 🔄
  • Made glycan graph caching in glycan_graph_memoize somewhat faster (cf5e7bd)
Fixed 🐛
  • Fixed some degree calculations in generate_graph_features (cf5e7bd)
  • Fixed graph_to_string very rarely not ordering branches in a deterministic/idempotent manner (9f42561)
Deprecated ⚠️

processing

Added ✨
  • Universal Input via canonicalize_iupac can now deal with more monosaccharide cases, such as Ribp, Glc1OMe, or single-monosaccharide glycans such as aDGlcpA (6754bbf, 426ff9d)
  • Universal Input via canonicalize_iupac can now more robustly handle modifications in CSDB-linear, such as in Ac(1-5)aXNeup(2-6)[Ac(1-2)]bDGalpN(1-4)bDGalp(1-4)bDGlcp, S-3)bDGlcpA(1-3)bDGalp(1-4)[Ac(1-2)]bDGlcpN(1-3)bDGalp(1-4)bDGlcp, -4)[S-2)]aD3,6anhGalp(1-3)[S-2)]bDGalp(1-, or Ac(1-2)[xXEt?N(1-P-6)]bDGlcpN(1-3)bDManp(1-4)bDGlcp, as well as more robustly strip reducing end anomeric indicator (99942e8, a152c47, d6c3d56, ced60fc, a6882db)
  • Universal Input via canonicalize_iupac can now parse more complex Oxford sequences, such as F(6)A2G(4)2S(3,3)2 (9f42561)
  • max_specify_glycan will now also specify these cases: ("Fuc(a1-?)GlcNAc", "Fuc(a1-3/4)GlcNAc"), ("Fuc(a1-?)]GlcNAc", "Fuc(a1-3/4)]GlcNAc"), ("Fuc(a1-?)Gal(", "Fuc(a1-2)Gal("), ("GalOS", "Gal3/6S"), ("GlcNAcOS", "GlcNAc6S"), ("Fuc(a1-?)[Gal(b1-?)]", "Fuc(a1-3/4)[Gal(b1-3/4)]"), ("Man(a1-?)Man", "Man(a1-2/3/6)Man") (f3b4389, 487dad4)
  • Oxford parsing in canonicalize_iupac will now detect hybrid glycans and will add extra mannoses to the a1-6 branch (c4ba7f9)
Changed 🔄
  • Universal Input via canonicalize_iupac now is more robust to modified reducing ends in IUPAC-extended glycans (6754bbf)
  • canonicalize_iupac has again been made more robust against typos and personal idiosyncrasies in nomenclature (39b1d25, 9f57fa5, e419fc4)
Fixed 🐛
  • Fixed some Oxford sequences (e.g., A1, A2) being misidentified as blood group glycolipids by canonicalize_iupac (9f42561)
  • Fixed canonicalize_composition mistaking sialic acid composition blocks as sulfate in compositions such as H2N2S1Sulf1 (bb822c8)
  • Fixed oxford_to_iupac wrongly ordering NeuOAc5Ac instead of Neu5AcOAc (eda4474)

tokenization

Added ✨
  • The modification keyword argument in mz_to_composition etc now also accepts procainamide as an argument (e4a2a40)
  • mz_to_composition and mz_to_structures now have a new max_charge keyword argument that sets the maximum applicable charge state as well as the ion mode (e1475f9)
  • mz_to_composition and mz_to_structures now also support ppm-level mass tolerances via the new tolerance_unit keyword argument that allows users to switch between Da and ppm (e1475f9)
Changed 🔄
  • mz_to_composition now also filters by provided glycan_class if a user provides a custom df_use (ab57479)
Fixed 🐛
  • composition_to_mass now correctly factors in the extra methylation (former ring oxygen) that happens in the combination of modification == 'reduced' and sample_prep == 'permethylated' (30e1a46)
  • Fixed incompatibility of condense_composition_matching with scikit-learn >= 1.9.0 (34d839f)
Deprecated ⚠️
  • Removed the keyword argument mode from mz_to_composition and mz_to_structures; will be handled by the new max_charge instead (e1475f9)
  • Removed the doubly_charged option from the extras keyword argument in mz_to_composition; will be handled by the new max_charge instead (e1475f9)

annotate

Added ✨
  • annotate_glycan now also exposes the keyword argument condense=False, analogous to annotate_dataset, to get only non-zero motifs (5d4fa70)
Changed 🔄
  • Arguments glycans and feature_set in quantify_motifs have been changed to keyword arguments with defaults glycans = None (will be inferred from first column if it contains glycans, otherwise needs to be supplied) and feature_set = ['known', 'exhaustive'] (c5db6a8)
  • Wildcard monosaccharide-only columns (e.g., dHex) in get_k_saccharide will now also aggregate signal from their specific instances (e.g., Fuc) across the dataset (was already working like that for Sia and for k>1) (6de40f9)
Fixed 🐛
  • quantify_motifs can now also be used with full datasets that still have the first column be a glycan string column (fa98caa)
  • get_k_saccharides will no longer mistakenly capture narrow linkage wildcards like b1-3/4 in its feature columns for k=1 (where there should only be monosaccharides) (c209f85)

draw

Added ✨
  • When supplied with narrow monosaccharide wildcards, GlycoDraw will now draw bisected monosaccharides (e.g., a blue-yellow split rectangle for GlcNAc/GalNAc) (f85ab80)
Changed 🔄
  • If draw_method = chem3d, GlycoDraw will now preferentially fetch a realistic conformer from GlycoShape/PDB via glycontact, if the user has glycontact installed (lazily imported), and only fall back to RDKit if none can be found (7a59d08)
Fixed 🐛
  • Made SVG parsing in annotate_figure more robust (e3dce0e)
  • Glycans with / in their sequence can now be saved with the glycan name as the filename in GlycoDraw without error (edba778)

analysis

Changed 🔄
  • Several functions are now more robust toward the specific glycan column naming for processing (aad159e)
  • get_jtk now internally uses fine-tuned data imputation for rhythmic data via MissForest improvements (4305377)
Fixed 🐛
  • Fixed column access in ALR-treatment of get_glycanova (aad159e)
  • Fixed warnings when using scikit-learn>=1.9.0 due to the deprecated penalty keyword arg in sklearn.linear_model.LogisticRegression used in multi_feature_scoring (34d839f)
  • Fixed edge case where get_glycanova could assign effect sizes to wrong glycans, if glycans had been dropped due to variance filtering (e419fc4)
Deprecated ⚠️
  • Deprecated glycan_col_name keyword argument in get_pvals_motif and characterize_monosaccharide; will be auto-detected (aad159e)

regex

Fixed 🐛
  • Specifying exact occurrences in get_match, as in get_match("[HexNAc]{2}", "Gal(b1-4)GlcNAc(b1-4)GlcNAc"), is now more robust/accurate (28894d8)
Deprecated ⚠️
  • Deprecated process_occurrence, process_main_branch, and process_question_mark as they are handled in-line now (28894d8)
  • Deprecated the lookahead keyword argument in fill_missing_in_list (handled automatically) (487dad4)

v1.8.1

Choose a tag to compare

@Bribak Bribak released this 24 Apr 08:05
7f615b0

Changelog

[1.8.1]

  • fixed deploy action internally still relying on nbdev2 (84134fb)

glycan_data

loader

Changed 🔄
  • Changed human_macrophages_N_2024-11-28-625934 and human_macrophages_O_2024-11-28-625934 glycomics datasets to human_macrophages_N_2024_11_28_625934 and human_macrophages_O_2024_11_28_625934 (6a673d0)
  • Recurated human_brain_GSL_PMID40207879 glycomics dataset with improved nomenclature conversion (cf8706a)
Fixed 🐛
  • Fixed one faulty sequence in df_glycan that caused graph generation to fail (b2f6ab9)

stats

Added ✨
  • Added hsic to calculate Hilbert-Schmidt Independence Criterion between variables, to measure dependency (c560fbb)

motif

analysis

Changed 🔄
  • Added distance matrix to beta diversity output in get_biodiversity (dca7820)
Fixed 🐛
  • Fixed column names slipping into column values when motifs = True combined with transform = ALR in get_pca (e802da1)
  • Made motif abundance re-normalization more robust in preprocess_data (ac6fa53)

draw

Changed 🔄
  • Improved branch spacing in GlycoDraw for highly branched glycans (6a673d0)

tokenization

Added ✨
  • Added a mass_tag float keyword argument to mz_to_composition and mz_to_structures for glycans tagged at the reducing end (7adaf75)
  • Added a modification string keyword argument to composition_to_mass and glycan_to_mass for glycans tagged at the reducing end (8160490)
  • mz_to_composition and mz_to_structures now also support the combination of multiply-charged ions with adducts (8160490)
Fixed 🐛
  • Fixed mass calculation of additionally acetylated glycans in glycan_to_mass (7adaf75)
Deprecated ⚠️
  • The reduced bool keyword argument in mz_to_composition and mz_to_structures has been replaced with the modification string keyword argument (8160490)
  • Deprecated the reducing_end keyword argument in match_composition_relaxed, as it was no longer being used (3f459b2)
Fixed 🐛
  • Fixed a bug in mask_rare_glycoletters in which rare linkages occasionally were not masked (3f459b2)

processing

Added ✨
  • Added some more lipid shorthands (e.g., Fuc-GD1a or Fuc-GA1) to Universal Input/canonicalize_iupac (cf8706a)
Changed 🔄
  • Universal Input via canonicalize_iupac can now deal with more pyranose indicators (e.g., Altp or Lyxp) (467673b)
  • Universal Input via canonicalize_iupac can now deal with more sulfate variants (e.g., 6-O-sulfo, [S-6]) (0e102a6, cbe20da)
Fixed 🐛
  • Fixed overeager modification of already correctly formatted 6PCho modifications in canonicalize_iupac (467673b)
  • Fixed max_specify_glycan not specifying the chitobiose core in N-glycans (3f459b2)

network

biosynthesis

Added ✨
  • Added get_biosynthetic_coherence function to estimate how well glycan abundances can be predicted from biosynthetic networks, to disentangle biosynthetic vs carrier variance (b7020fd)

v1.8.0

Choose a tag to compare

@Bribak Bribak released this 24 Mar 07:59
773f23d

Changelog

[1.8.0]

  • fixed Quarto accessing of pyproject.toml attributes for doc building (cd9b62f)

glycan_data

loader

Added ✨
Changed 🔄
  • Specified wildcards in glycomics_human_colorectal_O_PMC9254241 (e71550d)
Fixed 🐛
  • Made sure that incomplete API access in get_molecular_properties does not lead to outright failure (52c6cf9)
  • glycomics_data_loader and other LazyLoader instances are now robust against duplicate column names with the .1, .2 suffix (they will be stripped now) (44e8473, 1cdb270)

motif

annotate

Added ✨
  • get_k_saccharides and annotate_dataset can now dynamically create enrichment motifs of the type Sia(a2-3)Gal or Terminal_Sia(a2-3/6) if multiple sialic acid types are present in input data (522b7cf)
Fixed 🐛
  • Made sure curly bracket sequence content ("floaty bits") are correctly counted in count_unique_subgraphs_of_size_k (522b7cf)
  • Make sure all narrow linkage wildcards, even if not present in linkages, are being correctly parsed in count_unique_subgraphs_of_size_k (5220912)

graph

Changed 🔄
  • Added _prefilter_labels for more cheap checks to avoid graph operations and thus make compare_glycans and subgraph_isomorphism considerably faster (b865229)
  • Made glycan_to_graph function much faster (up to 10x) (750cdb1)
  • Made graph_to_string_int function ~40% faster (750cdb1)
Deprecated ⚠️
  • Deprecated evaluate_adjacency; will be handled in-line in glycan_to_graph (750cdb1)
  • Deprecated canonicalize_glycan_graph; will be handled in-line in graph_to_string_int (750cdb1)
  • Deprecated neighbor_is_branchpoint; no longer in use (e020ffb)

draw

Changed 🔄
  • HexN, dHexNAc, and HexA shapes now get drawn in fewer objects/more efficiently (10da7c5)
Fixed 🐛
  • Fixed displaying beta-linkages instead of alpha-linkages in annotate_figure (e71550d)
Deprecated ⚠️
  • Deprecated scale_in_range; has been in-lined instead (855a9f8)
  • Deprecated process_repeat; has been in-lined instead (855a9f8)

analysis

Changed 🔄
  • get_volcano can now also deal with input dataframes that have the Glycan column be the index instead (e71550d)
  • Equivalence p-values in get_differential_expression now also use the same sample-size adjusted alpha as regular p-values (3884125)
  • Specifying return_plot=True in get_heatmap will now also return the column names and the transformed dataframe, next to the plot object (3b72129)
  • Improved default plot styling for outputs from functions (855a9f8)
Fixed 🐛
  • CLR-transformation for paired data in preprocess_data now correctly uses the shared geometric mean as reference, to preserve within-pair differences (3884125)
  • Fixed equivalence p-values in get_differential_expression if sets=True (3884125)
  • CLR-transformed motif-level quantification in preprocess_data and get_pca used the glycan-level geometric mean as a reference, rather than the motif-level geometric mean, which is now fixed (c71c385)
  • get_roc now saves the figures for all classes, not just the last, in a set-up of filepath + multi-group comparison (855a9f8)
  • User-provided random_state values/generators are now correctly propagated through to multi_feature_scoring (855a9f8)

tokenization

Added ✨
  • mz_to_composition now has a new keyword argument deprioritized, which is a set of disfavored monosaccharides/modifications that will only be used if no composition can be found otherwise (i.e., less harsh than full exclusion via filter_out). This keyword argument is now also exposed in mz_to_structures (316f962)

tokenization

Changed 🔄
  • canonicalize_iupac now is even more robust regarding typo correction (acf05e1)

network

biosynthesis

Added ✨
  • Added build_network_from_glycans handler to do a BFS-search to get the bulk biosynthetic network going (b865229)
  • Added hierarchical option (now the new default) to the keyword argument options in plot_format in plot_network, for a more organized network display (8d03348)
  • extend_network now has the new auto_steps keyword argument, which (if to_extend is a target composition), will calculate the minimum number of steps, cross-check it against the provided maximum as steps, and then iteratively extend the most favorable leaf nodes toward the target composition (f8f2fa9)
Changed 🔄
  • construct_network is now more than twice as fast (a1c810c, b865229)
  • Dynamic wildcard construction in get_differential_biosynthesis now also creates the most parsimonious narrow wildcards, similar to annotate (e71550d)
  • Renamed the Feature column in get_differential_biosynthesis to Glycan (e71550d)
  • extend_network now accepts compositions in any format in the to_extend keyword argument, using Universal Input (18f7ba5)
  • extend_network now early-exits if the composition provided in to_extend already exists within the network, outputting the existing matching structures in the network (18f7ba5)
  • monolink_to_enzyme is now comma-separated instead of tab-separated and is more complete (10dc46e)
Fixed 🐛
  • Fixed reaction hover label in plot_network (8d03348)
  • Fixed a bug in add_high_man_removal which set the edge labels with a lambda function instead of a string (f2b5f99)
Deprecated ⚠️
  • Deprecated find_shared_virtuals, adjacencyMatrix_to_network, get_virtual_nodes, get_neighbors, create_adjacency_matrix; now all handled in-line (a1c810c, b865229)
  • Deprecated find_path, find_shortest_path, deorphanize_nodes, shells_to_edges, which is all now handled by the new build_network_from_glycans (b865229)

v1.7.1

Choose a tag to compare

@Bribak Bribak released this 16 Feb 09:16
d1807b4

Changelog

[1.7.1]

  • glycorender version bump from 0.2.3 to 0.2.5 (1933574)
  • upgraded nbdev2 to nbdev3 for the documentation (+ removed now unnecessary files) (eb3f727)
  • improved start-up time of the package (i.e., time at first import in a session) (10a39f0)

motif

draw

Changed 🔄
  • Generic substituents will now be properly formatted in GlycoDraw (89eb687)
  • Unknown base monosaccharides in GlycoDraw now correctly default to blank hexagons (89eb687)
  • Make sure GlycoDraw can draw !-containing sequences (e.g., Internal_LewisA) even with restrict_vocab=True (1933574)
Fixed 🐛
  • Make sure reducing_end_label is perfectly y-centered in GlycoDraw (7e9e980)
  • Fixed setting utf-8 as default encoding in annotate_figure (1933574)

processing

Added ✨
  • Added LacdiNAc to the common_names support in Universal Input (d1140d1)
  • Added max_specify_glycan function to infer sequence ambiguities/uncertainties as best as possible (e2cf92a)
Fixed 🐛
  • canonicalize_iupac is now more robust when handling variant modification dialects in IUPAC-condensed (i.e., not mistaking them for CSDB-linear), such as Galβ1-3(6SGlcNAcβ1-6)GalNAcol (046ea12)
  • min_process_glycans and get_lib now correctly handle glycans with floating modifications, such as {6S}{Neu5Ac(a2-3)}Gal(b1-4)GlcNAc(b1-6)[Gal(b1-3)]GalNAc (68f1e1b)

analysis

Changed 🔄
  • characterize_monosaccharide is now much faster (0de71c5)
Fixed 🐛
  • Fixed temporary file handling in annotate_volcano=True in get_volcano (1933574)

annotate

Added ✨
  • Added new get_minimal_ksaccharide_ambiguity function to find the minimal needed narrow linkage wildcard to encompass all variants in dataset (8a0bbce)
Changed 🔄
  • feature_set options exhaustive and the terminal variants now fully lean into narrow linkage wildcards for dynamically generated wildcards (e.g., a2-3/6), instead of the broader a2-? versions, which are scoped based on the provided data (8a0bbce)
  • get_terminal_structures can now be used for any size value, not only 1 and 2 (ef353fb)
  • annotate_dataset will now internally use get_terminal_structures for the terminal3 feature-set keyword (ef353fb)
Fixed 🐛
  • Fixed topologically incorrect disaccharides in get_terminal_structures output (ef353fb)

ml

models

  • When using prep_model with trained=True on SweetNet-type models, the function now auto-corrects the num_classes value, if a wrong output dimension is provided (i.e., if it clashes with the trained model) (ccf2d34)
Fixed 🐛
  • Fixed warning message in train_ml_model about not specifying feature_calc (0de71c5)

v1.7.0

Choose a tag to compare

@Bribak Bribak released this 19 Nov 10:04
303308e

Changelog

[1.7.0]

  • Added some more lazy loading of drawing-related imports to improve package start-up time (b806bdd)
  • glycowork now requires at least version 0.2.3 of glycorender[png] (d51c0eb)
  • glycowork now requires Python>=3.10, as 3.9 is no longer supported by the Python Foundation (2b8838c)
  • switched type hints to native type hints (supported from Python 3.10) (2b8838c, 9b7b192)

motif

processing

Added ✨
  • Added the verbose keyword argument (default = True) to glytoucan_to_glycan, to suppress the output of non-matched IDs (aa0b9a4)
  • Added the kcf_to_iupac function to convert the KCF nomenclature into IUPAC-condensed (bfe947c, 37fad0d)
  • Added the glycoctxml_to_iupac function to convert the GlycoCT XML nomenclature into IUPAC-condensed (3d967b6)
  • Added the is_composition utility function to quickly check whether a string is a composition or a glycan (dcd3ffe)
Changed 🔄
  • Improved the detection of LinearCode sequences in canonicalize_iupac with the new looks_like_linearcode helper function (e26566e)
  • Improved the hook for triggering Oxford nomenclature conversion in canonicalize_iupac to be less permissive and faster (c5ae5e5, 25186e5)
  • Renamed linearcode1d_to_iupac to glyseeker_to_iupac (d44aba1)
  • Improved the hook for checking GlyTouCan IDs in canonicalize_iupac by making it more specific (aa0b9a4)
  • Improved wurcs_to_iupac handling of complex sequences to have more robust WURCS handling in canonicalize_iupac (0f5a5d8)
  • Improved glycoworkbench_to_iupac handling of variable reducing ends to have more robust GlycoWorkbench handling in canonicalize_iupac (0f5a5d8)
  • canonicalize_iupac can now also convert KCF sequences, thanks to the new kcf_to_iupac function (bfe947c, 37fad0d)
  • canonicalize_iupac can now also convert GlycoCT XML sequences, thanks to the new glycoctxml_to_iupac function (3d967b6)
  • Arbitrary chemical substituents from CSDB-linear glycans are now being correctly handled in canonicalize_iupac, even if the specific substituent is not yet supported (b71af07)

draw

Added ✨
  • Added the reducing_end_label keyword argument to GlycoDraw to display any connected text to the right of the glycan (such as "protein", which will be connected via a regular linkage) (c36a1a6)
  • Added the GlycanDrawing class to allow GlycoDraw to output glycorender aesthetics in a Jupyter notebook context (d51c0eb)
Changed 🔄
  • Improved branch spacing for complex glycans (i.e., less overlap) within GlycoDraw (39ba99c)
  • Added the restrict_vocab keyword argument to GlycoDraw (default: False) to support drawing of exotic glycans while still facilitating a restricted vocabulary for annotating glycans in figures (fcddb36)

annotate

Added ✨
  • Added the get_glycan_similarity function to calculate cosine similarities between glycan motif fingerprints between two glycan sequences (bd4d071)
Changed 🔄
  • get_k_saccharides now has limited support to extract information from inputs that are a mix of sequences and compositions; namely monosaccharide counts, when up_to = True (dcd3ffe)

ml

model_training

Added ✨
  • Added WarmupScheduler class to (by default) have a warming-up period of learning rate schedule (for training stability) in training_setup (469649a)
Changed 🔄
  • train_model now supports training GIFFLAR-type glycan models (1a7e720, 18f42e5)
  • training_setup has the new warmup_epochs keyword argument (default = 5) that determines the length of the learning rate warm-up schedule (469649a)
  • train_model now performs gradient clipping for improved training stability (469649a)

models

Added ✨
  • The GIFFLAR model can now be requested from prep_model via "GIFFLAR" as model_type (1a7e720, 18f42e5)

inference

Changed 🔄
  • Added the multilabel keyword argument to glycans_to_emb, to support inference for multilabel outputs (0437583)

glycan_data

loader

Added ✨
  • Added parse_lines utility function to parse copy-pasted content from an Excel column into a list (b806bdd)

v1.6.4

Choose a tag to compare

@Bribak Bribak released this 01 Oct 09:31
b0b4205

Changelog

[1.6.4]

  • The required version for glyles, when using the [chem] or [all] optional installs, has been bumped up to 1.2.3a0 to resolve dependency conflicts (27eb990)
  • Added new glycomics, glycoproteomics, and lectin microarray datasets (5558f1e)
  • Added link to canonicalize web app into README (3c862db)

motif

draw

Added ✨
  • GlycoDraw now has the new keyword argument highlight_linkages, which will draw selected linkages in red and thicker (6b60a53)

analysis

Added ✨
  • get_pca now has the new keyword argument size, to let users control the size of points with a scalar column in the provided meta-data (ab7669c)

graph

Fixed 🐛
  • Fixed an issue in subgraph_isomorphism_with_negation, where motif graphs were only shallowly copied, potentially causing graph mutation during processing and leading to too permissive matching in annotate_dataset and higher-level functions (8b75aae)

processing

Changed 🔄
  • Using canonicalize_iupac on a monosaccharide contained in lib now has an early return, preventing overlapping name spaces with the common names (040cbc8)
  • Added support for old 'z' uncertainty notation in canonicalize_iupac via replace_dic (b079ece)
  • Support C2-inference for beta-linked sialic acid in canonicalize_iupac (563ea6d)
  • Support CarbBank IUPAC dialect in canonicalize_iupac (563ea6d)
  • Generalized handling and sorting of Neu5,9Ac type modifications in canonicalize_iupac (757a24f)
  • Improved handling of CSDB-linear modifications in canonicalize_iupac (757a24f)

v1.6.3

Choose a tag to compare

@Bribak Bribak released this 25 Jul 13:28
28d8ab5

Changelog

[1.6.3]

  • glycowork is now compatible with specifying narrow modification ambiguities (e.g., Gal(b1-3)GalNAc4/6S) (ec290e8)
  • made the bokeh dependency runtime-optional by importing it just-in-time for plot_network (ea9929e)

glycan_data

stats

Added ✨
  • Alpha biodiversity calculation in alpha_biodiversity_stats now performs Welch's ANOVA instead of ANOVA if scipy>=1.16 (ab73368)
  • ALR transformation functions now also expose the random_state keyword argument for reproducible seeding (23cafe7)

motif

processing

Added ✨
  • COMMON_ENANTIOMER dict to track the implicit enantiomer state (e.g., we write Gal instead of D-Gal but we do note the deviation L-Gal) (bb7575c)
  • GLYCONNECT_TO_GLYTOUCAN dict to support GlyConnect IDs as input to Universal Input / canonicalize_iupac (ea9929e)
Changed 🔄
  • canonicalize_iupac and its parsers will now leave the D-/L- prefixes in monosaccharides, which will then be centrally homogenized with COMMON_ENANTIOMER, for a more refined and detailed output (bb7575c)
  • canonicalize_iupac now considers more IUPAC variations, such as Neu5,9Ac instead of Neu5,9Ac2 (a764897)
  • canonicalize_iupac no longer strips trailing -Cer (d8c948b)
  • canonicalize_iupac now handles alpha and beta (d8c948b)
  • glycoworkbench_to_iupac is now trigged by presence of either End-- or u-- (d8c948b)
  • wurcs_to_iupac now supports more tokens (d9d6e57)
  • canonicalize_iupac now supports Gal4,6Pyr modifications (487c68a)
  • wurcs_to_iupac can now process sulfur linkages (e.g., Glc(b1-S-4)Glc) (88b2d54)
  • wurcs_to_iupac is now more robust to prefixes (e.g., L-, 6-deoxy-, etc) (ac171c5)
  • wurcs_to_iupac can now deal with ultra-long glycans (i.e., a-z, A-Z, aa-az, and aA-aZ) (487c68a)

tokenization

Changed 🔄
  • glycan_to_composition is now compatible with the new narrow modification ambiguities (e.g., Gal(b1-3)GalNAc4/6S) (ec290e8)

graph

Changed 🔄
  • compare_glycans is now compatible with the new narrow modification ambiguities (e.g., Gal(b1-3)GalNAc4/6S) (ec290e8)

draw

Fixed 🐛
  • fixed overlap in floating substituents in GlycoDraw if glycan had fewer branching levels than unique floating substituents (daade78)

analysis

Added ✨
  • ANOVA-based time series analysis in get_time_series now performs Welch's ANOVA instead of ANOVA if scipy>=1.16 (ab73368)
  • All analysis endpoint functions can now be directly seeded, without having to pre-transform data, with the newly exposed random_state keyword argument (23cafe7)

v1.6.2

Choose a tag to compare

@Bribak Bribak released this 09 Jun 08:14
6a87763

Changelog

[1.6.2]

glycan_data

loader

Changed 🔄
  • huggingface_hub will now only be imported upon running download_model, making it technically run-time optional and improving package start-up time (d87e8af)

draw

Changed 🔄
  • openpyxl will now only be imported upon running plot_glycans_excel, making it technically run-time optional and improving package start-up time (d87e8af)

processing

Added ✨
  • canonicalize_iupac now removes extraneous quote marks around input glycans (fbe454c)
  • Added more milk oligosaccharide common names to the Universal Input pipeline as recognized by canonicalize_iupac (39e8a19)
Changed 🔄
  • canonicalize_iupac will now recognize GLYCAM sequences terminating in -OME (6430ebb)
Fixed 🐛
  • Fixed capitalisation in mapping of IGG N-glycan codes to account for .lower() call in canonicalize_iupac (48fb211)
  • Fixed variant LDManHep handling in canonicalize_iupac (6430ebb)

v1.6.1

Choose a tag to compare

@Bribak Bribak released this 29 May 05:17
2f673a8

[1.6.1]

  • Moved xgboost dependency into the optional [ml] install (0c62acf)
  • glycowork now no longer has a svglib dependency, due to improvements in glycorender, requiring glycorender[png]==0.2.0 (4cad68f)

motif

graph

Changed 🔄
  • glycan_to_nxGraph_int will now automatically convert provided lib dicts into HashableDict objects, if they aren't already (fe5cd74)
  • compare_glycans used with two strings now has another early-return condition if the two glycans have different numbers of branches, enhancing efficiency (fe5cd74)

processing

Changed 🔄
  • canonicalize_iupac can now handle some more variations, such as double-anomeric linkages ((a2-1b)), and will leave modification-containing-seeming monosaccharides (e.g., Psif, Sorf) intact (5b25c1f)

v1.6.0

Choose a tag to compare

@Bribak Bribak released this 27 May 09:10
ab4bc2b

Changelog

[1.6.0]

  • All glycan graphs are now directed graphs (nx.Graph --> nx.DiGraph), flowing from the root (reducing end) to the tips (non-reducing ends), which has led to code changes in quite few functions. Some functions run faster now, yet outputs are unaffected (03dfad6)
  • Added huggingface_hub>=0.16.0 as a new dependency to facilitate more robust model distribution (22f6b8f)
  • Moved drawSvg~=2.0 and openpyxl from the optional [draw] install to the dependencies of base glycowork. That allows for the usage of GlycoDraw in, e.g., Jupyter environments etc, even if glycowork[draw] has not been installed. Since these dependencies are unproblematic, no special install needs to be followed for the base glycowork install (60e51da)
  • Deprecated the optional [draw] install completely, by replacing the problematic cairosvg dependency with our new & custom renderer glycorender, which is now a new base dependency of glycowork (7c4fbe1)
  • Moved Pillow dependency into glycorender (793e71f)
  • Deprecated mpld3 and matplotlib-inline dependencies; added new bokeh and IPython base dependencies for better interactive plotting in a Jupyter environment (972c34b, 13b0699)
  • Formally added numpy and matplotlib to base dependencies (ba40c73)
  • Exposed canonicalize_iupac to the glycoworkGUI (ba40c73)
  • Implemented submodule lazy loading to speed up package imports & start-up (9bf18f7)

glycan_data

loader

Added ✨
  • Added HashableDict class to allow for caching of functions with dicts as inputs (03dfad6)
  • Added GlycoDataFrame class to extend pd.DataFrame by adding the .glyco_filter method, to easily filter glycan dataframes by the occurrence/count of sequence motifs (9764b3e)
  • Added new curated glycoproteomics dataset: sorghum_N_PMID39137587 (13b0699)
  • Updated glycan_binding, df_glycan, df_species to be bigger, better, and cleaner (e302075)
Changed 🔄
  • Refined motif definition of Internal_LewisX/Internal_Lewis_A/i_antigen in motif_list, to exclude LewisY/LewisB/I_antigen from matching/overlapping (07c9c12)
  • Renamed Hyluronan in motif_list into Hyaluronan (07c9c12)
  • Removed Nglycolyl_GM2 from motif_list; it's captured by GM2 (07c9c12)
  • Further curated glycomics datasets stored in glycomics_data_loader by introducing the b1-? --> b1-3/4 narrow linkage ambiguities (9eeaa3a, 436bf09)
  • download_model will now download model weights and representations from the HuggingFace Hub (22f6b8f)
  • df_species and df_glycan are now of type GlycoDataFrame; build_custom_df now returns a dataframe of type GlycoDataFrame (9764b3e)
  • DataFrameSerializer will now also correctly serialize cells in which (i) lists of strings or (ii) dictionaries have been converted into one string (Excel/pandas interplay of complex cells), where we use ast to try to literally evaluate them back into lists of strings (i) / dictionaries (ii) (806a47c, e302075)

stats

Fixed 🐛
  • Fixed a DeprecationWarning about implicit indexing in alr_transformation when a dict is used for custom_scale (9bf18f7)

motif

processing

Added ✨
  • GlyTouCanIDs are now another supported nomenclature in the context of Universal Input and can be used as inputs for functions etc, supported via improvements in canonicalize_iupac (eafb218)
  • Added sanitize_iupac to detect and fix chemical impossibilities (like two monosaccharides connected via the same hydroxyl group) and fix it (407cd6f, 74d35a0)
  • Added GLYCAN_MAPPINGS dictionary to map commonly used glycan names to their IUPAC-condensed sequence (36d33b8)
  • Added linearcode1d_to_iupac to support sequences of type 01Y41Y41M(31M21M21M)61M(31M21M)61M21M in the Universal Input platform (d0eee40)
  • CSDB linear code is now another supported nomenclature in the context of Universal Input and can be used as inputs for functions etc, supported via improvements in canonicalize_iupac (8dd34b7, 36d2a61, 69c00e1, 2d8fdfd, cb97593, 0e07c56)
  • Added transform_repeat_glycan to support bringing repeat structures of type 1)Fruf(b2-3)Fruf(b2- into the glycowork format of Fruf(b2-3)Fruf(b2-1)Fruf (36d2a61, 2d8fdfd)
  • Added nglycan_stub_to_iupac to support sequences of type (Hex)3 (HexNAc)1 (NeuAc)1 + (Man)3(GlcNAc)2 in the Universal Input platform (69c00e1)
  • Added iupac_to_smiles alias for IUPAC_to_SMILES (cb97593)
  • Added GAG_disaccharide_to_iupac to support disaccharide structural code (DSC) for GAGs (e.g., D2A6) in the context of Universal Input (0770bcd)
  • Added more WURCS tokens for better support in the context of Universal Input, now stored in wurcs_tokens.json (436bf09, 84c5bcc, b30553f, b94cf6d, d1fd4c7, 14bbd4d, a109176)
  • Support monosaccharides without anomeric indicator and phospho-linkages in WURCS (14bbd4d)
Changed 🔄
  • Moved .motif.query.glytoucan_to_glycan into .motif.processing (eafb218)
  • canonicalize_iupac will now use sanitize_iupac to auto-fix chemical impossibilities in input glycans (407cd6f)
  • More GlycoWorkBench sequence variants can now be handled via glycoworkbench_to_iupac/canonicalize_iupac (9eeaa3a, 436bf09, 74d35a0, 87fd540)
  • canonicalize_iupac and most glycowork functions now also support common names, like "LacNAc" or "2'-FL", in the Universal Input framework, thanks to GLYCAN_MAPPINGS (36d33b8, ab42dbb)
  • get_class can now identify repeating unit glycans and returns "repeat" in this case (74d35a0)
  • canonicalize_iupac can now handle even more IUPAC-dialects, like aMan13(aMan16)Man, where the anomeric state is declared before the monosaccharide (24c8e81, ab42dbb)
  • canonicalize_iupac will now use glycan_to_nxGraph and graph_to_string for branch canonicalization, instead of choose_correct_isoform. On average, this works much better and is more reliable (7c52a0e)
  • canonicalize_iupac is now more robust to (5-6) type linkages and to the associated sugar alcohols, like Rib5P-ol (7a260ac)
  • canonicalize_iupac will now raise a ValueError instead of a warning if a glycan string has mismatching brackets (b69fced)
  • canonicalize_iupac can now handle even more IUPAC-dialects such as Neu5Ac-α-2,6-Gal-β-1,3-GlcNAc-β-Sp (cb2c898)
  • canonicalize_iupac can now handle α,β before linkage parentheses (70b2f61)
  • get_class will now correctly annotate plant N-glycans with core a1-3 Fuc (8dd34b7)
  • Rare GLYCAM variants without "-OH" at the end can now also be handled by glycam_to_iupac (207a050)
  • Support single-monosaccharide glycans in GlycoCT within glycoct_to_iupac (87fd540)
  • Support variant sulfate notations in oxford_to_iupac (b35fc0e)
  • Improved parsing of Sialic acid linkage specification in oxford_to_iupac (06ea51f)
  • Added Oxford preferred antenna parsing in oxford_to_iupac (013456f)
  • Added Sialic acid Acetyl modification parsing in oxford_to_iupac (c402bf2)
  • enabled usage of single strings, next to lists, in iupac_to_smiles (8c5aa64)
  • glycam_to_iupac can now handle KDN tokens and more exotic modifications (8c5aa64)
  • iupac_to_smiles can now auto-use Universal Input, if used with a single-string input
Deprecated ⚠️
  • Deprecated find_isomorphs and choose_correct_isoform; this will be done (and better) by the new canonicalize_glycan_graph instead (7c52a0e)

annotate

Changed 🔄
  • Renamed clean_up_heatmap to deduplicate_motifs (407cd6f)
  • Allow sets of glycans as inputs in get_k_saccharides, in addition to lists of glycans (74d35a0)
  • Made get_k_saccharides faster by re-using graphs and using the directed graphs in an optimized way (7c52a0e)
  • get_terminal_structures will now return an actual ValueError when setting size to be higher than 2 (fa451ba)
Fixed 🐛
  • Fixed an edge case in get_k_saccharides, in which choosing a size larger than the size of the largest glycan in the input caused an error (db7847d)
  • Fixed get_k_saccharides with higher values of size, which occasionally produced invalid strings, by refactoring count_unique_subgraphs_of_size_k and switching it to use the changed graph_to_string_int, to ensure motif validity (db7847d)
  • Fixed preprocess_data, which was attempting to transform 0-containing dataframes when no transform argument was provided (878701a)
  • Fixed an issue in get_molecular_properties in which failed requests with placeholder set to False could lead to a size mismatch in preparing the output dataframe (106d0b0)
Deprecated ⚠️
  • Deprecated link_find; will be done by an optimized get_k_saccharides instead (since link_find relied on find_isomorphs) (7c52a0e)

draw

Added ✨
  • Added get_branches_from_graph to process directed glycan graphs into components for GlycoDraw (e56d015)
  • Added the reverse_highlight keyword argument to GlycoDraw, if you want to highlight everything except a certain motif (which means you can highlight discontiguous sequence stretches) (f5e3b2f)
  • GlycoDraw will now inject ALT text / metadata into all its outputs (displayed or saved as .pdf/.svg/.png) for improved accessibility and to aid curation efforts. The ALT text will be automatically generated and includes appropriate tags, the glycan sequence, and used drawing options. But it can also be overriden, if desired, via the new alt_text keyword argument in GlycoDraw (793e71f)
Changed 🔄
  • Quantitative highlighting in GlycoDraw via the per_residue keyword argument will now use individual SNFG-colors instead of a uniform highlight color (07c9c12)
  • Refactored get_coordinates_and_labels to be more efficient and generalizable; with this and the new get_branches_from_graph, GlycoDraw is now capable of drawing even more complex structures accurately (e56d015, 36fbba9)
  • Next to .svg and .pdf, it is now also possible to save .png files with GlycoDraw (36fb...
Read more