Semaglutide Molecular Weight, Sequence & CAS Reference

This technical reference sheet provides precise biochemical data on semaglutide for laboratory research settings. Below are the verified amino acid sequence, molecular weight, CAS registry details, side-chain acylation specs, and analytical parameters required for quantitative in vitro and preclinical experimentation.

GMP-compliant U.S. facilities
ISO 17025 third-party COAs
100% domestic — no imports
Fast tracked domestic shipping
Shop research peptides

Quick answer

This technical reference sheet provides precise biochemical data on semaglutide for laboratory research settings. Below are the verified amino acid sequence, molecular weight, CAS registry details, side-chain acylation specs, and analytical parameters required for quantitative in vitro and preclinical experimentation.

Reviewed by PX1 Research scientific team

Key takeaways

  • [Semaglutide](/research-peptides/semaglutide) is a long-acting, synthetically modified peptide analog derived from human glucagon-like peptide-1 (GLP-1(7-37)).
  • The empirical molecular formula of [semaglutide](/research-peptides/semaglutide) free base is C187H291N45O59.
  • The primary amino acid sequence of [semaglutide](/research-peptides/semaglutide) consists of 31 amino acid residues modified at three critical positions relative to the endogenous human GLP-1(7-37) sequence.
  • The chemical structure of the Lys26 side chain represents a sophisticated piece of peptide engineering.

1. Biochemical Overview and Primary Specifications

Semaglutide is a long-acting, synthetically modified peptide analog derived from human glucagon-like peptide-1 (GLP-1(7-37)). In laboratory research settings, it serves as a critical reference compound for investigating incretin receptor signaling pathways, G-protein coupled receptor (GPCR) activation, and metabolic peptide dynamics. Understanding the exact semaglutide molecular weight sequence and structural profile is essential for accurate molar calculations, quantitative assays, and reproducible in vitro experimental design.

Unlike native GLP-1, which exhibits a brief biological half-life due to rapid cleavage by dipeptidyl peptidase-4 (DPP-4), semaglutide incorporates specific structural substitutions and a fatty acid side chain. These modifications significantly alter its physicochemical properties, molecular mass, and binding affinity profiles in preclinical models. Research institutions evaluating analytical standards across our catalog of all peptides rely on accurate mass-to-charge ratios and chemical formulas to calibrate mass spectrometers and liquid chromatography systems.

2. Molecular Formula, Mass, and CAS Reference

The empirical molecular formula of semaglutide free base is C187H291N45O59. The theoretical monoisotopic molecular weight is approximately 4111.12 Da, while the average molecular weight evaluated across mass spectrometry protocols is 4113.58 g/mol (4113.6 Da). These values represent the free base peptide prior to salt formation or hydration adjustments.

The Chemical Abstracts Service (CAS) registry number assigned to semaglutide free base is 910463-68-2. In research supply chains and analytical verification logs, this CAS reference distinguishes the specific acylated analog from related native GLP-1 sequences or un-acylated peptide backbones. When ordering custom synthesis or lot-matched batches for comparative preclinical studies, verifying the CAS registry details alongside the lot-specific Certificate of Analysis ensures complete structural identity.

Researchers conducting mass spectroscopic verification must account for the molecular weight contribution of both the primary amino acid backbone and the extended C18 diacid acyl chain connected via a hydrophilic spacer. Discrepancies between observed m/z peaks and theoretical values often stem from counterion adducts or varying ionization states during electrospray ionization (ESI-MS) analysis.

3. Primary Amino Acid Sequence and Modification Mapping

The primary amino acid sequence of semaglutide consists of 31 amino acid residues modified at three critical positions relative to the endogenous human GLP-1(7-37) sequence. The primary sequence using standard three-letter amino acid code nomenclature is:

His-Aib-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys(AEEAc-AEEAc-gamma-Glu-17-carboxyheptadecanoyl)-Glu-Phe-Ile-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly

The sequence modifications are strategically engineered for specific biochemical functions:

• Position 8 (Aib8): Substitution of native Alanine with alpha-aminoisobutyric acid (Aib), a non-proteinogenic amino acid. This substitution introduces steric hindrance that blocks enzymatic cleavage by DPP-4 at the Lys7-Glu9 peptide bond.

• Position 34 (Arg34): Substitution of native Lysine with Arginine. This substitution prevents unintended acylation at position 34 during chemical synthesis, ensuring targeted attachment of the side chain exclusively at position 26.

• Position 26 (Lys26 Acylation): The epsilon-amino group of Lysine-26 is conjugated to a synthetic spacer consisting of two 8-amino-3,6-dioxaoctanoic acid (AEEAc) units linked to a gamma-glutamic acid (gamma-Glu) linker, terminated by a 17-carboxyheptadecanoyl (C18 diacid) fatty acid chain.

4. Side-Chain Chemistry and Albumin Binding Mechanism

The chemical structure of the Lys26 side chain represents a sophisticated piece of peptide engineering. The extended spacer—[AEEAc-AEEAc-gamma-Glu]—provides conformational flexibility and hydrophilic character, allowing the terminal C18 fatty diacid to project outward from the peptide backbone.

In cell culture media, serum-supplemented assays, and preclinical animal models, this C18 diacid group binds non-covalently to site II on serum albumin. Albumin binding creates a reversible circulating reservoir, protecting the peptide from renal clearance and extending its half-life in laboratory models. When performing in vitro receptor binding assays or cell-based cAMP reporter assays, researchers must account for serum albumin concentration in the buffer, as high bovine serum albumin (BSA) levels reduce the free effective peptide concentration available for receptor engagement.

For additional technical documentation on comparative peptide structures and receptor signaling profiles, explore our comprehensive research hub, which details structural kinetics across multiple peptide classes.

5. Counterions, Salt Forms, and Net Peptide Content

Synthetic peptides manufactured via solid-phase peptide synthesis (SPPS) are purified using reverse-phase high-performance liquid chromatography (RP-HPLC) employing mobile phases containing trifluoroacetic acid (TFA). Consequently, research-grade semaglutide is frequently isolated as a trifluoroacetate (TFA) salt, where basic amino acid residues (His1, Lys26, Arg30, Arg32) non-covalently bind TFA counterions.

Alternatively, semaglutide may undergo salt exchange to yield an acetate salt form. The choice of salt form directly impacts total gross mass, water absorption kinetics, and net peptide content (NPC):

• TFA Salt Form: Contains variable TFA counterion content (typically 10–20% by mass). While suitable for general analytical mapping, residual TFA can exhibit cytotoxicity in sensitive primary cell cultures.

• Acetate Salt Form: Produced by exchanging TFA counterions with acetate ions. Preferred for cell-based in vitro assays sensitive to TFA trace residues.

Net Peptide Content (NPC) represents the actual percentage of pure peptide mass within the lyophilized cake relative to counterions (TFA/acetate) and bound residual moisture. Lyophilized semaglutide typical NPC ranges between 75% and 85%. Researchers must adjust gross weight measurements using the lot-specific NPC percentage provided on the PX1 COA documentation before preparing target stock concentrations. To quickly compute molar concentrations accounting for net peptide purity, researchers can utilize the online reconstitution calculator.

6. Analytical Spectroscopic and Chromatography Specifications

To verify batch-to-batch consistency and high structural fidelity, research suppliers utilize rigorous analytical methodologies. PX1 Research subjects all semaglutide lots to dual verification via reverse-phase HPLC and High-Resolution Mass Spectrometry (HRMS) in an ISO 17025 accredited laboratory.

Analytical specifications evaluated prior to release include:

• Purity by RP-HPLC: Greater than 98.0% peak area integration at 214 nm and 280 nm UV detection wavelengths.

• Mass Verification by ESI-MS: Observed mass peak matching the theoretical molecular weight of 4113.6 Da within a tight tolerance (+/- 1.0 Da).

• Residual Solvents: Gas chromatography (GC) analysis ensuring organic solvent levels fall strictly below standard laboratory safety limits.

• Endotoxin Content: Bacterial endotoxin testing via Chromogenic Limulus Amebocyte Lysate (LAL) assay, verifying levels strictly below < 0.05 EU/mg for preclinical research applications.

Details on analytical protocols and full chromatograms can be found on our semaglutide research guide page.

7. Comparative Structural Analysis Across Incretin Peptides

Understanding how semaglutide differs structurally from other metabolic and gastrointestinal research peptides provides valuable context for comparative bioassays. Within the same receptor targeting class, liraglutide possesses a C16 fatty acid attached via a single gamma-Glu spacer at Lys26 and retains native Alanine at position 8, resulting in a lower molecular weight (3751.2 Da) and shorter half-life in research models. Conversely, dual-agonist compounds such as tirzepatide feature a 39-amino-acid sequence with a C20 fatty diacid moiety on a Lys residue, yielding a significantly higher molecular weight (4813.5 Da) and distinct dual GLP-1/GIP receptor affinity profiles.

When expanding metabolic research protocols to include non-incretin secretagogue analogs or dual GLP-1/GLP-2 receptor probes like GLP2-T, cross-referencing molecular weights, salt forms, and side-chain acylation profiles ensures proper concentration normalization across experimental arms.

8. Handling, Reconstitution, and Solution Stability Parameters

Lyophilized semaglutide should be stored at -20°C or -80°C in a desiccated environment protected from light. Under these conditions, the dry peptide powder maintains structural stability for extended periods.

For laboratory reconstitution, researchers should adhere to standard sterile handling protocols:

1. Reconstitution Solvents: Use sterile phosphate-buffered saline (PBS, pH 7.4) or bacteriostatic water containing 0.9% benzyl alcohol depending on assay requirements.

2. Dissolution Technique: Allow the vial to equilibrate to room temperature before adding solvent. Gently swirl or invert the vial. Avoid vigorous vortexing or rapid agitation, as mechanical shear stress can induce peptide aggregation and hydrophobic self-assembly into fibril structures.

3. Solubilization Characteristics: Due to the hydrophobic C18 fatty acid chain, semaglutide exhibits pH-dependent solubility. Optimal solubility occurs in neutral to slightly alkaline aqueous buffers (pH 7.4–8.0). Acidic solutions (pH < 6.0) may cause transient precipitation or reduced solubility.

4. Aliquoting and Storage: Once reconstituted, aliquot the stock solution into single-use polypropylene microtubes to prevent repeated freeze-thaw cycles. Reconstituted aqueous solutions stored at 4°C are stable for up to 28 days when preserved with appropriate bacteriostatic agents, or up to 3 months at -80°C.

For bulk supply inquiries or custom packaging requirements tailored to high-throughput laboratory automation, visit our wholesale portal.

Frequently Asked Questions

What is the exact molecular weight of semaglutide free base?

The theoretical average molecular weight of semaglutide free base is 4113.58 g/mol (4113.6 Da), with an empirical molecular formula of C187H291N45O59.

What is the CAS registry number for semaglutide?

The Chemical Abstracts Service (CAS) registry number for semaglutide free base is 910463-68-2.

Why is alpha-aminoisobutyric acid (Aib) incorporated at position 8?

Aib is substituted for Alanine at position 8 to sterically protect the His7-Glu8 cleavage site against enzymatic degradation by dipeptidyl peptidase-4 (DPP-4) in biological fluids.

How does net peptide content (NPC) affect sample mass calculations?

Net peptide content represents the actual proportion of pure peptide relative to counterions (TFA or acetate) and residual moisture. If a lot has an 80% NPC, 1.0 mg of lyophilized powder contains 0.8 mg of active semaglutide peptide.

Is semaglutide supplied as a TFA or acetate salt?

PX1 Research offers research-grade semaglutide primarily as a TFA salt for general analytical work or as an acetate salt upon request for sensitive in vitro cell culture models.

What analytical methods are used to verify PX1 semaglutide purity?

Every lot undergoes reverse-phase HPLC to verify >98% purity, Electrospray Ionization Mass Spectrometry (ESI-MS) to confirm mass identity, and Chromogenic LAL testing to ensure endotoxin levels remain below < 0.05 EU/mg.

What buffer pH is optimal for semaglutide reconstitution?

Semaglutide demonstrates optimal aqueous solubility in neutral to slightly alkaline buffers (pH 7.4–8.0), such as standard PBS. Acidic conditions below pH 6.0 should be avoided to prevent precipitation.

How does the Lys26 side chain modification alter semaglutide behavior?

The Lys26 side chain consists of two AEEAc hydrophilic spacers, a gamma-Glu linker, and a C18 fatty diacid. This structural assembly enables high-affinity reversible binding to serum albumin, extending half-life in preclinical research models.

Related pages

All products are sold strictly for laboratory and research use only. Not for human or veterinary use, diagnosis, treatment or consumption. Statements have not been evaluated by the FDA.