An authoritative technical breakdown of the semaglutide peptide sequence, primary amino acid modifications, and side-chain engineering designed for GLP-1 receptor binding research. Learn how specific structural alterations resist enzymatic degradation and extend half-life in laboratory models.
An authoritative technical breakdown of the semaglutide peptide sequence, primary amino acid modifications, and side-chain engineering designed for GLP-1 receptor binding research. Learn how specific structural alterations resist enzymatic degradation and extend half-life in laboratory models.
The semaglutide peptide sequence is a synthetic 31-amino-acid peptide derivative modeled after native human glucagon-like peptide-1, specifically GLP-1(7-37). Its single-letter amino acid notation with modified residue nomenclature is represented as H-Aib-EGTFTSDVSSYLEGQAA-K(AEEA-AEEA-γ-Glu-17-carboxyheptadecanoyl)-EFIAWLVRGRG. In standard three-letter amino acid code, the sequence reads: His-Aib-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys(linker-fatty acid)-Glu-Phe-Ile-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly.
Synthesized for high molecular stability, the chemical formula of semaglutide is C187H291N45O59, yielding a calculated monoisotopic molecular weight of approximately 4113.58 g/mol. Unlike endogenously produced incretins that undergo rapid renal clearance and enzymatic inactivation, the specific modifications engineered into the semaglutide primary sequence fundamentally alter its pharmacokinetic profile in animal models and isolated cell cultures. Researchers evaluating GLP-1 mimetics rely on precise sequence verification to ensure target receptor selectivity and reproducible binding kinetics across assay platforms.
In endogenous GLP-1(7-37), the second amino acid residue is L-alanine. Dipeptidyl peptidase-4 (DPP-4) is a ubiquitous serine exopeptidase that specifically cleaves dipeptides from the N-terminus of proteins containing proline or alanine at the second position. In biological matrices, native GLP-1 experiences enzymatic cleavage at the Ala8-Glu9 peptide bond within 1.5 to 2 minutes, rendering the truncated fragment functionally inactive at the receptor target.
To prevent rapid proteolytic cleavage in preclinical assay systems, the alanine residue at position 8 (referred to as position 2 in the truncated active sequence) is substituted with alpha-aminoisobutyric acid (Aib), an unnatural non-proteinogenic amino acid featuring two methyl groups attached to the alpha-carbon. In vitro stability assays demonstrate that the steric hindrance introduced by Aib prevents DPP-4 recognition and cleavage without compromising the N-terminal interaction required for activating the GLP-1 receptor agonist mechanisms. This structural preservation allows long-duration incubations during cell signaling and reporter gene experiments.
Beyond N-terminal protection, the semaglutide peptide sequence contains a critical side-chain modification at Lysine-26 (Lys26). Native GLP-1 contains a lysine residue at position 26 that is modified in semaglutide through covalent conjugation to a hydrophobic moiety via a multi-component synthetic hydrophilic spacer.
The complete side-chain structure attached to the epsilon-amino group of Lys26 consists of a bis-aminodiethoxyacetyl linker (AEEA-AEEA), followed by a gamma-glutamic acid (γ-Glu) spacer, anchored to a terminal C18 fatty diacid (17-carboxyheptadecanoyl / octadecanedioic acid). Preclinical binding studies confirm that this extended C18 diacid chain facilitates reversible, high-affinity non-covalent binding to serum albumin in rodent and non-human primate plasma. This albumin-binding mechanism reduces free renal filtration, prolongs plasma half-life in vivo, and stabilizes the tertiary peptide conformation against peripheral endopeptidase degradation.
Evaluating the structural evolution of therapeutic incretin candidates requires comparative analysis of amino acid substitutions, acylation lengths, and target receptor cross-reactivity. While earlier analogs like liraglutide utilized a single C16 palmitoyl fatty acid chain attached to Lys26 via a simple γ-Glu spacer without N-terminal Aib substitution, semaglutide incorporates the optimized C18 diacid and double AEEA linker alongside Aib2 protection.
More recent multi-receptor research molecules expand on this framework: tirzepatide incorporates a dual GIP/GLP-1 sequence featuring a C20 fatty diacid attached to a lysine residue via a similar di-Glu spacer, while retatrutide introduces additional sequence variations that enable triple agonist affinity across GLP-1, GIP, and glucagon receptors (GCGR). Understanding these subtle sequence variations is vital for laboratory teams comparative modeling of receptor internalization, cAMP accumulation, and signal transduction cascades in vitro.
In analytical chemistry workflows, verifying the purity and precise molecular weight of the semaglutide peptide sequence requires high-performance liquid chromatography (HPLC) paired with electrospray ionization mass spectrometry (ESI-MS). The presence of the lipophilic C18 side chain alters liquid chromatography retention times relative to non-acylated peptides, requiring specialized reverse-phase C4 or C18 columns with organic solvent gradients (such as acetonitrile with 0.1% trifluoroacetic acid).
When analyzed via ESI-MS, semaglutide characteristically displays multiple protonated charge states ([M+3H]3+, [M+4H]4+, and [M+5H]5+) corresponding to its molecular mass of ~4113.58 Da. Structural verification protocols must confirm the absence of truncated deletion sequences (e.g., missing Aib or single amino acid omissions) and unacylated Lys26 intermediates, both of which can distort preclinical binding affinity assays and skew cell culture data.
Preclinical cell-based assays using Chinese Hamster Ovary (CHO) or HEK293 cell lines expressing human or rodent GLP-1 receptors reveal that the semaglutide peptide sequence exhibits nanomolar binding affinity (Ki) for GLP-1R. Despite the bulky C18 side chain, the N-terminal region maintains high potency for orthosteric receptor activation.
Upon receptor binding, semaglutide stimulates adenylate cyclase via Gs protein coupling, triggering intracellular cyclic adenosine monophosphate (cAMP) accumulation and downstream activation of protein kinase A (PKA) and exchange protein directly activated by cAMP (EPAC2). In primary rodent pancreatic islet models, this signaling cascade promotes glucose-dependent insulin secretion studies and beta-cell survival pathways. Investigators measuring gene expression profiles or phosphorylated protein targets rely on reference-standard peptide sequences to eliminate batch-to-batch operational variance.
Proper handling and dissolution of lipophilic acylated peptides are necessary to prevent aggregation and preserve peptide integrity during experimental protocols. Lyophilized semaglutide should be stored long-term at -20°C or -80°C in a desiccated container protected from light exposure.
When preparing solutions for laboratory assays, researchers should follow established guidelines detailed in our semaglutide reconstitution guide. Reconstitution typically involves sterile bacteriostatic water or phosphate-buffered saline (PBS, pH 7.4). Due to the hydrophobic nature of the C18 fatty acid, initial wetting with sterile buffer followed by gentle swirler agitation—avoiding vigorous vortexing—ensures complete dissolution without generating shear-induced aggregates. Reconstituted aliquots should be used promptly or frozen once to prevent freeze-thaw degradation cycles.
Achieving reproducible experimental results requires working with research peptides manufactured under rigorous quality control standards. Impurities such as TFA salts, residual synthesis solvents, incomplete sequence truncations, or bacterial endotoxins can induce cytotoxic effects in cell culture or artifactual immune responses in animal models.
PX1 Research supplies USA-manufactured research peptides subjected to rigorous third-party testing. Every batch undergoes high-performance liquid chromatography (HPLC) to verify >99% purity and ESI-MS to confirm absolute sequence identity. Furthermore, lot-specific Certificate of Analysis (COA) documents detail chromogenic Limulus Amebocyte Lysate (LAL) endotoxin testing (<0.01 EU/mg) and mass spectral data. Explore our complete catalog of compounds at /research-peptides/all-peptides or review institutional purchasing options via our /wholesale portal and access open technical resources at the PX1 /research center.
What is the exact amino acid sequence of semaglutide?
Semaglutide is a 31-amino-acid peptide with the sequence His-Aib-Glu-Gly-Thr-Phe-Thr-Ser-Asp-Val-Ser-Ser-Tyr-Leu-Glu-Gly-Gln-Ala-Ala-Lys(AEEA-AEEA-γ-Glu-17-carboxyheptadecanoyl)-Glu-Phe-Ile-Ala-Trp-Leu-Val-Arg-Gly-Arg-Gly.
What is the molecular weight and formula of the semaglutide peptide?
Semaglutide has a chemical formula of C187H291N45O59 and a molecular weight of approximately 4113.58 g/mol.
Why is 2-aminoisobutyric acid (Aib) incorporated into the sequence?
Aib is substituted for L-alanine at position 2 (position 8 relative to full GLP-1) to create steric hindrance that prevents enzymatic cleavage by Dipeptidyl Peptidase-4 (DPP-4), significantly increasing peptide stability in laboratory models.
What is attached to Lysine-26 in the semaglutide structure?
Lysine-26 is conjugated to a synthetic spacer containing two AEEA (amino-ethoxy-ethoxy-acetic acid) groups, a gamma-glutamic acid linker, and a terminal C18 fatty diacid (octadecanedioic acid), which facilitates albumin binding.
How is the semaglutide peptide sequence verified analytically?
Sequence identity and purity are verified using Reverse-Phase High-Performance Liquid Chromatography (RP-HPLC) for chemical purity (>99%) and Electrospray Ionization Mass Spectrometry (ESI-MS) for molecular mass confirmation.
What are the recommended storage conditions for lyophilized semaglutide?
Lyophilized semaglutide research powder should be stored sealed at -20°C or -80°C away from moisture and light. Reconstituted solutions should be aliquoted and kept frozen to avoid repeated freeze-thaw cycles.
How does semaglutide differ structurally from tirzepatide?
Semaglutide is a selective GLP-1 mono-agonist featuring a 31-amino-acid sequence with a C18 fatty acid chain at Lys26. Tirzepatide is a dual GIP/GLP-1 receptor agonist featuring a 39-amino-acid sequence modified with a C20 fatty acid chain.
Are PX1 Research peptides tested for endotoxins?
Yes. Every lot supplied by PX1 Research undergoes third-party LAL chromogenic endotoxin testing to guarantee levels strictly below <0.01 EU/mg, alongside HPLC and MS verification documented on the lot COA.
All products are sold strictly for laboratory and research use only. Not for human or veterinary use, diagnosis, treatment or consumption. Statements have not been evaluated by the FDA.