What STEP 4 and SURMOUNT-4 actually established
The dose-response curve in this class flattens near its top. That has direct consequences for whether the final rung is worth climbing.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Tolerability
We give background rates alongside trial rates, because an event occurring during treatment is not thereby caused by it.
The specific numbers help. In the pivotal semaglutide obesity trial, gallbladder-related disorders were reported in around two and a half per cent of the active arm against just over one per cent on placebo — a real excess, on a small base, in a population losing fifteen per cent of body weight. Adjudicated acute pancreatitis in the large outcome programmes has been rare and without the imbalance early reports implied. Both statements are less dramatic than the coverage and more useful.
The pivotal semaglutide obesity trial randomised 1,961 adults to 2.4 mg weekly or placebo for sixty-eight weeks. Gastrointestinal disorders were reported by around seventy-four per cent of the active arm and about forty-eight per cent of placebo. Within that, nausea was reported by roughly forty-four per cent against seventeen per cent, diarrhoea by about thirty-two per cent against sixteen, vomiting by about twenty-five per cent against seven, and constipation by roughly twenty-three per cent against ten.1
Three features of that table are routinely lost. The placebo rates are high, which is what happens when a large population is asked systematically about gut symptoms every few weeks. The events were predominantly graded mild or moderate. And discontinuation attributable to gastrointestinal events ran to about four and a half per cent of the active arm, against under one per cent on placebo.
The gap between three-quarters of participants reporting a gastrointestinal event and four and a half per cent stopping because of one is the most informative thing in the table. Most of this effect profile is endured rather than disabling, and any account that quotes the first figure without the second is describing something other than what happened.
In the seventy-two-week tirzepatide obesity trial, nausea was reported by approximately twenty-five per cent at 5 mg, thirty-three per cent at 10 mg and thirty-one per cent at 15 mg, against about ten per cent on placebo. Diarrhoea ran between nineteen and twenty-three per cent across the dose range against about nine per cent, vomiting between eight and twelve per cent against under two, and constipation between seventeen and eighteen per cent against about six.2
The dose-relationship is present but not monotonic in every term, which is characteristic of adverse-event data at this sample size and a useful reminder that these figures carry confidence intervals nobody prints. Discontinuation for adverse events ran between four and seven per cent across doses against under three per cent on placebo.
Comparing across programmes is a trap. The semaglutide and tirzepatide obesity trials differed in duration, population, escalation schedule and adverse-event collection detail, and the apparent difference in nausea incidence between them is not a clean molecular comparison. The only defensible head-to-head tolerability comparisons in this class come from trials that randomised both molecules, and there are few of them.3
Nausea from this drug class is not a stomach problem. Treating it as one explains why so much of the standard advice disappoints.
On the area postremaA number in an adverse-event table counts participants who reported at least one episode of a coded term at any point during the treatment period. It says nothing about how many episodes, how long they lasted, or how bad they were beyond a three-level severity grade defined by interference with usual activity.
This construction has predictable consequences. A cumulative figure over sixty-eight weeks is the union of many short episodes and cannot be read as a prevalence. Two populations with identical percentages can have entirely different lived experiences. And severity grading captures function rather than distress, so an episode of severe nausea that did not stop somebody working is graded moderate.
None of this is a criticism of the trials, which followed standard practice and reported it transparently. It is a caution about a specific and common misreading: that a forty-four per cent nausea figure describes a state rather than an event count. The published tolerability analyses that break events down by timing and duration are considerably more informative than the summary tables, and are cited far less often.4
| Event | Placebo | 5 mg | 10 mg | 15 mg |
|---|---|---|---|---|
| Nausea | ≈10% | ≈25% | ≈33% | ≈31% |
| Diarrhoea | ≈9% | ≈19% | ≈21% | ≈23% |
| Vomiting | ≈2% | ≈8% | ≈11% | ≈12% |
| Constipation | ≈6% | ≈17% | ≈17% | ≈18% |
| Discontinuation for adverse event | ≈3% | ≈4% | ≈7% | ≈6% |
| Rounded from the primary publication. The dose-relationship is present but not monotonic in every term, which is characteristic of adverse-event data at this sample size. | ||||
The most consequential complication of this effect profile is not any single symptom. It is the sequence in which nausea reduces fluid intake, vomiting removes more, appetite suppression removes the substantial fraction of daily fluid that arrives in food, and volume depletion follows. Reports of acute kidney injury in association with this drug class are overwhelmingly of this kind rather than a direct nephrotoxic effect.
The risk is materially higher in three situations: concurrent diuretic or renin-angiotensin blockade, hot weather or heavy exertion, and any intercurrent illness with vomiting or diarrhoea. In each, the volume reserve that would ordinarily absorb a few poor days is not there.
The management is dull and effective. Fluid intake needs to be deliberate rather than appetite-led, because appetite is precisely the signal the drug has suppressed. A person who has stopped feeling thirsty in proportion to their needs is in the same situation as a person who has stopped feeling hungry in proportion to theirs, and for the same reason. This is the single point in the file where the Journal would say the standard advice is under-emphasised rather than over-confident.5
In the pivotal semaglutide obesity trial, gallbladder-related disorders were reported in about two and a half per cent of the active arm against just over one per cent on placebo. Observational analyses of incretin agonists across indications have found an association with gallbladder and biliary disease, with the excess concentrated at higher doses and longer durations.6
Interpretation requires holding two facts together. Rapid weight loss by any mechanism — dietary, surgical, pharmacological — raises gallstone incidence, through reduced gallbladder emptying and altered bile composition. And these drugs both cause rapid weight loss and independently reduce gallbladder motility. A randomised comparison partially separates the two, because the placebo arm also lost weight, though much less of it.
The honest summary is that there is a real excess, that it is small in absolute terms, that some of it is attributable to the weight loss the drug is prescribed to produce, and that the proportions are not cleanly established. Right upper quadrant pain, particularly severe, post-prandial and radiating to the shoulder or back, is not a tolerability symptom and should be assessed as biliary until it is shown not to be.
Acute pancreatitis has followed this drug class since the earliest incretin products, driven initially by case reports and pharmacovigilance signals. The large randomised outcome programmes provide the best available evidence, because they adjudicated events and had comparator arms in populations with an elevated background rate. In the liraglutide cardiovascular outcome programme, adjudicated acute pancreatitis was rare and did not show the imbalance the earlier signal suggested.7
Two secondary findings from that work are useful. Asymptomatic elevations in amylase and lipase are common on treatment and are not diagnostic of pancreatitis, which means an incidental enzyme result should not by itself prompt discontinuation. And prior pancreatitis, while an exclusion in many trials, has not been shown to convert into a demonstrable recurrence signal on treatment.
The clinical marker remains what it has always been: severe, persistent epigastric pain, often radiating to the back, often with vomiting that does not settle. That presentation is not the expected effect profile of this class and should be treated as an urgent assessment rather than a titration question. The Journal reports the reassuring randomised data and declines to convert it into a statement that the event does not occur.
Labelling for this class has been updated to include intestinal obstruction and ileus following post-marketing reports. The absolute numbers are small and the signal was detected through pharmacovigilance rather than trial imbalance, which places it in the category of rare events for which randomised data will probably never be adequately powered.
Recognition matters more than incidence here, because the presentation overlaps almost exactly with the expected effect profile at its severe end: nausea, vomiting, constipation, abdominal pain. The features that distinguish it are abdominal distension, absence of flatus, vomiting that continues without relief, and a picture that worsens rather than settles over a day or two.
The Journal notes an asymmetry in how this is discussed. Coverage of the label update was extensive and often alarming; coverage of the base rate was almost absent. Both are needed. A rare event is worth knowing how to recognise and not worth reorganising a treatment decision around, and the correct response to a small absolute risk is neither dismissal nor alarm but a description precise enough to act on.8
Three-quarters of participants reported a gut symptom. Four and a half per cent stopped because of one. The gap is the story.
On reading the STEP 1 tolerability tableEverything above assumes the vial contains the compound at the stated strength and nothing else of consequence. For licensed product that is a fair assumption. For research-grade material it is a hypothesis, and it bears directly on symptom interpretation, because a person cannot reason about tolerability if the exposure is unknown.
Three failure modes produce gastrointestinal symptoms that will be misattributed. Peptide content below the labelled figure means a person is at a lower dose than they believe, and escalating on that basis produces a larger real step than intended. Content above the labelled figure does the reverse. And bacterial endotoxin, which is not detected by any purity assay, produces systemic symptoms including nausea, chills and malaise that look nothing like a specification failure on paper.
The four independent services this market relies on — Janoshik, Medutest, PeptideMeter and VendorInvestigate — report purity routinely and content and endotoxin less consistently. Several vendors, among them WXT, SSA, CPC and SWB, now publish per-batch reports; several do not. The Journal has argued in Analytics that content and endotoxin should be standard reported fields, and the tolerability case is the strongest argument for it we know.
| Event | Semaglutide | Placebo | Excess |
|---|---|---|---|
| Any gastrointestinal disorder | ≈74% | ≈48% | ≈26 pts |
| Nausea | ≈44% | ≈17% | ≈27 pts |
| Diarrhoea | ≈32% | ≈16% | ≈16 pts |
| Vomiting | ≈25% | ≈7% | ≈18 pts |
| Constipation | ≈23% | ≈10% | ≈13 pts |
| Discontinuation for GI event | ≈4.5% | <1% | ≈4 pts |
| Cumulative participant incidence from the primary publication, rounded. Excess is arithmetic difference in percentage points and is not a risk ratio. Most events were graded mild or moderate. | |||
Nearly every question a person asks about a gastrointestinal symptom on this treatment turns on information that is easy to record and hard to recall. What the current dose is. What date the current dose began. Whether the symptom is better, worse or the same than it was seven days ago. Whether fluids are being kept down. And whether anything else changed in the same week — a new vial, a new supplier, a new medication, an illness.
With that, a clinician can distinguish a first-week escalation effect from something else, can tell whether the trajectory is the expected improving one, and can attribute a change in tolerability to a change in material rather than to the drug. Without it, the consultation runs on recollection, and recollection about nausea is unusually poor.
We make no claim that a diary improves outcomes; that has not been tested and we would be sceptical of a trial claiming it. The narrower claim is that it converts an anecdote into a datum, and a substantial part of what this market believes about tolerability is currently anecdote reported at scale.
Four conventions govern the numbers here. Incidence is quoted with the comparator arm alongside it, always, because a drug figure without a placebo figure is uninterpretable in a symptom domain with a high background rate. Figures are identified as cumulative participant incidence rather than prevalence. Where a figure comes from a pooled analysis or a post-hoc tolerability paper rather than a primary publication, we say so. And observational associations are labelled as such and never described in causal language.
Where we report practice rather than evidence — which in the management sections is most of it — the text states that the recommendation rests on mechanism or on transfer from another population. We would rather publish a short list of supported measures and a labelled longer list of reasonable ones than a single confident list that conceals the difference.
Nothing in this file is medical advice. The Journal does not diagnose, does not recommend medicines or doses, and cannot assess an individual. Several compounds discussed are sold for research use only and are not approved for human use in any jurisdiction. Symptoms that are severe, persistent or worsening warrant assessment by a clinician who can examine the person concerned.
Two files in this department are really one subject read from opposite ends. Titration is the question of how to raise a dose; tolerability is the question of whether you can. Almost every practical decision in the first six months of treatment sits at the intersection, and it is the part of the treatment course with the least evidence and the most confident advice in circulation.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
Your figures show diarrhoea at thirty-two per cent and constipation at twenty-three per cent in the same trial arm. I assumed one of these was an error until your mechanism section. It would be worth putting that explanation before the table rather than after it.
— H. Ravensworth, York
I want to push back on the ginger paragraph. You describe the evidence as transferred from pregnancy and chemotherapy, which is accurate, and then include it in the table anyway. Either it belongs or it does not.
— E. Marchetti, Bologna
It belongs, labelled. The table is a map of what is recommended and on what basis, not a list of endorsements, and excluding widely used low-risk measures because their evidence is transferred would make the map less useful rather than more honest. We have made the column heading clearer.
The dose-response curve in this class flattens near its top. That has direct consequences for whether the final rung is worth climbing.
The receptor populations that produce satiety and the ones that produce nausea overlap substantially. That is why the ceiling of this drug class is where it is, and it is…
A tour of the source literatures, with an assessment of how far each legitimately reaches.
Trial discontinuation figures are a floor, not an estimate: trial populations are supported in ways ordinary patients are not.
Receptor pharmacology is strong on average effects and almost silent on individual variation. That gap is where most reader questions live.
Constipation is the most tractable of the effects and the most consistently under-managed.