SURPASS-3 extension data: what happens after the trial stops
A design note rather than a result: what the comparator was, and what that permits you to conclude.
TheCompound Journal
Reporting on incretins, compounding & the peptide supply chain
Interruption
The schedules in circulation, described accurately, with their evidentiary status attached.
The word microdosing has migrated into this field from an unrelated context and carries an implication it has not earned: that a small dose produces a distinct and desirable class of effect. For these agents there is no evidence of any such thing. A low dose produces a smaller version of the same effect, with the caveat that the dose-response relationships for weight, for glycaemia and for gastrointestinal tolerability have different shapes, so the ratio of benefit to burden does change across the range. That is worth knowing and it is not what the borrowed word implies.
The parent programmes establish the losses from which the withdrawal arms fall: approximately 14.9 per cent at sixty-eight weeks for semaglutide 2.4 mg in adults without diabetes, approximately 20.9 per cent at seventy-two weeks for tirzepatide 15 mg, and — the only continuous two-year comparator anybody has — approximately 15.2 per cent sustained at week 104 with treatment maintained throughout.12
Read together, the withdrawal evidence supports four statements and does not support a fifth. Regain begins promptly after cessation, within weeks rather than months. It proceeds at a decelerating rate, with the steepest portion in the first three to six months. It does not, within twelve months of follow-up, return participants fully to their original baseline; residual reductions of roughly five to ten per cent persist at one year in all three datasets. And continued treatment maintains and usually extends the loss, with the extension diminishing as the plateau is approached.
The statement not supported is that the drugs cause weight regain, or that stopping leaves a person worse off than never having started. Nothing in these datasets shows overshoot above the original baseline at a group level. Every arm that stopped remained below where it began at the end of follow-up.
The Journal makes this point repeatedly because the contrary claim circulates widely and is often accompanied by a mechanistic story about metabolic damage. The withdrawal trials are the direct test of that claim and they do not support it. What they do support is the unremarkable proposition that a treatment for a chronic condition works while it is being taken.
The homeostatic response to weight loss was characterised in a study that predates this drug class and remains the clearest account of it. After a substantial diet-induced reduction, circulating leptin fell and remained suppressed, ghrelin rose and remained elevated, satiety hormones including peptide YY and cholecystokinin fell, and subjective hunger was significantly greater than baseline — and all of these persisted at twelve months, long after active dieting had ended.3
Alongside the hormonal changes sits a reduction in resting energy expenditure larger than the loss of metabolically active tissue predicts, an effect usually termed adaptive thermogenesis. Its magnitude and persistence are contested, but in the most extreme documented cohort — participants in a televised competitive weight-loss programme followed for six years — resting metabolic rate remained substantially suppressed relative to prediction long after most of the weight had returned.4
Both findings explain regain after drug withdrawal without invoking anything specific to incretins. A drug that suppresses appetite is holding a system away from a defended state. Remove the drug and the system, which has been signalling for restoration throughout, gets what it has been asking for. That is not a drug effect; it is the condition the drug was treating becoming visible again.
Nothing in these datasets shows a group overshooting its original baseline. Every arm that stopped ended below where it began.
On the metabolic-damage claimMean regain trajectories conceal a distribution as wide as the one that characterises the initial response. In the reported withdrawal arms, some participants returned to within a percentage point of their original baseline within a year while others retained most of their loss with no pharmacological support whatever. The interquartile ranges in the published figures are broad, and the trials were not designed to explain them.
Nothing measured at randomisation predicts an individual regain trajectory usefully. Baseline body mass index, the magnitude of the initial loss, age, sex and diabetes status all shift the mean modestly and leave the spread largely intact. The behavioural variables that plausibly matter — what a person was eating and doing during the loss, and whether any of it persisted — were not measured with the granularity required to test them.
This is the same structural gap that runs through the whole field: the mean effects are well characterised and the variance is not. It has a specific practical consequence here. A person deciding whether to stop cannot be given a personal probability of holding their weight, because no such probability has been established, and any source offering one is offering a group mean dressed as a prediction.
| Reason | Randomised evidence on outcome | Typical notice | Resumption likely? |
|---|---|---|---|
| Protocol-driven withdrawal | Three designs | Planned | Not applicable |
| Reached target weight | None | Planned | Sometimes |
| Intolerable side effects | Discontinuation rates only | Days | Sometimes, lower dose |
| Cost or coverage loss | None | Weeks or none | Often, when coverage returns |
| Supply interruption | None | None | Usually, at reset tolerability |
| Discontinuation rates for adverse events are reported in every pivotal trial; outcomes after discontinuation for the other reasons are not, because the trials did not enrol people who stopped for them. | |||
This section is short because the evidence is. The Journal has searched the trial registries and the published literature for any randomised comparison of an intermittent schedule against a standard weekly schedule for any GLP-1 receptor agonist or dual agonist, at any dose, for any indication. We have found none. We have also found no observational cohort large enough to characterise outcomes on such a schedule with the standard confounders addressed.
What exists is dose-ranging data from the phase 2 programmes, which establishes that lower average exposures produce smaller weight effects, and pharmacokinetic modelling, which establishes what average exposure and what peak-to-trough ratio a given interval would produce. Neither tells you whether a fortnightly schedule maintains weight in somebody who has already lost it, which is the question actually being asked.
An absence of evidence is not evidence of harm and the Journal does not present it as such. It is, however, the entire evidentiary position, and readers encountering confident protocols for intermittent use should know that the confidence is not coming from data. Nothing in this section is advice, and the compounds sold for research use only that appear in some of these protocols are not approved for human use.
For a drug eliminated with first-order kinetics, the accumulation ratio at steady state is approximately one divided by one minus the exponential of minus the elimination rate constant times the dosing interval. For a seven-day half-life given weekly, that yields a ratio of about two. Given fortnightly, the interval is two half-lives, the residual fraction at the next dose is a quarter, and the accumulation ratio falls to about one and a third.
Two consequences follow. Average concentration on a fortnightly schedule at the same nominal dose is roughly a third lower than weekly, not half, because accumulation differs. And the peak-to-trough ratio rises from modest to fourfold, so the exposure pattern is qualitatively different: the person spends part of each cycle at an exposure that would be sub-therapeutic on a weekly schedule and part at a higher peak.
What the arithmetic cannot tell you is whether that pattern is better, worse or equivalent for maintaining weight, because the relationship between exposure pattern and weight effect is not known — only the relationship between average exposure and weight effect at steady state. The modelling is solid and it answers a different question from the one people bring to it.
This is the most practically consequential item in the whole subject and the one least often stated in advance. Gastrointestinal tolerability to these agents develops over weeks of continued exposure and decays when exposure is removed. After four weeks without the drug, plasma concentrations are a small fraction of steady state and the tolerability accommodation has substantially reset. Resuming at the previous maintenance dose therefore presents the system with an exposure step it has not experienced for a month.
The clinical convention — resume at a lower dose and re-escalate — follows from the pharmacokinetics rather than from caution.5 Product labelling for several agents in the class advises consideration of re-initiation at a lower dose after an extended interruption, and the threshold at which this applies differs between products, which is a detail worth checking against the specific label rather than a general rule.
The shortage period demonstrated the consequence of ignoring this at scale. Large numbers of people lost access for six to ten weeks, resumed where they had left off, and experienced nausea and vomiting considerably worse than during their original escalation. It was predictable, it was predicted by anybody who had read the label carefully, and it was almost never communicated.
Every trial in this class delivers a behavioural intervention alongside the drug: energy-restriction targets, activity targets, and regular contact with a study team. That contact is itself an intervention of measurable effect, which is why placebo arms in these programmes lose two to three per cent of body weight rather than nothing. Where the behavioural component was deliberately intensified, the placebo arm lost around 5.7 per cent over sixty-eight weeks, which is a useful upper bound on what contact and counselling alone achieved in these populations.6
It matters for the withdrawal question in a way that is usually elided. The semaglutide off-treatment extension withdrew the drug and the lifestyle support together, so its regain figure describes the removal of a package.7 The STEP 4 and SURMOUNT-4 withdrawal arms kept the lifestyle component running, so their regain figures describe the removal of a molecule with support maintained.58 Those are different experiments and the second is the more conservative.
Anybody comparing regain figures across the three should therefore expect the extension to look worse, and it does. The Journal states which withdrawal design a figure comes from every time it quotes one, because the alternative is pooling two different experiments into a single number that describes neither. The same caution applies to the frequent comparison with dietary weight-loss regain, where the behavioural intervention is the whole of the treatment.
A seven-day half-life tapers itself. What a taper buys is behavioural, and it should be argued for on those terms.
On coming offThe Journal’s position is that three trials would resolve almost everything currently argued about in this area, and that all three are straightforward. The first is a dose-reduction design: after a lead-in to target, randomise to full dose, one step down, two steps down, or placebo, and follow for a year with weight as the primary endpoint. It would establish the shape of the descending dose-response curve and would cost a fraction of a pivotal programme.
The second is an interval design: after a lead-in, randomise to weekly, fortnightly and three-weekly administration at the same nominal dose. It would answer the intermittent-schedule question directly and would settle whether the exposure pattern matters independently of average exposure.
The third is a taper design: randomise abrupt cessation against a stepped reduction over twelve weeks, with appetite, eating behaviour and weight measured for a year afterwards. It would test the only argument for tapering that is worth testing.
None of the three is under way as far as the Journal can establish. Readers who know otherwise should write to letters@compoundjournal.com; a registered protocol for any of them would be news in this department.
| Study | Design | Lead-in | Randomised follow-up | Lifestyle support after |
|---|---|---|---|---|
| STEP 1 extension | Off-treatment observation | 68 weeks on drug | 52 weeks off | Withdrawn |
| STEP 4 | Randomised switch to placebo | 20 weeks to 2.4 mg | 48 weeks | Continued |
| SURMOUNT-4 | Randomised switch to placebo | 36 weeks to max tolerated | 52 weeks | Continued |
| S-LiTE | Post-diet maintenance, 4 arms | 8-week low-energy diet | 52 weeks | Continued |
| STEP 5 | Continuous treatment, no withdrawal | — | 104 weeks on drug | Continued |
| The first three are the withdrawal evidence base. STEP 5 is included because it is the only two-year continuous-treatment comparator and is frequently cited alongside the withdrawal data as though it were part of it. | ||||
Four things accompany every regain number in these pages. Which withdrawal design it comes from, because an off-treatment extension and a randomised placebo switch are different experiments. Whether the lifestyle intervention continued in the arm being described. What the denominator is — regain as a percentage of body weight, as a percentage of the weight lost, or as a final position relative to original baseline, three quantities that are routinely quoted interchangeably. And the follow-up duration, because the regain curve decelerates and a figure at six months is not a figure at a year.
The third of those is where most of the misreporting happens. A statement that participants regained two-thirds is a proportion of loss; a statement that they regained eleven per cent is a proportion of body weight; a statement that they finished 5.6 per cent below baseline is a final position. All three can describe the same arm and they are not interchangeable.
Where a source we are quoting has not stated its denominator, we say that rather than inferring it. Readers who find a regain figure in these pages without its design and its denominator have found an error, and the standards desk would like to hear about it at standards@compoundjournal.com.
This is reporting on a body of trial evidence and it is not advice about whether or how to stop taking a medicine. The decision to discontinue an agent prescribed for glycaemic control, cardiovascular risk or kidney disease is materially different from the decision to discontinue one prescribed for weight, and in every case it belongs with a clinician who has seen the person and knows why the drug was started.
Two further notes. Compounds sold for research use only are not approved for human use in any jurisdiction, and nothing here should be read as guidance about using them or about stopping their use. And where this piece describes what clinicians report doing about maintenance dosing, that is description of practice and not a schedule anybody should adopt from a magazine.
The Journal takes correspondence on this subject at letters@compoundjournal.com and factual challenges at standards@compoundjournal.com. Letters describing a personal experience of stopping are read with attention and are published, where they are published, as accounts rather than as evidence — a distinction this department tries hard to preserve in both directions.
What the shortage years demonstrated, at a scale no trial will ever match, is that this class is now being stopped and restarted routinely by circumstance rather than by decision. That is the discontinuation that actually happens, and there is no literature on it at all. The Journal regards documenting it as reporting rather than research, and will keep doing so.
Selected from correspondence received on this article. Writers are identified by initial, surname and city, verified before printing. Replies are from the desk that filed the piece or from the standards editor. Write to letters@compoundjournal.com.
I stopped eight months ago after reaching a weight I was happy with, and I have regained four of the twenty-two kilograms I lost. Every article I read told me to expect two-thirds back. I am not complaining, but I would like to know whether I am unusual or whether the two-thirds figure was always a mean concealing an enormous range.
— R. Whitlam, Adelaide, SA
The second. The published interquartile ranges around those means are wide, and outcomes like yours are well within them. The trials were not designed to explain why some people hold weight after cessation and others do not, and nothing measured at randomisation predicts it usefully. You are not an anomaly; you are part of a distribution nobody quotes.
You describe the maintenance strategy of stepping down one dose and holding for eight to twelve weeks as something clinicians report doing, and then say it is not a recommendation. That distinction will be lost on most readers, and printing the protocol makes you a source for it whether you intend to be or not.
— C. Rautenbach, Pretoria
This is the hardest editorial question this department faces and we do not think you are wrong. Our position is that a practice this widespread is better described accurately, with its evidentiary status stated, than left to circulate in fragments. We accept that the distinction does work that a reader may not do.
Three months after stopping, my HbA1c had barely moved and I concluded I had got away with it. Six months after stopping, it was back where it started. Your point about the lag is the single most useful sentence I have read on this subject.
— D. Mazzarella, Catania
Your fortnightly arithmetic table is correct but I think it understates the practical point. A fourfold peak-to-trough swing is not merely lower average exposure; it is a different drug experience, with the last few days of each cycle spent at a concentration the person has effectively titrated off.
— A. Mbeki, Lusaka
Well put, and better than our own phrasing. We have adopted the point in the text with attribution to a reader.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
Efficacy was never the question in this appraisal. Duration of treatment was.
Real-world persistence figures, with their definitions stated, because the definitions are doing most of the work.
The practice is near-universal, clinically sensible, and supported by observational data rather than randomised comparison. We say which is which.
A design note rather than a result: what the comparator was, and what that permits you to conclude.
A design note rather than a result: what the comparator was, and what that permits you to conclude.